Overview
What this challenge is about.
Design a multi-region failover for an enterprise RAG service using Terraform, run a chaos test, and measure RTO/RPO. Earn a verifiable certificate.
The scenario
The startup (around 55 staff) bills around USD 6M ARR with 8 anchor customers; a single multi-hour outage could trigger an SLA refund and a churn-conversation cascade.
The Brief
What you'll do, and what you'll demonstrate.
Design and prove a multi-region failover for a RAG service that achieves a measured RTO under 30 minutes and RPO under 5 minutes.
Earning criteria — what you'll demonstrate
- Design a multi-region active-passive architecture for an AI service
- Reason about RTO/RPO targets for stateful AI workloads (vector stores)
- Run a chaos test that proves the design under failure
- Author an SRE runbook that holds up at 3am
Program Fit
Where this fits in your program.
Sharpens the same skills your degree expects you to demonstrate.
Cloud Computing for Data and ML
Master · Data Engineering
Strong alignment
This challenge maps to Cloud Computing for Data and ML at the Master level. It sharpens the same practical skills your coursework expects — but in a real industry context with actual constraints and deliverables.
Skills
Skills you'll demonstrate.
Each one shows up on your verified credential.
- Multi Region Architecture
Apply multi region architecture to solve real industry problems and demonstrate production-level capability.
- Disaster Recovery
Apply disaster recovery to solve real industry problems and demonstrate production-level capability.
- Infrastructure As Code
Apply infrastructure as code to solve real industry problems and demonstrate production-level capability.
- Vector Databases
Apply vector databases to solve real industry problems and demonstrate production-level capability.
- Chaos Engineering
Apply chaos engineering to solve real industry problems and demonstrate production-level capability.
- Cloud Services
Apply cloud services to solve real industry problems and demonstrate production-level capability.
Careers
Career paths this challenge builds toward
Completing this challenge demonstrates skills that transfer directly to these roles:
AI Solutions Architect
Multi-region failover design for stateful AI workloads is the kind of project that defines an AI solutions architect's first year at an enterprise-AI company.
This challenge sharpens
- multi-region-architecture
- disaster-recovery
- vector-databases
MLOps Engineer
Chaos engineering and SRE-ready runbooks are MLOps disciplines for keeping AI services up under real-world stress.
This challenge sharpens
- chaos-engineering
- infrastructure-as-code
- cloud-services
AI Engineer
Stateful RAG architecture across regions is the AI-engineer-meets-platform skillset that early-stage AI companies hire for.
This challenge sharpens
- vector-databases
- cloud-services
- disaster-recovery