Right-Size a Real-Time Recommendation Serving Cluster
Overview
What this challenge is about.
You analyze 7 days of streaming telemetry, tune HPA or KEDA scaling, run a load test, and deliver a rollout plan. Get a verifiable certificate.
The scenario
The startup (around 70 engineers, around 12M monthly active users) spends about USD 90,000 per month on the rec serving tier; the CFO wants a believable 25-30 percent reduction without breaking the experience.
The Brief
What you'll do, and what you'll demonstrate.
Cut off-peak serving cost by 30 percent on a real-time recommendation cluster without breaching the p99 latency SLO.
Earning criteria — what you'll demonstrate
- Analyze serving telemetry to find over-provisioning
- Choose an autoscaling strategy under latency-SLO constraints
- Run a load test that faithfully reproduces peak traffic
- Write a rollout plan with explicit rollback triggers
Program Fit
Where this fits in your program.
Sharpens the same skills your degree expects you to demonstrate.
Machine Learning at Scale
Master · Ai Systems
Strong alignment
This challenge maps to Machine Learning at Scale at the Master level. It sharpens the same practical skills your coursework expects — but in a real industry context with actual constraints and deliverables.
Skills
Skills you'll demonstrate.
Each one shows up on your verified credential.
- Model Serving
Apply model serving to solve real industry problems and demonstrate production-level capability.
- Kubernetes
Apply kubernetes to solve real industry problems and demonstrate production-level capability.
- Autoscaling
Apply autoscaling to solve real industry problems and demonstrate production-level capability.
- Load Testing
Apply load testing to solve real industry problems and demonstrate production-level capability.
- Cost Optimization
Apply cost optimization to solve real industry problems and demonstrate production-level capability.
- Python
Write clean, efficient Python for data processing, automation, and backend services.
Careers
Career paths this challenge builds toward
Completing this challenge demonstrates skills that transfer directly to these roles:
MLOps Engineer
Right-sizing a real-time serving cluster under latency constraints is the daily reality of MLOps engineers at any consumer-ML company.
This challenge sharpens
- model-serving
- autoscaling
- kubernetes
Data Engineer
Telemetry analysis and capacity planning bridge directly into the data-engineer's broader work on pipeline cost discipline.
This challenge sharpens
- python
- cost-optimization
- model-serving
AI Solutions Architect
Designing autoscaling under SLO constraints is the architect's job when sizing real-time AI workloads for enterprise customers.
This challenge sharpens
- model-serving
- autoscaling
- cost-optimization