Overview
What this challenge is about.
Build a monitoring stack for a RAG assistant with drift, cost, and latency alerts, then write a team playbook. Get a verifiable certificate.
The scenario
The bank's AI team (around 90 engineers) had a public refusal-rate incident on its customer chatbot last quarter and now has board-level interest in monitoring maturity.
The Brief
What you'll do, and what you'll demonstrate.
Stand up an end-to-end monitoring stack for one LLM-backed product and write the playbook to onboard five more.
Earning criteria — what you'll demonstrate
- Design monitoring for both classical ML drift and LLM-specific quality
- Implement an end-to-end collection-to-dashboard pipeline
- Set up alerts that page the right people without crying wolf
- Write a playbook that scales monitoring across a product portfolio
Program Fit
Where this fits in your program.
Sharpens the same skills your degree expects you to demonstrate.
ML Engineering and Production ML
Master · Ai Systems
Strong alignment
This challenge maps to ML Engineering and Production ML at the Master level. It sharpens the same practical skills your coursework expects — but in a real industry context with actual constraints and deliverables.
Skills
Skills you'll demonstrate.
Each one shows up on your verified credential.
- Model Monitoring
Apply model monitoring to solve real industry problems and demonstrate production-level capability.
- Data Drift Detection
Apply data drift detection to solve real industry problems and demonstrate production-level capability.
- Llm Evaluation
Apply llm evaluation to solve real industry problems and demonstrate production-level capability.
- Grafana
Apply grafana to solve real industry problems and demonstrate production-level capability.
- Alerting
Apply alerting to solve real industry problems and demonstrate production-level capability.
- Observability
Apply observability to solve real industry problems and demonstrate production-level capability.
Careers
Career paths this challenge builds toward
Completing this challenge demonstrates skills that transfer directly to these roles:
MLOps Engineer
Owning monitoring for LLM-backed products from drift to cost is the MLOps platform role that enterprise AI teams urgently need post-incident.
This challenge sharpens
- model-monitoring
- data-drift-detection
- observability
AI Engineer
Wiring LLM-as-judge sampling and refusal-rate metrics is core AI-engineer work at any team shipping production LLM features.
This challenge sharpens
- llm-evaluation
- model-monitoring
- alerting
AI Solutions Architect
Designing the monitoring stack and writing the cross-portfolio playbook is the architectural work AI solutions architects own at enterprise customers.
This challenge sharpens
- model-monitoring
- observability
- grafana