Overview
What this challenge is about.
Run 18 prompt configurations on code snippets, test top 3 on 2,000 samples, and pick the cost-efficient winner. Earn a verifiable certificate.
The scenario
The lab (research org, around 200 staff, 7-figure monthly model bill on this single pipeline) holds a quarterly infra review where every workload above USD 50K/month must justify its cost ratio.
The Brief
What you'll do, and what you'll demonstrate.
Find a Pareto-optimal prompt + model configuration that cuts spend by 40 percent on a 2M-call/week scoring pipeline without losing human-agreement quality.
Earning criteria — what you'll demonstrate
- Design a factorial prompt + model experiment under a fixed call budget
- Quantify the cost-quality trade-off rigorously (correlation, CIs, cost-per-call)
- Choose between prompt strategies (zero-shot, few-shot, CoT) on evidence
- Communicate optimization findings to an infrastructure-review audience
Program Fit
Where this fits in your program.
Sharpens the same skills your degree expects you to demonstrate.
Aligned coursework coming soon.
Skills
Skills you'll demonstrate.
Each one shows up on your verified credential.
- Prompt Optimization
Apply prompt optimization to solve real industry problems and demonstrate production-level capability.
- Cost Quality Tradeoff
Apply cost quality tradeoff to solve real industry problems and demonstrate production-level capability.
- Experiment Design
Apply experiment design to solve real industry problems and demonstrate production-level capability.
- Evaluation
Apply evaluation to solve real industry problems and demonstrate production-level capability.
- Structured Output
Apply structured output to solve real industry problems and demonstrate production-level capability.
- Ab Testing
Apply ab testing to solve real industry problems and demonstrate production-level capability.
Careers
Career paths this challenge builds toward
Completing this challenge demonstrates skills that transfer directly to these roles:
Prompt Engineer
Running structured prompt + model sweeps under a real production budget is exactly what senior prompt engineers do at AI-heavy companies.
This challenge sharpens
- prompt-optimization
- cost-quality-tradeoff
- experiment-design
MLOps Engineer
Owning the optimization + monitoring loop on a 2M-call/week pipeline is MLOps-engineer territory at any company spending serious money on LLM APIs.
This challenge sharpens
- cost-quality-tradeoff
- experiment-design
- evaluation
Applied AI Scientist
Factorial experiment design and rigorous cost-quality reporting is the rigor applied AI scientists bring to internal-tooling decisions.
This challenge sharpens
- experiment-design
- evaluation
- ab-testing