Run a Human-Preference Study Comparing Two Coding Assistants
Overview
What this challenge is about.
Run a blinded paired-comparison study with 12 developers and 8 coding tasks, analyze results, and earn a verifiable certificate.
The scenario
The startup (around 20 staff, around 12,000 active IDE users) is paying around USD 18,000 a month across two vendors and needs to consolidate to one without hurting product quality.
The Brief
What you'll do, and what you'll demonstrate.
Run a pre-registered human-preference study comparing two coding assistants and produce a vendor-decision recommendation.
Earning criteria — what you'll demonstrate
- Design a pre-registered human-preference study
- Justify sample size before collecting data
- Analyze paired-comparison data with frequentist and Bayesian methods
- Present a vendor decision under statistical uncertainty
Program Fit
Where this fits in your program.
Sharpens the same skills your degree expects you to demonstrate.
AI Measurement and Evaluation
Master · Responsible Ai
Strong alignment
This challenge maps to AI Measurement and Evaluation at the Master level. It sharpens the same practical skills your coursework expects — but in a real industry context with actual constraints and deliverables.
Skills
Skills you'll demonstrate.
Each one shows up on your verified credential.
- Experiment Design
Apply experiment design to solve real industry problems and demonstrate production-level capability.
- Statistical Evaluation
Apply statistical evaluation to solve real industry problems and demonstrate production-level capability.
- Human Evaluation
Apply human evaluation to solve real industry problems and demonstrate production-level capability.
- Pre Registration
Apply pre registration to solve real industry problems and demonstrate production-level capability.
- Llm Evaluation
Apply llm evaluation to solve real industry problems and demonstrate production-level capability.
- Stakeholder Communication
Apply stakeholder communication to solve real industry problems and demonstrate production-level capability.
Careers
Career paths this challenge builds toward
Completing this challenge demonstrates skills that transfer directly to these roles:
Applied AI Scientist
Designing a pre-registered evaluation for a real vendor decision is the applied AI scientist's contribution to product orgs.
This challenge sharpens
- experiment-design
- statistical-evaluation
- llm-evaluation
Data Scientist
Paired-comparison analysis with honest uncertainty is bread-and-butter data-scientist craft.
This challenge sharpens
- statistical-evaluation
- experiment-design
- pre-registration
AI Product Manager
Turning an evaluation into a defensible vendor decision is exactly the AI PM's contribution to the procurement conversation.
This challenge sharpens
- stakeholder-communication
- human-evaluation
- llm-evaluation