Design a Capability Evaluation for an Open-Weights Coding Model
Overview
What this challenge is about.
Design 40 coding tasks across 4 safety buckets for an open-weights model, report pass/refusal rates, and earn a verifiable certificate.
The scenario
The nonprofit (around 30 staff, foundation-funded, EU + UK reach) publishes evaluations that get cited in parliamentary hearings; the next briefing focuses on open-weights coding models.
The Brief
What you'll do, and what you'll demonstrate.
Run a documented capability evaluation of an open-weights coding model across benign, dual-use, refusal-bait, and hard buckets.
Earning criteria — what you'll demonstrate
- Design a multi-bucket capability evaluation with refusal tracking
- Use only public, ethics-safe task sources for dual-use buckets
- Report capability results with statistical honesty
- Translate evaluation results into policy-relevant observations
Program Fit
Where this fits in your program.
Sharpens the same skills your degree expects you to demonstrate.
AI Safety and Alignment
Master · Responsible Ai
Strong alignment
This challenge maps to AI Safety and Alignment at the Master level. It sharpens the same practical skills your coursework expects — but in a real industry context with actual constraints and deliverables.
Skills
Skills you'll demonstrate.
Each one shows up on your verified credential.
- Capability Evaluation
Apply capability evaluation to solve real industry problems and demonstrate production-level capability.
- Safety Evaluation
Apply safety evaluation to solve real industry problems and demonstrate production-level capability.
- Llm Evaluation
Apply llm evaluation to solve real industry problems and demonstrate production-level capability.
- Research Writing
Apply research writing to solve real industry problems and demonstrate production-level capability.
- Policy Communication
Apply policy communication to solve real industry problems and demonstrate production-level capability.
- Python
Write clean, efficient Python for data processing, automation, and backend services.
Careers
Career paths this challenge builds toward
Completing this challenge demonstrates skills that transfer directly to these roles:
AI Safety Researcher
Public capability evaluations with policy framing are exactly the role's contribution to the AI governance conversation.
This challenge sharpens
- capability-evaluation
- safety-evaluation
- policy-communication
ML Researcher
Designing a multi-bucket evaluation set with rigorous statistics is bread-and-butter ML research work.
This challenge sharpens
- llm-evaluation
- capability-evaluation
- research-writing
Applied AI Scientist
Translating evaluations into policy-relevant observations is the applied-AI bridge into policy and governance teams.
This challenge sharpens
- llm-evaluation
- policy-communication
- research-writing