Skip to contentSkip to content
Verified credentials. On-chain. Forever.Learn more
Ewance
Sign in
Cover image for Design a Capability Evaluation for an Open-Weights Coding Model
Research

Design a Capability Evaluation for an Open-Weights Coding Model

FreeVerified credential3 weeksAdvanced

Overview

What this challenge is about.

Design 40 coding tasks across 4 safety buckets for an open-weights model, report pass/refusal rates, and earn a verifiable certificate.

The scenario

The nonprofit (around 30 staff, foundation-funded, EU + UK reach) publishes evaluations that get cited in parliamentary hearings; the next briefing focuses on open-weights coding models.

CredentialBlockchain-anchored
ShareableLinkedIn-ready
LanguageEnglish
PaceSelf-paced

The Brief

What you'll do, and what you'll demonstrate.

Run a documented capability evaluation of an open-weights coding model across benign, dual-use, refusal-bait, and hard buckets.

Earning criteria — what you'll demonstrate

  • Design a multi-bucket capability evaluation with refusal tracking
  • Use only public, ethics-safe task sources for dual-use buckets
  • Report capability results with statistical honesty
  • Translate evaluation results into policy-relevant observations

Program Fit

Where this fits in your program.

Sharpens the same skills your degree expects you to demonstrate.

AI Safety and Alignment

Master · Responsible Ai

Strong alignment

This challenge maps to AI Safety and Alignment at the Master level. It sharpens the same practical skills your coursework expects — but in a real industry context with actual constraints and deliverables.

Careers

Career paths this challenge builds toward

Completing this challenge demonstrates skills that transfer directly to these roles:

AI Safety Researcher

Public capability evaluations with policy framing are exactly the role's contribution to the AI governance conversation.

This challenge sharpens

  • capability-evaluation
  • safety-evaluation
  • policy-communication

ML Researcher

Designing a multi-bucket evaluation set with rigorous statistics is bread-and-butter ML research work.

This challenge sharpens

  • llm-evaluation
  • capability-evaluation
  • research-writing

Applied AI Scientist

Translating evaluations into policy-relevant observations is the applied-AI bridge into policy and governance teams.

This challenge sharpens

  • llm-evaluation
  • policy-communication
  • research-writing

One more thing

You can put a credential on your CV by Friday.