Skip to contentSkip to content
Verified credentials. On-chain. Forever.Learn more
Ewance
Sign in
Cover image for Run an Alignment Probe on a Coding Assistant
Research

Run an Alignment Probe on a Coding Assistant

FreeVerified credential3 weeksAdvanced

Overview

What this challenge is about.

Design 240 probe prompts to test an AI coder, score outputs, and write a red-team report. End with a verifiable certificate.

The scenario

The lab (~400 staff) publishes model cards as part of its release process; an external red-team report adds credibility ahead of enterprise adoption.

CredentialBlockchain-anchored
ShareableLinkedIn-ready
LanguageEnglish
PaceSelf-paced

The Brief

What you'll do, and what you'll demonstrate.

Probe a 14B coding assistant for over-refusal, insecure code generation, and data leakage, and publish a model-card-ready red-team report.

Earning criteria — what you'll demonstrate

  • Design probe prompts for distinct alignment failure modes
  • Apply consistent rubrics for hand-scoring LLM outputs
  • Build a severity ranking that combines likelihood and impact
  • Write a public-facing red-team report

Program Fit

Where this fits in your program.

Sharpens the same skills your degree expects you to demonstrate.

Large Language Models

Master · Generative Ai

Strong alignment

This challenge maps to Large Language Models at the Master level. It sharpens the same practical skills your coursework expects — but in a real industry context with actual constraints and deliverables.

Careers

Career paths this challenge builds toward

Completing this challenge demonstrates skills that transfer directly to these roles:

AI Safety Researcher

Designing and running an alignment red-team is the core day-to-day of safety researchers at any frontier AI lab.

This challenge sharpens

  • red-teaming
  • alignment-evaluation
  • risk-assessment

ML Researcher

Disciplined probe design and inter-rater scoring is the methodological foundation of empirical LLM research.

This challenge sharpens

  • prompt-design
  • llm-evaluation
  • alignment-evaluation

Research Scientist

Writing a model-card-ready red-team report is exactly the publishable output expected of a junior research scientist at an AI lab.

This challenge sharpens

  • report-writing
  • red-teaming
  • alignment-evaluation

One more thing

You can put a credential on your CV by Friday.