Skip to contentSkip to content
Verified credentials. On-chain. Forever.Learn more
Ewance
Sign in
Cover image for Spec Trust-and-Safety Eval Harness for an LLM-Powered Customer-Support Bot
Design

Spec Trust-and-Safety Eval Harness for an LLM-Powered Customer-Support Bot

FreeVerified credential3 weeksAdvanced

Overview

What this challenge is about.

Spec Trust-and-Safety Eval Harness for an LLM-Powered Customer-Support Bot. Advanced challenge in design. Designing real products under real constraints, ear...

CredentialBlockchain-anchored
ShareableLinkedIn-ready
LanguageEnglish
PaceSelf-paced

The Brief

What you'll do, and what you'll demonstrate.

Spec and reference-implement a nightly trust-and-safety harness for an LLM customer-support bot covering jailbreaks, PII, and toxicity.

This is not a design exercise. It is the work a product designer does between a brief and a shipped interface. That distinction matters to every hiring manager who has seen candidates redesign Spotify's homepage and none who have worked under real product constraints.

When you finish, you will have something most graduates do not: a real-world deliverable, verified by Ewance, that you can show to a hiring manager and say "I did this. Here is the proof."

Earning criteria — what you'll demonstrate

  • Design a multi-axis safety evaluation harness for LLM products
  • Curate jailbreak and PII test sets at useful scale
  • Integrate a toxicity classifier into automated gating
  • Document a harness so engineering can pick it up next sprint

Program Fit

Where this fits in your program.

Sharpens the same skills your degree expects you to demonstrate.

Aligned coursework coming soon.

Careers

Career paths this challenge builds toward

Completing this challenge demonstrates skills that transfer directly to these roles:

AI Safety Researcher

Designing a nightly safety eval harness for an LLM product is the AI safety researcher's textbook job at any enterprise-AI vendor.

This challenge sharpens

  • llm-evaluation
  • red-teaming
  • pii-detection

MLOps Engineer

Gating deploys on regression-detection thresholds is the MLOps engineer's craft applied to safety axes.

This challenge sharpens

  • regression-detection
  • harness-design
  • python

Prompt Engineer

Curating jailbreak test sets and reasoning about LLM failure axes is core prompt-engineer territory in safety-conscious product teams.

This challenge sharpens

  • llm-evaluation
  • red-teaming
  • harness-design

One more thing

You can put a credential on your CV by Friday.