Spec Trust-and-Safety Eval Harness for an LLM-Powered Customer-Support Bot
Overview
What this challenge is about.
Spec Trust-and-Safety Eval Harness for an LLM-Powered Customer-Support Bot. Advanced challenge in design. Designing real products under real constraints, ear...
The Brief
What you'll do, and what you'll demonstrate.
Spec and reference-implement a nightly trust-and-safety harness for an LLM customer-support bot covering jailbreaks, PII, and toxicity.
This is not a design exercise. It is the work a product designer does between a brief and a shipped interface. That distinction matters to every hiring manager who has seen candidates redesign Spotify's homepage and none who have worked under real product constraints.
When you finish, you will have something most graduates do not: a real-world deliverable, verified by Ewance, that you can show to a hiring manager and say "I did this. Here is the proof."
Earning criteria — what you'll demonstrate
- Design a multi-axis safety evaluation harness for LLM products
- Curate jailbreak and PII test sets at useful scale
- Integrate a toxicity classifier into automated gating
- Document a harness so engineering can pick it up next sprint
Program Fit
Where this fits in your program.
Sharpens the same skills your degree expects you to demonstrate.
Aligned coursework coming soon.
Skills
Skills you'll demonstrate.
Each one shows up on your verified credential.
- Llm Evaluation
Apply llm evaluation to solve real industry problems and demonstrate production-level capability.
- Red Teaming
Apply red teaming to solve real industry problems and demonstrate production-level capability.
- Pii Detection
Apply pii detection to solve real industry problems and demonstrate production-level capability.
- Regression Detection
Apply regression detection to solve real industry problems and demonstrate production-level capability.
- Python
Write clean, efficient Python for data processing, automation, and backend services.
- Harness Design
Apply harness design to solve real industry problems and demonstrate production-level capability.
Careers
Career paths this challenge builds toward
Completing this challenge demonstrates skills that transfer directly to these roles:
AI Safety Researcher
Designing a nightly safety eval harness for an LLM product is the AI safety researcher's textbook job at any enterprise-AI vendor.
This challenge sharpens
- llm-evaluation
- red-teaming
- pii-detection
MLOps Engineer
Gating deploys on regression-detection thresholds is the MLOps engineer's craft applied to safety axes.
This challenge sharpens
- regression-detection
- harness-design
- python
Prompt Engineer
Curating jailbreak test sets and reasoning about LLM failure axes is core prompt-engineer territory in safety-conscious product teams.
This challenge sharpens
- llm-evaluation
- red-teaming
- harness-design