Design Eval Suite for a Multimodal Brainstorming Assistant
Overview
What this challenge is about.
Design an eval suite for a multimodal AI assistant, run it on two models, and earn a verifiable certificate.
The scenario
The startup (around 50 people, planning a public launch) needs an eval suite that survives model swaps, founder pressure, and external safety review; bespoke notebooks aren't going to cut it.
The Brief
What you'll do, and what you'll demonstrate.
Design and prototype a CI-runnable evaluation suite for a multimodal brainstorming assistant covering quality, factuality, safety, and creativity.
Earning criteria — what you'll demonstrate
- Design a multimodal evaluation suite balancing automated + rubric scores
- Build safety + factuality probe sets that survive future model changes
- Engineer eval as code, runnable in CI
- Communicate eval-suite trade-offs to product + safety leadership
Program Fit
Where this fits in your program.
Sharpens the same skills your degree expects you to demonstrate.
Generative AI
Master · Generative Ai
Strong alignment
This challenge maps to Generative AI at the Master level. It sharpens the same practical skills your coursework expects — but in a real industry context with actual constraints and deliverables.
Skills
Skills you'll demonstrate.
Each one shows up on your verified credential.
- Llm Evaluation
Apply llm evaluation to solve real industry problems and demonstrate production-level capability.
- Multimodal Evaluation
Apply multimodal evaluation to solve real industry problems and demonstrate production-level capability.
- Safety Evaluation
Apply safety evaluation to solve real industry problems and demonstrate production-level capability.
- Rubric Design
Apply rubric design to solve real industry problems and demonstrate production-level capability.
- Python
Write clean, efficient Python for data processing, automation, and backend services.
- Ci Integration
Apply ci integration to solve real industry problems and demonstrate production-level capability.
Careers
Career paths this challenge builds toward
Completing this challenge demonstrates skills that transfer directly to these roles:
AI Product Manager
Designing the eval suite that gates a consumer launch is exactly the day-one work of an AI PM at any consumer-AI company.
This challenge sharpens
- rubric-design
- llm-evaluation
- safety-evaluation
AI Safety Researcher
Building versioned safety + factuality probe sets that survive model swaps is core AI safety work in product-led organizations.
This challenge sharpens
- safety-evaluation
- rubric-design
- multimodal-evaluation
MLOps Engineer
Shipping evaluation as a CI-runnable harness is the MLOps craft of making model-quality gates automatic and reliable.
This challenge sharpens
- ci-integration
- python
- llm-evaluation