Evaluate an Agent Suite on the SWE-Bench-Style Coding Benchmark
Overview
What this challenge is about.
Test 3 open-source agent frameworks on 50 coding tasks, compare pass@1, cost, and speed, then recommend one. Get a verifiable certificate.
The scenario
The applied-AI org (anonymized, large US tech company, around 60 engineers in the unit, internal-coding-agent program) wants to standardize tooling so contributors aren't all building bespoke agent harnesses.
The Brief
What you'll do, and what you'll demonstrate.
Pick the open-source coding-agent framework that gives the org the best pass@1 per dollar on the benchmark, with an ADR the leadership can adopt.
Earning criteria — what you'll demonstrate
- Benchmark agent frameworks on a real coding task suite
- Reason about agent cost and latency, not just accuracy
- Author an ADR that survives technical leadership review
- Diagnose where agent failures actually come from (planning vs. tool-use vs. model)
Program Fit
Where this fits in your program.
Sharpens the same skills your degree expects you to demonstrate.
Aligned coursework coming soon.
Skills
Skills you'll demonstrate.
Each one shows up on your verified credential.
- Llm Agents
Apply llm agents to solve real industry problems and demonstrate production-level capability.
- Agent Evaluation
Apply agent evaluation to solve real industry problems and demonstrate production-level capability.
- Benchmarking
Apply benchmarking to solve real industry problems and demonstrate production-level capability.
- Tool Use
Apply tool use to solve real industry problems and demonstrate production-level capability.
- Python
Write clean, efficient Python for data processing, automation, and backend services.
- Cost Analysis
Apply cost analysis to solve real industry problems and demonstrate production-level capability.
Careers
Career paths this challenge builds toward
Completing this challenge demonstrates skills that transfer directly to these roles:
AI Engineer
Cross-framework agent benchmarking with an ADR-quality writeup is the kind of project that signals senior-AI-engineer judgment in interviews.
This challenge sharpens
- llm-agents
- agent-evaluation
- benchmarking