Evaluate an Agent Suite on the SWE-Bench-Style Coding Benchmark
Overview
What this challenge is about.
Evaluate an Agent Suite on the SWE-Bench-Style Coding Benchmark. Advanced challenge in analysis. Analyzing real datasets and building models that drive decis...
The Brief
What you'll do, and what you'll demonstrate.
Pick the open-source coding-agent framework that gives the org the best pass@1 per dollar on the benchmark, with an ADR the leadership can adopt.
This is not a data exercise. It is the work an analyst does when stakeholders need answers from messy data. That distinction matters to every hiring manager who has seen candidates describe statistical methods and none who have extracted insight from messy, real-world data.
When you finish, you will have something most graduates do not: a real-world deliverable, verified by Ewance, that you can show to a hiring manager and say "I did this. Here is the proof."
Earning criteria — what you'll demonstrate
- Benchmark agent frameworks on a real coding task suite
- Reason about agent cost and latency, not just accuracy
- Author an ADR that survives technical leadership review
- Diagnose where agent failures actually come from (planning vs. tool-use vs. model)
Program Fit
Where this fits in your program.
Sharpens the same skills your degree expects you to demonstrate.
Aligned coursework coming soon.
Skills
Skills you'll demonstrate.
Each one shows up on your verified credential.
- Llm Agents
Apply llm agents to solve real industry problems and demonstrate production-level capability.
- Agent Evaluation
Apply agent evaluation to solve real industry problems and demonstrate production-level capability.
- Benchmarking
Apply benchmarking to solve real industry problems and demonstrate production-level capability.
- Tool Use
Apply tool use to solve real industry problems and demonstrate production-level capability.
- Python
Write clean, efficient Python for data processing, automation, and backend services.
- Cost Analysis
Apply cost analysis to solve real industry problems and demonstrate production-level capability.
Careers
Career paths this challenge builds toward
Completing this challenge demonstrates skills that transfer directly to these roles:
AI Engineer
Cross-framework agent benchmarking with an ADR-quality writeup is the kind of project that signals senior-AI-engineer judgment in interviews.
This challenge sharpens
- llm-agents
- agent-evaluation
- benchmarking