Skip to contentSkip to content
Verified credentials. On-chain. Forever.Learn more
Ewance
Sign in
Cover image for Evaluate an Agent Suite on the SWE-Bench-Style Coding Benchmark
Analysis

Evaluate an Agent Suite on the SWE-Bench-Style Coding Benchmark

FreeVerified credential2 weeksAdvanced

Overview

What this challenge is about.

Test 3 open-source agent frameworks on 50 coding tasks, compare pass@1, cost, and speed, then recommend one. Get a verifiable certificate.

The scenario

The applied-AI org (anonymized, large US tech company, around 60 engineers in the unit, internal-coding-agent program) wants to standardize tooling so contributors aren't all building bespoke agent harnesses.

CredentialBlockchain-anchored
ShareableLinkedIn-ready
LanguageEnglish
PaceSelf-paced

The Brief

What you'll do, and what you'll demonstrate.

Pick the open-source coding-agent framework that gives the org the best pass@1 per dollar on the benchmark, with an ADR the leadership can adopt.

Earning criteria — what you'll demonstrate

  • Benchmark agent frameworks on a real coding task suite
  • Reason about agent cost and latency, not just accuracy
  • Author an ADR that survives technical leadership review
  • Diagnose where agent failures actually come from (planning vs. tool-use vs. model)

Program Fit

Where this fits in your program.

Sharpens the same skills your degree expects you to demonstrate.

Aligned coursework coming soon.

One more thing

You can put a credential on your CV by Friday.