model profile
Rank · overall#17
Score75.6
Benchmark scores
7 benchmarksSWE-bench Verified82%% Resolved · 2026-07-09
Terminal-Bench 2.176.2%Accuracy · Jul 9, 2026 · ± 1.2
Chatbot Arena (Coding)1521Coding Elo
Arena Agent1.17%Net Improvement · ± 0.61
WebDev Arena1538WebDev Elo · ± 9
SWE-bench Pro61.5%% Resolved
MCP Atlas88.1Score
Across harnesses
1 agentsPer-benchmark score · composite per harness. Tap a row to open that board.
Scraped and aggregated from public leaderboards · 2026-08-16