model profile
Rank · overall#29
Score73.0
Benchmark scores
7 benchmarksSWE-bench Verified79.8%% Resolved · 2026-06-30
Terminal-Bench 2.174.6%Accuracy · Jul 9, 2026 · ± 1.6
Chatbot Arena (Coding)1520Coding Elo
Arena Agent7.14%Net Improvement · ± 2.29
WebDev Arena1540WebDev Elo · ± 8
FrontierMath (Tier 4)29.3%Accuracy · ± 7.2
BrowseComp86.6%Accuracy
Across harnesses
1 agentsPer-benchmark score · composite per harness. Tap a row to open that board.
Scraped and aggregated from public leaderboards · 2026-08-16