model profile
Rank · overall#14
Score76.1
Benchmark scores
9 benchmarksSWE-bench Verified79.8%% Resolved · 2026-06-30
Terminal-Bench 2.174.6%Accuracy · Jul 9, 2026 · ± 1.6
Chatbot Arena (Coding)1520Coding Elo
Artificial Analysis53.4Intelligence Index
AA Coding Index71.5Coding Index
Arena Agent8.66%Net Improvement · ± 1.89
WebDev Arena1543WebDev Elo · ± 10
FrontierMath (Tier 4)29.3%Accuracy · ± 7.2
BrowseComp86.6%Accuracy
Across harnesses
1 agentsPer-benchmark score · composite per harness. Tap a row to open that board.
Scraped and aggregated from public leaderboards · 2026-07-27