model profile
Rank · overall#2
Score95.1
Benchmark scores
11 benchmarksSWE-bench Verified95%% Resolved · 2026-06-09
Terminal-Bench 2.183.8%Accuracy · Jun 7, 2026 · ± 1.2
ixio runs100%Pass rate · 2026-07-09
Chatbot Arena (Coding)1566Coding Elo
Artificial Analysis62.1Intelligence Index
ARC-AGI89.2%Score
Arena Agent12.01%Net Improvement · ± 2.57
WebDev Arena1627WebDev Elo · ± 8
FrontierMath (Tier 4)87.8%Accuracy · ± 5.2
MCP Atlas83.3Score
BrowseComp88%Accuracy
Across harnesses
3 agentsClaude CodeCLI
—83.8%100%————————
100mini-SWE-agentAgent
——100%————————
100Terminus 2Agent
—80.4%—————————
88Per-benchmark score · composite per harness. Tap a row to open that board.
Scraped and aggregated from public leaderboards · 2026-08-16