model profile
Rank · overall#15
Score76.2
Benchmark scores
12 benchmarksSWE-bench Verified82%% Resolved · 2026-04-16
Terminal-Bench 2.168.9%Accuracy · May 1, 2026 · ± 1.4
Chatbot Arena (Coding)1559Coding Elo
ARC-AGI75.8%Score
Arena Agent8.17%Net Improvement · ± 1.43
WebDev Arena1558WebDev Elo · ± 6
Stratix Cup22.7Tournament score
FrontierMath (Tier 4)31.7%Accuracy · ± 7.4
MCP Atlas79.1Score
SkillsBench61.2Score
MathArena Apex40.6%Accuracy
BrowseComp79.3%Accuracy
Across harnesses
2 agentsPer-benchmark score · composite per harness. Tap a row to open that board.
Scraped and aggregated from public leaderboards · 2026-08-16