model profile
Rank · overall#109
Score58.7
Benchmark scores
4 benchmarksSWE-bench Verified79.4%% Resolved · 2026-05-19
Chatbot Arena (Coding)1505Coding Elo
Arena Agent0.2%Net Improvement · ± 0.87
FrontierMath (Tier 4)34.1%Accuracy · ± 7.5
Across harnesses
0 agentsThis model appears only in model-level benchmarks — no harness ran it directly.
Scraped and aggregated from public leaderboards · 2026-08-16