model profile
Rank · overall#101
Score60.1
Benchmark scores
4 benchmarksSWE-bench Verified79.4%% Resolved · 2026-05-19
Chatbot Arena (Coding)1505Coding Elo
Arena Agent1.46%Net Improvement · ± 0.93
FrontierMath (Tier 4)34.1%Accuracy · ± 7.5
Across harnesses
0 agentsThis model appears only in model-level benchmarks — no harness ran it directly.
Scraped and aggregated from public leaderboards · 2026-09-05