coding-agent leaderboard · teams
Which team makes the best coding model?
67teams/478models/2026-08-16updated
Every team is only as good as its single best model. So each one here is ranked by that model's composite — a percentile-blended score across every public benchmark we scrape, fair across scales — not by how many models it ships. Flip to Open to rank by best open-weight model.
Team leaderboard
67 teams
#TeamBest model
1Claude Opus 544972GPT-5.6 Sol54933Grok 4.618894Qwen3.8-Max46885Kimi K320886Muse Spark23847nemotron-3-ultra-550b-a55b14808DeepSeek-V433759Gemini-3.7-Flash377510step-3.5-flash57411GLM-5.3177312AI Co-Mathematician27213MiroThinker-H127114mistral-small-4-119b-2603197015ERNIE-5.126916Seed2.0 Pro36917MiniMax M386718LongCat-Flash-Chat26519seed-2.1-pro-preview16520Solar Pro 436521Hunyuan-Hy3116322Agents-A186223MiMo-V2.5-Pro56224Exa Agent36125Amazon-Nova-Chat-11-1055926MAI-1-Preview85727Inkling-Small25528Parallel Ultra8x65429INTELLECT-315330Motif 315331GLM-5.2 (Max) Z.ai ·25232Thinking Machines Inkling15033Command A (03-2025)64934OLMo-3-32b-think54935Yi-Lightning54836Athene-v2-Chat-72B24837Sarvam-105B24738Perplexity Agent Advanced54639OpenSeeker-v224440Tongyi DeepResearch54141QED-Nano14142GLM-4.7-Flash14043Jamba-1.5-Large24044A.X-K214045Tavily + GPT-5.4 harness14046Reka-Core-2024090443947laguna-m.123948K-EXAONE 2.013849Gemma-2-9B-it-SimPO13750Granite-3.1-8B-Instruct53551InternLM2.5-20B-chat13452WebExplorer-8B (RL)13353Starling-LM-7B-beta13254DeepDive-32B13255KAT-Coder-Pro-V113256Zephyr-ORPO-141b-A35b-v0.123257DBRX-Instruct-Preview13158OpenChat-3.5-010623059Starling-LM-7B-alpha12960Nous-Hermes-2-Mixtral-8x7B-DPO22961Snowflake Arctic Instruct12862Vicuna-33B12763mercury-212764SOLAR-10.7B-Instruct-v1.012765pplx-70B-online12666MPT-30B-chat12667Dolphin-2.2.1-Mistral-7B126
Anthropic
OpenAI
xAI
Alibaba
open
Moonshot AI
open
Meta
NVIDIA
open
DeepSeek
open
Google
StepFun
open
Z.ai
open
Google DeepMind
MiroMind
Mistral AI
open
Baidu
ByteDance
MiniMax
open
Meituan
open
Bytedance
Upstage
Tencent
Academic Research
Xiaomi
open
Exa
Amazon
Microsoft
Thinking Machines
Parallel
Prime Intellect
open
Motif Technologies
MIT
Thinky
Cohere
Ai2
open
01 AI
NexusFlow
Sarvam AI
Perplexity
PolarSeeker
Alibaba Cloud / Tongyi Lab
LM-Provers
Zhipu AI
AI21 Labs
open
SK Telecom
Tavily
Reka AI
Poolside
open
LG AI Research
Princeton
open
IBM
open
InternLM
HKUST NLP Group
Nexusflow
open
THUDM / Tsinghua University
Proprietary
HuggingFace
open
Databricks
OpenChat
open
UC Berkeley
NousResearch
open
Snowflake
open
LMSYS
Inception AI
Upstage AI
Perplexity AI
MosaicML
Cognitive
open
A team's score is the composite of its top-ranked model — shipping more models never helps unless one is genuinely better. Best model → opens that model's full profile.