coding-agent leaderboard · teams
Which team makes the best coding model?
65teams/486models/2026-09-05updated
Every team is only as good as its single best model. So each one here is ranked by that model's composite — a percentile-blended score across every public benchmark we scrape, fair across scales — not by how many models it ships. Flip to Open to rank by best open-weight model.
Team leaderboard
65 teams
#TeamBest model
1GPT-6 Astra55972Claude Fable 5.146963Muse Spark 1.324874Kimi K320835Grok 4.618806Qwen3.8-Max51807nemotron-3-ultra-550b-a55b12808Gemini 3.1 Pro38769DeepSeek-V4337510step-3.5-flash57411Hy4 preview127112GLM-5.3197113MiroThinker-H127114mistral-small-4-119b-2603197015AI Co-Mathematician27016ERNIE-5.126917Seed2.0 Pro36918MiMo-V2.5-Pro56519Solar Pro 426520LongCat-Flash-Chat26421MiniMax M386422seed-2.1-pro-preview16223Agents-A186224Exa Agent36025Amazon-Nova-Chat-11-1055926GLM-5.2 (Max) Z.ai ·25827Thinking Machines Inkling25728MAI-1-Preview85629Parallel Ultra8x65430Inkling-Small25331INTELLECT-315332Command A (03-2025)54833OLMo-3-32b-think54834Yi-Lightning54835Athene-v2-Chat-72B24836Sarvam-105B24737Perplexity Agent Advanced54638OpenSeeker-v224439QED-Nano14140Tongyi DeepResearch54141GLM-4.7-Flash14042Jamba-1.5-Large24043Tavily + GPT-5.4 harness14044Reka-Core-2024090443945K2 Horizon 375B A23B13946laguna-m.123747Gemma-2-9B-it-SimPO13748Granite-3.1-8B-Instruct53549InternLM2.5-20B-chat13450WebExplorer-8B (RL)13351Starling-LM-7B-beta13252DeepDive-32B13253Zephyr-ORPO-141b-A35b-v0.123254KAT-Coder-Pro-V113155DBRX-Instruct-Preview13156OpenChat-3.5-010623057Starling-LM-7B-alpha12958Nous-Hermes-2-Mixtral-8x7B-DPO22959Snowflake Arctic Instruct12860Vicuna-33B12761mercury-212762SOLAR-10.7B-Instruct-v1.012763pplx-70B-online12664MPT-30B-chat12665Dolphin-2.2.1-Mistral-7B126
OpenAI
Anthropic
Meta
Moonshot AI
open
xAI
Alibaba
open
NVIDIA
open
Google
DeepSeek
open
StepFun
open
Tencent
Z.ai
open
MiroMind
Mistral AI
open
Google DeepMind
Baidu
ByteDance
Xiaomi
open
Upstage
Meituan
open
MiniMax
open
Bytedance
Academic Research
Exa
Amazon
MIT
Thinky
Microsoft
Parallel
Thinking Machines
Prime Intellect
open
Cohere
Ai2
open
01 AI
NexusFlow
Sarvam AI
Perplexity
PolarSeeker
LM-Provers
Alibaba Cloud / Tongyi Lab
Zhipu AI
AI21 Labs
open
Tavily
Reka AI
MBZUAI Institute of Foundation Models
Poolside
open
Princeton
open
IBM
open
InternLM
HKUST NLP Group
Nexusflow
open
THUDM / Tsinghua University
HuggingFace
open
Proprietary
Databricks
OpenChat
open
UC Berkeley
NousResearch
open
Snowflake
open
LMSYS
Inception AI
Upstage AI
Perplexity AI
MosaicML
Cognitive
open
A team's score is the composite of its top-ranked model — shipping more models never helps unless one is genuinely better. Best model → opens that model's full profile.