ixio
coding-agent leaderboard · teams

Which team makes the best coding model?

67teams/478models/2026-08-16updated

Every team is only as good as its single best model. So each one here is ranked by that model's composite — a percentile-blended score across every public benchmark we scrape, fair across scales — not by how many models it ships. Flip to Open to rank by best open-weight model.

Team leaderboard

67 teams
#TeamBest model
1
Anthropic
Claude Opus 544
97
2
OpenAI
GPT-5.6 Sol54
93
3
xAI
Grok 4.618
89
4
Alibaba
open
Qwen3.8-Max46
88
5
Moonshot AI
open
Kimi K320
88
6
Meta
Muse Spark23
84
7
NVIDIA
open
nemotron-3-ultra-550b-a55b14
80
8
DeepSeek
open
DeepSeek-V433
75
9
Google
Gemini-3.7-Flash37
75
10
StepFun
open
step-3.5-flash5
74
11
Z.ai
open
GLM-5.317
73
12
Google DeepMind
AI Co-Mathematician2
72
13
MiroMind
MiroThinker-H12
71
14
Mistral AI
open
mistral-small-4-119b-260319
70
15
Baidu
ERNIE-5.12
69
16
ByteDance
Seed2.0 Pro3
69
17
MiniMax
open
MiniMax M38
67
18
Meituan
open
LongCat-Flash-Chat2
65
19
Bytedance
seed-2.1-pro-preview1
65
20
Upstage
Solar Pro 43
65
21
Tencent
Hunyuan-Hy311
63
22
Academic Research
Agents-A18
62
23
Xiaomi
open
MiMo-V2.5-Pro5
62
24
Exa
Exa Agent3
61
25
Amazon
Amazon-Nova-Chat-11-105
59
26
Microsoft
MAI-1-Preview8
57
27
Thinking Machines
Inkling-Small2
55
28
Parallel
Parallel Ultra8x6
54
29
Prime Intellect
open
INTELLECT-31
53
30
Motif Technologies
Motif 31
53
31
MIT
GLM-5.2 (Max) Z.ai ·2
52
32
Thinky
Thinking Machines Inkling1
50
33
Cohere
Command A (03-2025)6
49
34
Ai2
open
OLMo-3-32b-think5
49
35
01 AI
Yi-Lightning5
48
36
NexusFlow
Athene-v2-Chat-72B2
48
37
Sarvam AI
Sarvam-105B2
47
38
Perplexity
Perplexity Agent Advanced5
46
39
PolarSeeker
OpenSeeker-v22
44
40
Alibaba Cloud / Tongyi Lab
Tongyi DeepResearch5
41
41
LM-Provers
QED-Nano1
41
42
Zhipu AI
GLM-4.7-Flash1
40
43
AI21 Labs
open
Jamba-1.5-Large2
40
44
SK Telecom
A.X-K21
40
45
Tavily
Tavily + GPT-5.4 harness1
40
46
Reka AI
Reka-Core-202409044
39
47
Poolside
open
laguna-m.12
39
48
LG AI Research
K-EXAONE 2.01
38
49
Princeton
open
Gemma-2-9B-it-SimPO1
37
50
IBM
open
Granite-3.1-8B-Instruct5
35
51
InternLM
InternLM2.5-20B-chat1
34
52
HKUST NLP Group
WebExplorer-8B (RL)1
33
53
Nexusflow
open
Starling-LM-7B-beta1
32
54
THUDM / Tsinghua University
DeepDive-32B1
32
55
Proprietary
KAT-Coder-Pro-V11
32
56
HuggingFace
open
Zephyr-ORPO-141b-A35b-v0.12
32
57
Databricks
DBRX-Instruct-Preview1
31
58
OpenChat
open
OpenChat-3.5-01062
30
59
UC Berkeley
Starling-LM-7B-alpha1
29
60
NousResearch
open
Nous-Hermes-2-Mixtral-8x7B-DPO2
29
61
Snowflake
open
Snowflake Arctic Instruct1
28
62
LMSYS
Vicuna-33B1
27
63
Inception AI
mercury-21
27
64
Upstage AI
SOLAR-10.7B-Instruct-v1.01
27
65
Perplexity AI
pplx-70B-online1
26
66
MosaicML
MPT-30B-chat1
26
67
Cognitive
open
Dolphin-2.2.1-Mistral-7B1
26

A team's score is the composite of its top-ranked model — shipping more models never helps unless one is genuinely better. Best model → opens that model's full profile.