ixio
coding-agent leaderboard · teams

Which team makes the best coding model?

65teams/486models/2026-09-05updated

Every team is only as good as its single best model. So each one here is ranked by that model's composite — a percentile-blended score across every public benchmark we scrape, fair across scales — not by how many models it ships. Flip to Open to rank by best open-weight model.

Team leaderboard

65 teams
#TeamBest model
1
OpenAI
GPT-6 Astra55
97
2
Anthropic
Claude Fable 5.146
96
3
Meta
Muse Spark 1.324
87
4
Moonshot AI
open
Kimi K320
83
5
xAI
Grok 4.618
80
6
Alibaba
open
Qwen3.8-Max51
80
7
NVIDIA
open
nemotron-3-ultra-550b-a55b12
80
8
Google
Gemini 3.1 Pro38
76
9
DeepSeek
open
DeepSeek-V433
75
10
StepFun
open
step-3.5-flash5
74
11
Tencent
Hy4 preview12
71
12
Z.ai
open
GLM-5.319
71
13
MiroMind
MiroThinker-H12
71
14
Mistral AI
open
mistral-small-4-119b-260319
70
15
Google DeepMind
AI Co-Mathematician2
70
16
Baidu
ERNIE-5.12
69
17
ByteDance
Seed2.0 Pro3
69
18
Xiaomi
open
MiMo-V2.5-Pro5
65
19
Upstage
Solar Pro 42
65
20
Meituan
open
LongCat-Flash-Chat2
64
21
MiniMax
open
MiniMax M38
64
22
Bytedance
seed-2.1-pro-preview1
62
23
Academic Research
Agents-A18
62
24
Exa
Exa Agent3
60
25
Amazon
Amazon-Nova-Chat-11-105
59
26
MIT
GLM-5.2 (Max) Z.ai ·2
58
27
Thinky
Thinking Machines Inkling2
57
28
Microsoft
MAI-1-Preview8
56
29
Parallel
Parallel Ultra8x6
54
30
Thinking Machines
Inkling-Small2
53
31
Prime Intellect
open
INTELLECT-31
53
32
Cohere
Command A (03-2025)5
48
33
Ai2
open
OLMo-3-32b-think5
48
34
01 AI
Yi-Lightning5
48
35
NexusFlow
Athene-v2-Chat-72B2
48
36
Sarvam AI
Sarvam-105B2
47
37
Perplexity
Perplexity Agent Advanced5
46
38
PolarSeeker
OpenSeeker-v22
44
39
LM-Provers
QED-Nano1
41
40
Alibaba Cloud / Tongyi Lab
Tongyi DeepResearch5
41
41
Zhipu AI
GLM-4.7-Flash1
40
42
AI21 Labs
open
Jamba-1.5-Large2
40
43
Tavily
Tavily + GPT-5.4 harness1
40
44
Reka AI
Reka-Core-202409044
39
45
MBZUAI Institute of Foundation Models
K2 Horizon 375B A23B1
39
46
Poolside
open
laguna-m.12
37
47
Princeton
open
Gemma-2-9B-it-SimPO1
37
48
IBM
open
Granite-3.1-8B-Instruct5
35
49
InternLM
InternLM2.5-20B-chat1
34
50
HKUST NLP Group
WebExplorer-8B (RL)1
33
51
Nexusflow
open
Starling-LM-7B-beta1
32
52
THUDM / Tsinghua University
DeepDive-32B1
32
53
HuggingFace
open
Zephyr-ORPO-141b-A35b-v0.12
32
54
Proprietary
KAT-Coder-Pro-V11
31
55
Databricks
DBRX-Instruct-Preview1
31
56
OpenChat
open
OpenChat-3.5-01062
30
57
UC Berkeley
Starling-LM-7B-alpha1
29
58
NousResearch
open
Nous-Hermes-2-Mixtral-8x7B-DPO2
29
59
Snowflake
open
Snowflake Arctic Instruct1
28
60
LMSYS
Vicuna-33B1
27
61
Inception AI
mercury-21
27
62
Upstage AI
SOLAR-10.7B-Instruct-v1.01
27
63
Perplexity AI
pplx-70B-online1
26
64
MosaicML
MPT-30B-chat1
26
65
Cognitive
open
Dolphin-2.2.1-Mistral-7B1
26

A team's score is the composite of its top-ranked model — shipping more models never helps unless one is genuinely better. Best model → opens that model's full profile.