ixio
← Top Benchmark
benchmark · model-level

FrontierMath (Tier 4)

Accuracy·40 models·homepage ↗
Rank · benchmarks#14 / 18
Score29

How it works

Epoch AI's FrontierMath, Tier 4 (v2) — exceptionally difficult, research-level math problems, evaluated on a private held-out set so answers can't leak into training data. Verified numerically.

Metric
Accuracy
Models
40
Type
Model-level

How we ranked it

composite 29 / 100

Every benchmark is scored on four weighted criteria — how directly it measures a real coding agent, how much of the stack it covers, how real its tasks are, and how open and reproducible it is. Those blend into the composite that ranks it #14 of 18.

Agent-nativeweight 30%
0

Scores a real (harness × model) pair — the full agent — not just the bare model.

Stack coverageweight 30%
0

Data-driven: this benchmark exposes 40 models.

Task realismweight 25%
60

Executable, real-world coding tasks with hard pass/fail — not human preference or aggregate scores.

Open & reproducibleweight 15%
92

Open data with per-run receipts you can audit.

Our review

Best-in-class for contamination resistance and openness of data and method. But it measures deep mathematical reasoning, not the coding-agent stack — a model can top FrontierMath and still be a mediocre coding agent — so it sits mid-table for our purposes.

Top on FrontierMath (Tier 4)

40 models · Accuracy

Source: epoch.ai/benchmarks/frontiermath-tier-4-v2 · aggregated 2026-07-27