Scores refreshed Oct 9Epoch AI Every model, ranked.
Every model, ranked.
Scores, prices and context.
337 models from 46 labs. 274 are independently scored; the rest show up as soon as they are.
Leaders
Over time
The frontier keeps moving.
- Open weights 104
- Closed 119
- Best so far
Every model
All 337.
337 models · scores in % · not yet scored at the bottom
Where the numbers come from
Scores are Epoch AI's own runs and the public leaderboards they collect, not launch claims. The Capabilities Index puts them on one scale. Prices and context windows are OpenRouter's listing.
- GPQA Diamond
- Graduate-level science questions written so that searching the web doesn't help.
- SWE-bench Verified
- Fixing real GitHub issues in Python projects; the human-checked subset.
- FrontierMath
- Unpublished research-level maths problems (tiers 1–3).
- ARC-AGI-2
- Abstract visual puzzles that are easy for people and hard for models.
- Humanity's Last Exam
- Expert-written questions across dozens of fields.
- SimpleQA Verified
- Short factual questions: does the model know, or make it up.
- Mock AIME
- Competition maths in the style of the AIME.
- Terminal-Bench
- Getting real tasks done in a command-line terminal.