in the loop
Scores refreshed Oct 9Epoch AI

Every model, ranked.
Scores, prices and context.

337 models from 46 labs. 274 are independently scored; the rest show up as soon as they are.

Leaders

Over time

The frontier keeps moving.

  • Open weights 104
  • Closed 119
  • Best so far
60708090100110120130140150160170JAN 24JUL 24JAN 25JUL 25JAN 26JUL 26Kimi K3Claude Opus 5.5
Figure 1. Every scored model by release date and Capabilities Index. The line is the best score so far.

Every model

All 337.

AI models with their Capabilities Index, benchmark scores and input price
#
1Claude Opus 5.5Anthropic167.387.5$4.00
2GPT-6 AstraOpenAI166.494.4$10
3GPT-6.1 SolOpenAI166.193.9$2.00
4Claude Sonnet 5.5Anthropic165.094.1$2.00
5Claude Fable 5.1Anthropic164.7–$10
6Claude Opus 5Anthropic162.891.8$5.00
7GPT-6 SolOpenAI162.792.4$2.00
8GPT-5.5 ProOpenAI162.191.9$30
9Claude Fable 5Anthropic162.181.1$10
10GPT-5.6 SolOpenAI161.791.3$2.00
11GPT-5.6 TerraOpenAI159.691.1$2.00
12GPT-5.5OpenAI159.192.0$5.00
13GPT-5.4 ProOpenAI158.992.8$30
14Claude Opus 4.8Anthropic158.288.0$5.00
15Kimi K3Moonshot · open157.490.8$0.80
16Gemini 3.7 FlashGoogle157.393.1$0.75
17GPT-5.4OpenAI156.891.1$2.50
18GPT-5.3 CodexOpenAI156.8–$1.75
19Muse Spark 1.3Meta156.8–$1.25
20Gemini 3.8 FlashGoogle156.793.9$0.75
21Grok 4.6xAI156.492.0$2.00
22Qwen 3.8 MaxAlibaba156.490.2–
23GPT-5.6 LunaOpenAI156.488.8$0.20
24GPT-6 LunaOpenAI156.387.3$0.10
25Claude Opus 4.7Anthropic156.386.9$5.00
26Claude Sonnet 5Anthropic156.287.4$2.00
27GLM-5.3Z.ai155.687.9$0.039
28GPT-5.2 ProOpenAI155.4–$21
29DeepSeek V4 Pro 0813DeepSeek · open155.388.9$0.66
30Claude Opus 4.6Anthropic155.287.4$5.00
31Qwen3.8 Max (0902)Alibaba155.189.7$2.00
32DeepSeek V4.1 FlashDeepSeek · open154.9–$0.30
33Muse Spark 1.2Meta154.9–$1.25
34Gemini 3.1 ProGoogle154.892.6–
35DeepSeek V4 Flash 0731DeepSeek · open154.588.0$0.006
36Gemini 3.5 FlashGoogle154.590.4$1.50
37Gemini 3.6 FlashGoogle154.392.2$0.75
38Muse Spark 1.1Meta154.2–$1.25
39Grok 4.5xAI153.991.3$2.00
40Qwen3.7-MaxAlibaba153.787.9$1.48
41Grok 4.7xAI153.590.3$2.00
42GPT-5.2OpenAI153.488.5$1.75
43Gemini 3 ProGoogle152.990.1–
44Claude Sonnet 4.6Anthropic152.283.2$3.00
45Muse SparkMeta152.086.4–
46Grok 4.20xAI152.085.8$1.25
47GLM-5.3-FlashZ.ai · open151.986.9$0.15
48Gemini 3 FlashGoogle151.885.9–
49GLM-5.2Z.ai · open151.889.1$0.06
50Kimi K2.6Moonshot · open151.187.7$0.47

337 models · scores in % · not yet scored at the bottom

Where the numbers come from

Scores are Epoch AI's own runs and the public leaderboards they collect, not launch claims. The Capabilities Index puts them on one scale. Prices and context windows are OpenRouter's listing.

GPQA Diamond
Graduate-level science questions written so that searching the web doesn't help.
SWE-bench Verified
Fixing real GitHub issues in Python projects; the human-checked subset.
FrontierMath
Unpublished research-level maths problems (tiers 1–3).
ARC-AGI-2
Abstract visual puzzles that are easy for people and hard for models.
Humanity's Last Exam
Expert-written questions across dozens of fields.
SimpleQA Verified
Short factual questions: does the model know, or make it up.
Mock AIME
Competition maths in the style of the AIME.
Terminal-Bench
Getting real tasks done in a command-line terminal.