in the loop

DeepSeek-R1 (May 2025)
#111 of 274 scored.

DeepSeekMay 28, 2025Open weights

Capabilities Index

Range 138.6–142.8

141.3

Context

Tokens

—

Input

Per 1M tokens

—

Output

Per 1M tokens

—

Benchmarks

How it scores.

  • GPQA Diamond68.4%

    89th of 186

  • ARC-AGI-21.1%

    65th of 79

  • Mock AIME66.4%

    99th of 176

Everything else Epoch has for it

  • MATH level 596.6%
  • Lech Mazur Writing81.9%
  • Fiction.LiveBench75.0%
  • Aider polyglot71.4%
  • METR Time Horizons53.8%
  • WeirdML41.6%
  • DeepResearch Bench35.1%
  • SimpleBench29.0%
  • ARC-AGI21.2%

On the index

Around it.

  1. #108Claude Sonnet 4Anthropic · $3.00 in141.7
  2. #109Gemini 2.5 Flash (May 2025)Google141.5
  3. #110Mistral Medium 3.5Mistral · $1.50 in141.3
  4. #111DeepSeek-R1 (May 2025)DeepSeek141.3
  5. #112Claude 3.7 SonnetAnthropic141.2
  6. #113Gemini 2.5 Flash (Jun 2025)Google140.8
  7. #114Grok-3 minixAI140.3

Sources