in the loop

DeepSeek-R1
#124 of 274 scored.

DeepSeekJan 20, 2025Open weights

Capabilities Index

Range 136.3–140.3

139.0

Context

Takes text

64K

Input

Per 1M tokens

$0.70

Output

Per 1M tokens

$2.50

Benchmarks

How it scores.

  • GPQA Diamond62.3%

    98th of 186

  • ARC-AGI-21.3%

    62nd of 79

  • Mock AIME53.3%

    112th of 176

Everything else Epoch has for it

  • MATH level 593.0%
  • Lech Mazur Writing83.0%
  • Fiction.LiveBench69.4%
  • Aider polyglot56.9%
  • METR Time Horizons51.9%
  • WeirdML36.5%
  • Balrog34.9%
  • SimpleBench17.1%
  • ARC-AGI15.8%

On the index

Around it.

  1. #121Qwen3.5-9BAlibaba · $0.10 in139.5
  2. #122GPT-5 nanoOpenAI · $0.05 in139.4
  3. #123Qwen3-235B-A22BAlibaba139.3
  4. #124DeepSeek-R1DeepSeek · $0.70 in139.0
  5. #125Qwen3-235B-A22B-Instruct (Jul 2025)Alibaba138.9
  6. #126Qwen3-32BAlibaba · $0.08 in138.5
  7. #127Grok 3xAI138.3

Sources