DeepSeek-R1 (May 2025)
#111 of 274 scored.
DeepSeekMay 28, 2025Open weights
Capabilities Index
Range 138.6–142.8
141.3
Context
Tokens
—
Input
Per 1M tokens
—
Output
Per 1M tokens
—
Benchmarks
How it scores.
- GPQA Diamond68.4%
89th of 186
- ARC-AGI-21.1%
65th of 79
- Mock AIME66.4%
99th of 176
Everything else Epoch has for it
- MATH level 596.6%
- Lech Mazur Writing81.9%
- Fiction.LiveBench75.0%
- Aider polyglot71.4%
- METR Time Horizons53.8%
- WeirdML41.6%
- DeepResearch Bench35.1%
- SimpleBench29.0%
- ARC-AGI21.2%
On the index
Around it.
Sources