in the loop

DeepSeek-V3 (Mar 2025)
#137 of 274 scored.

DeepSeekMar 24, 2025Open weights

Capabilities Index

Range 132.1–137.8

135.9

Context

Tokens

—

Input

Per 1M tokens

—

Output

Per 1M tokens

—

Benchmarks

How it scores.

  • GPQA Diamond56.8%

    104th of 186

  • Mock AIME37.7%

    117th of 176

Everything else Epoch has for it

  • Lech Mazur Writing77.0%
  • MATH level 575.5%
  • Aider polyglot55.1%
  • Fiction.LiveBench50.0%
  • METR Time Horizons49.6%
  • DTBench41.3%
  • WeirdML36.1%
  • LMCA18.2%
  • SimpleBench12.6%

On the index

Around it.

  1. #134GPT-4.5OpenAI136.7
  2. #135Qwen3-30B-A3BAlibaba · $0.12 in136.2
  3. #136Qwen3-8BAlibaba136.2
  4. #137DeepSeek-V3 (Mar 2025)DeepSeek135.9
  5. #138o1-miniOpenAI135.8
  6. #139DeepSeek-R1-Distill-Qwen-14BDeepSeek135.4
  7. #140Gemini 2.0 Flash Thinking (Jan 2025)Google135.4

Sources