in the loop

DeepSeek-V3
#151 of 274 scored.

DeepSeekDec 26, 2024Open weights

Capabilities Index

Range 127.5–135.7

132.3

Context

Takes text

164K

Input

Per 1M tokens

$0.32

Output

Per 1M tokens

$0.89

Benchmarks

How it scores.

  • GPQA Diamond42.0%

    122nd of 186

  • Mock AIME15.8%

    132nd of 176

Everything else Epoch has for it

  • ARC AI293.7%
  • HellaSwag85.2%
  • BBH83.3%
  • MMLU82.9%
  • TriviaQA82.9%
  • Winogrande70.4%
  • PIQA69.4%
  • MATH level 564.8%
  • Aider polyglot48.4%
  • METR Time Horizons47.4%
  • FrontierMath-2025-02-28-Private3.0%
  • SimpleBench2.7%

On the index

Around it.

  1. #148Claude 3.5 Sonnet (October 2024)Anthropic133.6
  2. #149Magistral Small 1.0Mistral133.2
  3. #150Qwen2.5-MaxAlibaba132.5
  4. #151DeepSeek-V3DeepSeek · $0.32 in132.3
  5. #152Llama 4 MaverickMeta · $0.19 in132.2
  6. #153Mistral Small 3.2Mistral131.7
  7. #154Gemini 1.5 Pro (Sept 2024)Google131.7

Sources