in the loop

Grok 3
#127 of 274 scored.

xAIApr 9, 2025Closed

Capabilities Index

Range 135.3–139.8

138.3

Context

Tokens

—

Input

Per 1M tokens

—

Output

Per 1M tokens

—

Benchmarks

How it scores.

  • GPQA Diamond67.7%

    92nd of 186

  • ARC-AGI-20.0%

    72nd of 79

  • Mock AIME55.5%

    110th of 176

Everything else Epoch has for it

  • MATH level 588.8%
  • Lech Mazur Writing76.4%
  • Fiction.LiveBench58.3%
  • Aider polyglot53.3%
  • WeirdML37.2%
  • Balrog29.5%
  • SimpleBench23.3%
  • FrontierMath-2025-02-28-Private6.7%
  • ARC-AGI5.5%
  • FrontierMath-Tier-4-2025-07-01-Private0.0%

On the index

Around it.

  1. #124DeepSeek-R1DeepSeek · $0.70 in139.0
  2. #125Qwen3-235B-A22B-Instruct (Jul 2025)Alibaba138.9
  3. #126Qwen3-32BAlibaba · $0.08 in138.5
  4. #127Grok 3xAI138.3
  5. #128Qwen3-14BAlibaba · $0.10 in138.2
  6. #129gpt-oss-20bOpenAI · $0.018 in137.8
  7. #130QwQ-32BAlibaba137.6

Sources