in the loop

Grok-3 mini
#114 of 274 scored.

xAIJun 24, 2025Closed

Capabilities Index

Range 137.5–142.0

140.3

Context

Tokens

—

Input

Per 1M tokens

—

Output

Per 1M tokens

—

Benchmarks

How it scores.

  • GPQA Diamond68.3%

    90th of 186

  • ARC-AGI-20.4%

    70th of 79

  • Mock AIME77.8%

    82nd of 176

Everything else Epoch has for it

  • MATH level 590.9%
  • Lech Mazur Writing73.5%
  • Fiction.LiveBench66.7%
  • Aider polyglot49.3%
  • WeirdML42.6%
  • ARC-AGI16.5%
  • FrontierMath-2025-02-28-Private10.3%

On the index

Around it.

  1. #111DeepSeek-R1 (May 2025)DeepSeek141.3
  2. #112Claude 3.7 SonnetAnthropic141.2
  3. #113Gemini 2.5 Flash (Jun 2025)Google140.8
  4. #114Grok-3 minixAI140.3
  5. #115o3-miniOpenAI · $1.10 in140.3
  6. #116Kimi K2 (Jul 2025)Moonshot140.1
  7. #117Gemini 2.5 Flash (Apr 2025)Google140.0

Sources