in the loop

o3-mini
#115 of 274 scored.

OpenAIJan 31, 2025Closed

Capabilities Index

Range 137.7–141.8

140.3

Context

Takes text, file

200K

Input

Per 1M tokens

$1.10

Output

Per 1M tokens

$4.40

Benchmarks

How it scores.

  • GPQA Diamond69.4%

    87th of 186

  • FrontierMath18.6%

    72nd of 81

  • ARC-AGI-23.0%

    59th of 79

  • SimpleQA Verified15.3%

    69th of 77

  • Mock AIME76.9%

    84th of 176

Everything else Epoch has for it

  • MATH level 596.5%
  • Lech Mazur Writing61.7%
  • Aider polyglot60.4%
  • CadEval54.0%
  • Fiction.LiveBench50.0%
  • DTBench48.0%
  • WeirdML43.7%
  • ARC-AGI34.5%
  • Cybench22.5%
  • LMCA22.3%
  • Chess Puzzles12.7%
  • SimpleBench7.4%
  • GSO-Bench1.3%
  • FrontierMath-Tier-4-v2-Private0.0%
  • Mystery Game Puzzles0.0%

On the index

Around it.

  1. #112Claude 3.7 SonnetAnthropic141.2
  2. #113Gemini 2.5 Flash (Jun 2025)Google140.8
  3. #114Grok-3 minixAI140.3
  4. #115o3-miniOpenAI · $1.10 in140.3
  5. #116Kimi K2 (Jul 2025)Moonshot140.1
  6. #117Gemini 2.5 Flash (Apr 2025)Google140.0
  7. #118gpt-oss-120bOpenAI · $0.037 in139.9

Sources