in the loop

o4-mini
#83 of 274 scored.

OpenAIApr 16, 2025Closed

Capabilities Index

Range 143.2–147.4

145.6

Context

Takes image, text, file

200K

Input

Per 1M tokens

$1.10

Output

Per 1M tokens

$4.40

Benchmarks

How it scores.

  • GPQA Diamond72.8%

    82nd of 186

  • FrontierMath36.1%

    55th of 81

  • ARC-AGI-26.1%

    48th of 79

  • Humanity's Last Exam14.0%

    21st of 40

  • SimpleQA Verified19.6%

    66th of 77

  • Mock AIME81.7%

    78th of 176

Everything else Epoch has for it

  • MATH level 597.8%
  • Fiction.LiveBench77.8%
  • Lech Mazur Writing75.0%
  • Aider polyglot72.0%
  • GeoBench64.0%
  • METR Time Horizons63.9%
  • DTBench62.7%
  • CadEval62.0%
  • ARC-AGI58.7%
  • WeirdML52.6%
  • VPCT36.3%
  • LMCA31.2%
  • SimpleBench26.4%
  • GDPval25.3%
  • Chess Puzzles22.1%
  • FrontierMath-Tier-4-v2-Private4.9%
  • GSO-Bench3.6%
  • Mystery Game Puzzles0.0%

On the index

Around it.

  1. #80MiniMax-M2.7MiniMax · $0.21 in145.8
  2. #81GLM-5Z.ai · $0.60 in145.8
  3. #82GPT-5.4 NanoOpenAI · $0.20 in145.8
  4. #83o4-miniOpenAI · $1.10 in145.6
  5. #84GPT-5 miniOpenAI · $0.25 in145.5
  6. #85Gemini 2.5 Pro (Jun 2025)Google145.3
  7. #86Gemini 3.5 Flash-LiteGoogle · $0.30 in145.1

Sources