in the loop

o3
#69 of 274 scored.

OpenAIApr 16, 2025Closed

Capabilities Index

Range 145.0–148.7

146.9

Context

Takes image, text, file

200K

Input

Per 1M tokens

$2.00

Output

Per 1M tokens

$8.00

Benchmarks

How it scores.

  • GPQA Diamond75.8%

    78th of 186

  • SWE-bench Verified62.3%

    27th of 32

  • FrontierMath33.3%

    59th of 81

  • ARC-AGI-26.5%

    47th of 79

  • Humanity's Last Exam16.3%

    18th of 40

  • SimpleQA Verified49.4%

    26th of 77

  • Mock AIME84.4%

    71st of 176

Everything else Epoch has for it

  • MATH level 597.8%
  • Fiction.LiveBench88.9%
  • Lech Mazur Writing83.9%
  • Aider polyglot81.3%
  • DTBench74.7%
  • CadEval74.0%
  • GeoBench74.0%
  • METR Time Horizons65.4%
  • ARC-AGI60.8%
  • WeirdML52.4%
  • LMCA46.7%
  • DeepResearch Bench45.2%
  • SimpleBench43.7%
  • Chess Puzzles34.8%
  • GDPval30.8%
  • VPCT28.0%
  • OSWorld23.0%
  • Mystery Game Puzzles21.8%
  • CL-bench17.8%
  • GSO-Bench8.8%
  • FrontierMath-Tier-4-2025-07-01-Private3.5%

On the index

Around it.

  1. #66o3-proOpenAI · $20 in147.4
  2. #67Qwen3.7-PlusAlibaba · $0.32 in147.4
  3. #68MiniMax-M3MiniMax · $0.30 in146.9
  4. #69o3OpenAI · $2.00 in146.9
  5. #70Claude Sonnet 4.5Anthropic · $3.00 in146.8
  6. #71Qwen 3.5 Plus (hosted 397B-A17B)Alibaba146.8
  7. #72MiniMax-M2.5MiniMax · $0.27 in146.7

Sources