in the loop

Claude Opus 4.8
#14 of 274 scored.

AnthropicMay 28, 2026Closed

Capabilities Index

Range 156.0–160.8

158.2

Context

Takes text, image, file

1M

Input

Per 1M tokens

$5.00

Output

Per 1M tokens

$25

Benchmarks

How it scores.

  • GPQA Diamond88.0%

    28th of 186

  • FrontierMath80.0%

    15th of 81

  • ARC-AGI-272.1%

    17th of 79

  • SimpleQA Verified53.0%

    20th of 77

  • Mock AIME98.3%

    20th of 176

Everything else Epoch has for it

  • ARC-AGI92.5%
  • DTBench91.5%
  • Surface Evolver Bench87.5%
  • WeirdML82.9%
  • ProofBench69.0%
  • LMCA67.7%
  • DeepSWE59.0%
  • SimpleBench57.8%
  • FrontierMath-Tier-4-v2-Private56.1%
  • DeepResearch Bench50.2%
  • APEX-Agents48.9%
  • GSO-Bench47.1%
  • FrontierCode46.5%
  • PostTrainBench33.8%
  • Chess Puzzles30.6%
  • Mystery Game Puzzles29.5%
  • EBR-bench28.6%
  • OSWorld 2.020.6%
  • Furniture Assembly17.9%
  • Remote Labor Index8.3%

On the index

Around it.

  1. #11GPT-5.6 TerraOpenAI · $2.00 in159.6
  2. #12GPT-5.5OpenAI · $5.00 in159.1
  3. #13GPT-5.4 ProOpenAI · $30 in158.9
  4. #14Claude Opus 4.8Anthropic · $5.00 in158.2
  5. #15Kimi K3Moonshot · $0.80 in157.4
  6. #16Gemini 3.7 FlashGoogle · $0.75 in157.3
  7. #17GPT-5.4OpenAI · $2.50 in156.8

Sources