in the loop

Claude Opus 4.5
#53 of 274 scored.

AnthropicNov 24, 2025Closed

Capabilities Index

Range 148.0–152.8

150.1

Context

Takes file, image, text

200K

Input

Per 1M tokens

$5.00

Output

Per 1M tokens

$25

Benchmarks

How it scores.

  • GPQA Diamond81.4%

    60th of 186

  • SWE-bench Verified76.6%

    9th of 32

  • FrontierMath34.4%

    57th of 81

  • ARC-AGI-237.6%

    33rd of 79

  • Humanity's Last Exam21.4%

    14th of 40

  • SimpleQA Verified45.7%

    37th of 77

  • Mock AIME86.1%

    68th of 176

  • Terminal-Bench63.1%

    10th of 35

Everything else Epoch has for it

  • DTBench83.1%
  • Cybench82.0%
  • ARC-AGI80.0%
  • GeoBench75.0%
  • METR Time Horizons75.0%
  • OSWorld66.3%
  • WeirdML63.7%
  • DeepResearch Bench54.8%
  • SimpleBench54.4%
  • LMCA52.3%
  • GDPval45.5%
  • Balrog43.5%
  • ProofBench36.0%
  • GSO-Bench26.5%
  • CL-bench21.1%
  • EBR-bench14.3%
  • Mystery Game Puzzles14.1%
  • VPCT10.0%
  • Chess Puzzles7.4%
  • FrontierMath-Tier-4-v2-Private4.9%
  • Remote Labor Index3.8%
  • Furniture Assembly0.0%

On the index

Around it.

  1. #50Kimi K2.6Moonshot · $0.47 in151.1
  2. #51GPT-5 ProOpenAI · $15 in150.3
  3. #52Inkling-SmallThinking Machines · $0.45 in150.2
  4. #53Claude Opus 4.5Anthropic · $5.00 in150.1
  5. #54GPT-5OpenAI · $1.25 in150.0
  6. #55Kimi K2.7 CodeMoonshot · $0.67 in150.0
  7. #56GLM-5.1Z.ai · $0.97 in149.8

Sources