in the loop

Claude Sonnet 4.5
#70 of 274 scored.

AnthropicSep 29, 2025Closed

Capabilities Index

Range 145.1–148.3

146.8

Context

Takes text, image, file

1M

Input

Per 1M tokens

$3.00

Output

Per 1M tokens

$15

Benchmarks

How it scores.

  • GPQA Diamond76.4%

    75th of 186

  • SWE-bench Verified71.3%

    23rd of 32

  • FrontierMath23.9%

    66th of 81

  • ARC-AGI-213.6%

    42nd of 79

  • Humanity's Last Exam9.4%

    23rd of 40

  • SimpleQA Verified30.7%

    59th of 77

  • Mock AIME77.8%

    82nd of 176

  • Terminal-Bench46.5%

    16th of 35

Everything else Epoch has for it

  • MATH level 597.7%
  • DTBench72.0%
  • METR Time Horizons67.4%
  • ARC-AGI63.7%
  • OSWorld62.9%
  • Cybench60.0%
  • DeepResearch Bench52.6%
  • WeirdML47.7%
  • LMCA45.6%
  • SimpleBench45.2%
  • GDPval42.5%
  • ProofBench19.0%
  • GSO-Bench14.7%
  • VPCT9.7%
  • Mystery Game Puzzles8.6%
  • Chess Puzzles7.4%
  • FrontierMath-Tier-4-v2-Private2.4%
  • EBR-bench2.4%
  • Remote Labor Index2.1%

On the index

Around it.

  1. #67Qwen3.7-PlusAlibaba · $0.32 in147.4
  2. #68MiniMax-M3MiniMax · $0.30 in146.9
  3. #69o3OpenAI · $2.00 in146.9
  4. #70Claude Sonnet 4.5Anthropic · $3.00 in146.8
  5. #71Qwen 3.5 Plus (hosted 397B-A17B)Alibaba146.8
  6. #72MiniMax-M2.5MiniMax · $0.27 in146.7
  7. #73Qwen3.5 397B-A17BAlibaba · $0.55 in146.7

Sources