in the loop

Claude Sonnet 4
#108 of 274 scored.

AnthropicMay 22, 2025Closed

Capabilities Index

Range 138.9–143.0

141.7

Context

Takes image, text, file

200K

Input

Per 1M tokens

$3.00

Output

Per 1M tokens

$15

Benchmarks

How it scores.

  • GPQA Diamond72.3%

    83rd of 186

  • ARC-AGI-25.9%

    49th of 79

  • Humanity's Last Exam3.1%

    31st of 40

  • Mock AIME71.1%

    89th of 176

Everything else Epoch has for it

  • MATH level 584.4%
  • Lech Mazur Writing81.4%
  • METR Time Horizons62.0%
  • DTBench61.8%
  • Aider polyglot61.3%
  • Fiction.LiveBench46.9%
  • DeepResearch Bench46.6%
  • WeirdML46.1%
  • OSWorld43.9%
  • ARC-AGI40.0%
  • GeoBench37.0%
  • Cybench35.0%
  • SimpleBench34.6%
  • LMCA34.1%
  • The Agent Company33.1%
  • FrontierMath-2025-02-28-Private7.3%
  • GSO-Bench4.9%
  • VPCT1.0%
  • FrontierMath-Tier-4-2025-07-01-Private0.0%

On the index

Around it.

  1. #105Qwen3-MaxAlibaba142.4
  2. #106o1OpenAI · $15 in141.9
  3. #107Gemma 4 26B A4BGoogle · $0.068 in141.8
  4. #108Claude Sonnet 4Anthropic · $3.00 in141.7
  5. #109Gemini 2.5 Flash (May 2025)Google141.5
  6. #110Mistral Medium 3.5Mistral · $1.50 in141.3
  7. #111DeepSeek-R1 (May 2025)DeepSeek141.3

Sources