in the loop

Claude Opus 4
#100 of 274 scored.

AnthropicMay 22, 2025Closed

Capabilities Index

Range 140.1–144.3

142.7

Context

Tokens

—

Input

Per 1M tokens

—

Output

Per 1M tokens

—

Benchmarks

How it scores.

  • GPQA Diamond68.3%

    90th of 186

  • SWE-bench Verified70.7%

    24th of 32

  • ARC-AGI-28.6%

    46th of 79

  • Humanity's Last Exam6.2%

    27th of 40

  • Mock AIME64.4%

    102nd of 176

Everything else Epoch has for it

  • MATH level 585.0%
  • Lech Mazur Writing83.6%
  • Aider polyglot72.0%
  • DTBench69.3%
  • METR Time Horizons63.9%
  • Fiction.LiveBench61.1%
  • SimpleBench50.6%
  • GeoBench49.0%
  • DeepResearch Bench46.8%
  • LMCA44.0%
  • WeirdML43.7%
  • Cybench38.0%
  • ARC-AGI35.7%
  • FrontierMath-2025-02-28-Private7.9%
  • VPCT7.0%
  • FrontierMath-Tier-4-2025-07-01-Private6.9%
  • GSO-Bench6.9%

On the index

Around it.

  1. #97Qwen 3.6 FlashAlibaba · $0.19 in143.3
  2. #98Gemini 2.5 Flash (Sep 2025)Google143.0
  3. #99Gemma 4 31B ITGoogle · $0.09 in142.7
  4. #100Claude Opus 4Anthropic142.7
  5. #101GPT-5.5 InstantOpenAI142.5
  6. #102Qwen3.5-35B-A3BAlibaba · $0.08 in142.5
  7. #103Gemini 2.5 Pro (May 2025)Google142.5

Sources