in the loop

Claude Opus 4.1
#92 of 274 scored.

AnthropicAug 5, 2025Closed

Capabilities Index

Range 141.7–146.0

144.1

Context

Takes image, text, file

200K

Input

Per 1M tokens

$15

Output

Per 1M tokens

$75

Benchmarks

How it scores.

  • GPQA Diamond69.7%

    86th of 186

  • SWE-bench Verified73.4%

    20th of 32

  • FrontierMath12.6%

    75th of 81

  • Humanity's Last Exam7.1%

    25th of 40

  • Mock AIME68.9%

    95th of 176

  • Terminal-Bench38.0%

    21st of 35

Everything else Epoch has for it

  • Lech Mazur Writing84.7%
  • METR Time Horizons66.8%
  • DTBench66.7%
  • SimpleBench52.0%
  • DeepResearch Bench48.3%
  • WeirdML45.9%
  • LMCA43.6%
  • GDPval43.6%
  • Cybench42.0%
  • Mystery Game Puzzles13.0%
  • EBR-bench7.9%
  • VPCT2.5%
  • FrontierMath-Tier-4-v2-Private2.4%
  • Chess Puzzles2.1%

On the index

Around it.

  1. #89Gemini 3.1 Flash-LiteGoogle · $0.25 in144.5
  2. #90Grok 4 FastxAI144.2
  3. #91Gemini 2.5 Pro (Mar 2025)Google144.2
  4. #92Claude Opus 4.1Anthropic · $15 in144.1
  5. #93Qwen 3.5 Flash (hosted 35B-A3B)Alibaba144.0
  6. #94Qwen 3.6 35B-A3BAlibaba · $0.15 in143.9
  7. #95Qwen3-235B-A22B-Thinking (Jul 2025)Alibaba143.8

Sources