in the loop

Claude Opus 4.6
#30 of 274 scored.

AnthropicFeb 5, 2026Closed

Capabilities Index

Range 153.4–157.3

155.2

Context

Takes text, image, file

1M

Input

Per 1M tokens

$5.00

Output

Per 1M tokens

$25

Benchmarks

How it scores.

  • GPQA Diamond87.4%

    36th of 186

  • SWE-bench Verified78.7%

    4th of 32

  • FrontierMath66.0%

    27th of 81

  • ARC-AGI-269.2%

    19th of 79

  • Humanity's Last Exam31.1%

    10th of 40

  • SimpleQA Verified47.0%

    32nd of 77

  • Mock AIME94.4%

    37th of 176

  • Terminal-Bench79.8%

    5th of 35

Everything else Epoch has for it

  • ARC-AGI94.0%
  • Cybench93.0%
  • DTBench85.3%
  • METR Time Horizons78.9%
  • WeirdML78.0%
  • LMCA65.6%
  • SimpleBench61.1%
  • DeepResearch Bench55.3%
  • ProofBench50.0%
  • APEX-Agents46.3%
  • GSO-Bench41.2%
  • FrontierMath-Tier-4-v2-Private26.8%
  • FrontierCode26.6%
  • CL-bench20.7%
  • Mystery Game Puzzles17.4%
  • CL-bench Life17.0%
  • EBR-bench12.7%
  • Chess Puzzles12.7%
  • Remote Labor Index4.2%
  • Furniture Assembly0.0%

On the index

Around it.

  1. #27GLM-5.3Z.ai · $0.039 in155.6
  2. #28GPT-5.2 ProOpenAI · $21 in155.4
  3. #29DeepSeek V4 Pro 0813DeepSeek · $0.66 in155.3
  4. #30Claude Opus 4.6Anthropic · $5.00 in155.2
  5. #31Qwen3.8 Max (0902)Alibaba · $2.00 in155.1
  6. #32DeepSeek V4.1 FlashDeepSeek · $0.30 in154.9
  7. #33Muse Spark 1.2Meta · $1.25 in154.9

Sources