in the loop

Claude Sonnet 4.6
#44 of 274 scored.

AnthropicFeb 17, 2026Closed

Capabilities Index

Range 149.9–154.5

152.2

Context

Takes text, image, file

1M

Input

Per 1M tokens

$3.00

Output

Per 1M tokens

$15

Benchmarks

How it scores.

  • GPQA Diamond83.2%

    54th of 186

  • SWE-bench Verified75.2%

    14th of 32

  • ARC-AGI-260.4%

    24th of 79

  • SimpleQA Verified35.5%

    49th of 77

  • Mock AIME85.8%

    69th of 176

  • Terminal-Bench53.4%

    12th of 35

Everything else Epoch has for it

  • ARC-AGI86.5%
  • DTBench83.1%
  • OSWorld72.1%
  • WeirdML66.1%
  • FrontierMath-2025-02-28-Private56.8%
  • DeepResearch Bench54.9%
  • LMCA54.7%
  • ProofBench45.0%
  • APEX-Agents43.0%
  • DeepSWE29.9%
  • FrontierCode24.3%
  • ExploitBench23.6%
  • FrontierMath-Tier-4-2025-07-01-Private13.8%
  • OSWorld 2.09.3%
  • Chess Puzzles8.5%
  • Mystery Game Puzzles7.5%

On the index

Around it.

  1. #41Grok 4.7xAI · $2.00 in153.5
  2. #42GPT-5.2OpenAI · $1.75 in153.4
  3. #43Gemini 3 ProGoogle152.9
  4. #44Claude Sonnet 4.6Anthropic · $3.00 in152.2
  5. #45Muse SparkMeta152.0
  6. #46Grok 4.20xAI · $1.25 in152.0
  7. #47GLM-5.3-FlashZ.ai · $0.15 in151.9

Sources