in the loop

Claude 3.7 Sonnet
#112 of 274 scored.

AnthropicFeb 24, 2025Closed

Capabilities Index

Range 138.8–142.7

141.2

Context

Tokens

—

Input

Per 1M tokens

—

Output

Per 1M tokens

—

Benchmarks

How it scores.

  • GPQA Diamond73.0%

    81st of 186

  • SWE-bench Verified61.0%

    28th of 32

  • ARC-AGI-20.9%

    66th of 79

  • Humanity's Last Exam3.4%

    29th of 40

  • Mock AIME57.7%

    107th of 176

Everything else Epoch has for it

  • MATH level 591.2%
  • Fiction.LiveBench83.3%
  • Lech Mazur Writing81.1%
  • GeoBench68.0%
  • Aider polyglot64.9%
  • METR Time Horizons60.0%
  • CadEval54.0%
  • DeepResearch Bench43.6%
  • OSWorld35.8%
  • SimpleBench35.7%
  • The Agent Company30.9%
  • ARC-AGI28.6%
  • Cybench20.0%
  • VPCT8.5%
  • FrontierMath-2025-02-28-Private7.3%
  • GSO-Bench3.8%

On the index

Around it.

  1. #109Gemini 2.5 Flash (May 2025)Google141.5
  2. #110Mistral Medium 3.5Mistral · $1.50 in141.3
  3. #111DeepSeek-R1 (May 2025)DeepSeek141.3
  4. #112Claude 3.7 SonnetAnthropic141.2
  5. #113Gemini 2.5 Flash (Jun 2025)Google140.8
  6. #114Grok-3 minixAI140.3
  7. #115o3-miniOpenAI · $1.10 in140.3

Sources