in the loop

Claude 3.5 Sonnet (October 2024)
#148 of 274 scored.

AnthropicOct 22, 2024Closed

Capabilities Index

Range 128.8–137.6

133.6

Context

Tokens

—

Input

Per 1M tokens

—

Output

Per 1M tokens

—

Benchmarks

How it scores.

  • GPQA Diamond40.4%

    127th of 186

  • Humanity's Last Exam0.0%

    36th of 40

  • Mock AIME8.4%

    136th of 176

Everything else Epoch has for it

  • MMLU83.1%
  • Lech Mazur Writing80.3%
  • GeoBench62.0%
  • MATH level 557.0%
  • Aider polyglot51.6%
  • CadEval48.0%
  • DTBench46.4%
  • METR Time Horizons45.2%
  • WeirdML40.0%
  • Balrog32.6%
  • SimpleBench29.7%
  • The Agent Company24.0%
  • GSO-Bench4.6%
  • FrontierMath-2025-02-28-Private3.6%
  • FrontierMath-Tier-4-2025-07-01-Private0.0%
  • VPCT0.0%

On the index

Around it.

  1. #145Gemini 2.0 Flash (Dec 2024)Google134.7
  2. #146Mistral Medium 3Mistral · $0.40 in134.1
  3. #147Gemini 2.5 Flash-Lite (Jun 2025)Google133.9
  4. #148Claude 3.5 Sonnet (October 2024)Anthropic133.6
  5. #149Magistral Small 1.0Mistral133.2
  6. #150Qwen2.5-MaxAlibaba132.5
  7. #151DeepSeek-V3DeepSeek · $0.32 in132.3

Sources