in the loop

Claude 3.5 Sonnet
#159 of 274 scored.

AnthropicJun 20, 2024Closed

Capabilities Index

Range undefined–undefined

130.0

Context

Tokens

—

Input

Per 1M tokens

—

Output

Per 1M tokens

—

Benchmarks

How it scores.

  • GPQA Diamond38.7%

    128th of 186

  • Mock AIME6.4%

    147th of 176

Everything else Epoch has for it

  • MMLU82.0%
  • MATH level 551.7%
  • DTBench46.3%
  • METR Time Horizons40.2%
  • WeirdML31.0%
  • Cybench17.5%
  • SimpleBench13.0%
  • FrontierMath-2025-02-28-Private1.8%
  • FrontierMath-Tier-4-2025-07-01-Private0.0%
  • VPCT0.0%

On the index

Around it.

  1. #156Grok-2 (Dec 2024)xAI130.5
  2. #157Phi-4Microsoft · $0.07 in130.4
  3. #158Gemma 3 27BGoogle · $0.08 in130.0
  4. #159Claude 3.5 SonnetAnthropic130.0
  5. #160Llama 4 ScoutMeta · $0.10 in129.6
  6. #161GPT-4.1 nanoOpenAI · $0.10 in129.6
  7. #162Gemini 1.5 Flash (Sep 2024)Google129.4

Sources