in the loop

Claude 3 Opus
#177 of 274 scored.

AnthropicFeb 29, 2024Closed

Capabilities Index

Range 121.5–130.0

126.9

Context

Tokens

—

Input

Per 1M tokens

—

Output

Per 1M tokens

—

Benchmarks

How it scores.

  • GPQA Diamond29.5%

    147th of 186

  • SimpleQA Verified12.6%

    72nd of 77

  • Mock AIME4.6%

    153rd of 176

Everything else Epoch has for it

  • MMLU79.5%
  • Winogrande77.0%
  • MATH level 537.5%
  • DTBench36.0%
  • METR Time Horizons29.5%
  • LMCA20.0%
  • WeirdML19.2%
  • Cybench10.0%
  • SimpleBench8.2%
  • Chess Puzzles0.0%

On the index

Around it.

  1. #174Claude 3.5 HaikuAnthropic127.2
  2. #175Mistral Small 3Mistral · $0.05 in127.1
  3. #176Gemini 1.5 Pro (May 2024)Google126.9
  4. #177Claude 3 OpusAnthropic126.9
  5. #178GPT-4o miniOpenAI · $0.15 in126.6
  6. #179GPT-4 Turbo (Nov 2023)OpenAI126.5
  7. #180Llama 3.1-70BMeta125.9

Sources