in the loop

GPT-4o (Aug 2024)
#166 of 274 scored.

OpenAIAug 6, 2024Closed

Capabilities Index

Range 123.6–130.7

128.8

Context

Tokens

—

Input

Per 1M tokens

—

Output

Per 1M tokens

—

Benchmarks

How it scores.

  • GPQA Diamond32.3%

    134th of 186

  • FrontierMath0.4%

    80th of 81

  • SimpleQA Verified26.0%

    62nd of 77

  • Mock AIME6.3%

    148th of 176

Everything else Epoch has for it

  • MMLU79.1%
  • MATH level 553.3%
  • METR Time Horizons33.8%
  • CadEval26.0%
  • Aider polyglot23.1%
  • Chess Puzzles8.5%
  • SimpleBench1.4%

On the index

Around it.

  1. #163Qwen2.5-72BAlibaba129.0
  2. #164GPT-4o (May 2024)OpenAI129.0
  3. #165GPT-4o (Nov 2024)OpenAI128.8
  4. #166GPT-4o (Aug 2024)OpenAI128.8
  5. #167Llama 3.1-405BMeta128.8
  6. #168Mistral Large 2 (Nov 2024)Mistral128.5
  7. #169Qwen2.5-32BAlibaba128.5

Sources