in the loop

GPT-4o (Nov 2024)
#165 of 274 scored.

OpenAINov 20, 2024Closed

Capabilities Index

Range 123.6–131.2

128.8

Context

Tokens

—

Input

Per 1M tokens

—

Output

Per 1M tokens

—

Benchmarks

How it scores.

  • GPQA Diamond30.5%

    140th of 186

  • SWE-bench Verified31.0%

    32nd of 32

  • ARC-AGI-20.0%

    72nd of 79

  • Humanity's Last Exam0.0%

    36th of 40

  • Mock AIME6.2%

    149th of 176

Everything else Epoch has for it

  • MMLU84.1%
  • Lech Mazur Writing81.8%
  • GeoBench71.0%
  • MATH level 549.8%
  • METR Time Horizons40.8%
  • WeirdML25.1%
  • Aider polyglot18.2%
  • Cybench12.5%
  • VPCT10.0%
  • GDPval9.9%
  • The Agent Company8.6%
  • ARC-AGI4.5%
  • FrontierMath-2025-02-28-Private0.6%
  • GSO-Bench0.0%

On the index

Around it.

  1. #162Gemini 1.5 Flash (Sep 2024)Google129.4
  2. #163Qwen2.5-72BAlibaba129.0
  3. #164GPT-4o (May 2024)OpenAI129.0
  4. #165GPT-4o (Nov 2024)OpenAI128.8
  5. #166GPT-4o (Aug 2024)OpenAI128.8
  6. #167Llama 3.1-405BMeta128.8
  7. #168Mistral Large 2 (Nov 2024)Mistral128.5

Sources