in the loop

GPT-4 Turbo (Apr 2024)
#173 of 274 scored.

OpenAIApr 9, 2024Closed

Capabilities Index

Range 121.9–128.8

127.3

Context

Tokens

—

Input

Per 1M tokens

—

Output

Per 1M tokens

—

Benchmarks

How it scores.

  • GPQA Diamond28.8%

    148th of 186

  • FrontierMath0.7%

    78th of 81

  • Mock AIME6.6%

    145th of 176

Everything else Epoch has for it

  • MMLU75.1%
  • MATH level 546.7%
  • METR Time Horizons36.7%
  • WeirdML18.0%
  • SimpleBench10.1%
  • Chess Puzzles1.1%

On the index

Around it.

  1. #170Mistral Large 2 (Jul 2024)Mistral127.5
  2. #171Mistral Small 3.1Mistral127.5
  3. #172Llama 3.3 70BMeta127.3
  4. #173GPT-4 Turbo (Apr 2024)OpenAI127.3
  5. #174Claude 3.5 HaikuAnthropic127.2
  6. #175Mistral Small 3Mistral · $0.05 in127.1
  7. #176Gemini 1.5 Pro (May 2024)Google126.9

Sources