in the loop

GPT-5.1
#57 of 274 scored.

OpenAINov 13, 2025Closed

Capabilities Index

Range 148.2–151.1

149.6

Context

Takes image, text, file

400K

Input

Per 1M tokens

$1.25

Output

Per 1M tokens

$10

Benchmarks

How it scores.

  • GPQA Diamond83.5%

    52nd of 186

  • SWE-bench Verified68.0%

    25th of 32

  • ARC-AGI-217.6%

    40th of 79

  • Humanity's Last Exam19.8%

    16th of 40

  • SimpleQA Verified48.0%

    30th of 77

  • Mock AIME88.6%

    58th of 176

  • Terminal-Bench47.6%

    15th of 35

Everything else Epoch has for it

  • DTBench83.5%
  • ARC-AGI72.8%
  • WeirdML60.8%
  • FrontierMath-2025-02-28-Private54.4%
  • LMCA51.6%
  • SimpleBench43.8%
  • DeepResearch Bench42.8%
  • VPCT38.0%
  • Chess Puzzles28.4%
  • CL-bench23.7%
  • FrontierMath-Tier-4-2025-07-01-Private20.8%
  • CL-bench Life17.3%
  • GSO-Bench13.7%
  • Mystery Game Puzzles10.8%

On the index

Around it.

  1. #54GPT-5OpenAI · $1.25 in150.0
  2. #55Kimi K2.7 CodeMoonshot · $0.67 in150.0
  3. #56GLM-5.1Z.ai · $0.97 in149.8
  4. #57GPT-5.1OpenAI · $1.25 in149.6
  5. #58Qwen 3.8 27BAlibaba · $0.42 in149.4
  6. #59Qwen 3.6 Max (Preview)Alibaba149.2
  7. #60Grok 4.3 BetaxAI149.2

Sources