in the loop

GPT-5.4
#17 of 274 scored.

OpenAIMar 5, 2026Closed

Capabilities Index

Range 155.0–158.9

156.8

Context

Takes text, image, file

1.05M

Input

Per 1M tokens

$2.50

Output

Per 1M tokens

$15

Benchmarks

How it scores.

  • GPQA Diamond91.1%

    17th of 186

  • SWE-bench Verified76.9%

    8th of 32

  • FrontierMath78.6%

    17th of 81

  • ARC-AGI-274.0%

    16th of 79

  • Humanity's Last Exam33.0%

    8th of 40

  • SimpleQA Verified45.1%

    38th of 77

  • Mock AIME97.8%

    24th of 176

  • Terminal-Bench81.8%

    2nd of 35

Everything else Epoch has for it

  • ARC-AGI93.7%
  • DTBench90.7%
  • WeirdML77.7%
  • METR Time Horizons74.3%
  • LMCA61.1%
  • ProofBench56.0%
  • APEX-Agents52.4%
  • DeepSWE51.8%
  • FrontierMath-Tier-4-v2-Private49.0%
  • Chess Puzzles41.1%
  • DeepResearch Bench35.1%
  • GSO-Bench31.4%
  • Mystery Game Puzzles30.6%
  • CL-bench27.9%
  • EBR-bench25.4%
  • CL-bench Life21.7%
  • PostTrainBench19.0%
  • MirrorCode15.6%
  • Furniture Assembly10.7%

On the index

Around it.

  1. #14Claude Opus 4.8Anthropic · $5.00 in158.2
  2. #15Kimi K3Moonshot · $0.80 in157.4
  3. #16Gemini 3.7 FlashGoogle · $0.75 in157.3
  4. #17GPT-5.4OpenAI · $2.50 in156.8
  5. #18GPT-5.3 CodexOpenAI · $1.75 in156.8
  6. #19Muse Spark 1.3Meta · $1.25 in156.8
  7. #20Gemini 3.8 FlashGoogle · $0.75 in156.7

Sources