in the loop

GPT-4.1
#133 of 274 scored.

OpenAIApr 14, 2025Closed

Capabilities Index

Range 133.6–138.6

136.8

Context

Takes image, text, file

1.05M

Input

Per 1M tokens

$2.00

Output

Per 1M tokens

$8.00

Benchmarks

How it scores.

  • GPQA Diamond55.9%

    106th of 186

  • SWE-bench Verified48.5%

    31st of 32

  • FrontierMath6.0%

    77th of 81

  • ARC-AGI-20.4%

    70th of 79

  • Humanity's Last Exam0.6%

    35th of 40

  • SimpleQA Verified31.1%

    58th of 77

  • Mock AIME38.3%

    116th of 176

Everything else Epoch has for it

  • MATH level 583.0%
  • GeoBench72.0%
  • Fiction.LiveBench63.9%
  • Aider polyglot52.4%
  • DTBench47.1%
  • CadEval42.0%
  • WeirdML39.0%
  • LMCA30.1%
  • SimpleBench12.4%
  • ARC-AGI5.5%
  • Chess Puzzles1.1%
  • FrontierMath-Tier-4-2025-07-01-Private0.0%

On the index

Around it.

  1. #130QwQ-32BAlibaba137.6
  2. #131DeepSeek-R1-Distill-Qwen-32BDeepSeek137.4
  3. #132Qwen3-30B-A3B-Instruct (Jul 2025)Alibaba137.4
  4. #133GPT-4.1OpenAI · $2.00 in136.8
  5. #134GPT-4.5OpenAI136.7
  6. #135Qwen3-30B-A3BAlibaba · $0.12 in136.2
  7. #136Qwen3-8BAlibaba136.2

Sources