in the loop

Gemini 2.5 Pro (Jun 2025)
#85 of 274 scored.

GoogleJun 5, 2025Closed

Capabilities Index

Range 143.8–146.8

145.3

Context

Tokens

—

Input

Per 1M tokens

—

Output

Per 1M tokens

—

Benchmarks

How it scores.

  • GPQA Diamond80.4%

    64th of 186

  • SWE-bench Verified57.6%

    30th of 32

  • FrontierMath24.6%

    65th of 81

  • ARC-AGI-24.9%

    52nd of 79

  • Humanity's Last Exam17.7%

    17th of 40

  • Mock AIME84.7%

    70th of 176

  • Terminal-Bench32.6%

    26th of 35

Everything else Epoch has for it

  • Fiction.LiveBench91.7%
  • Lech Mazur Writing83.8%
  • Aider polyglot83.1%
  • DTBench70.7%
  • METR Time Horizons55.4%
  • SimpleBench54.9%
  • WeirdML54.0%
  • DeepResearch Bench42.8%
  • ARC-AGI41.0%
  • LMCA40.9%
  • GDPval23.3%
  • VPCT19.6%
  • Chess Puzzles15.8%
  • GSO-Bench3.9%
  • Remote Labor Index0.8%
  • FrontierMath-Tier-4-v2-Private0.0%

On the index

Around it.

  1. #82GPT-5.4 NanoOpenAI · $0.20 in145.8
  2. #83o4-miniOpenAI · $1.10 in145.6
  3. #84GPT-5 miniOpenAI · $0.25 in145.5
  4. #85Gemini 2.5 Pro (Jun 2025)Google145.3
  5. #86Gemini 3.5 Flash-LiteGoogle · $0.30 in145.1
  6. #87DeepSeek-V3.2-ExpDeepSeek · $0.27 in145.0
  7. #88Qwen3.7 FlashAlibaba · $0.03 in144.6

Sources