in the loop

Gemini 3.1 Pro
#34 of 274 scored.

GoogleFeb 19, 2026Closed

Capabilities Index

Range 152.4–157.3

154.8

Context

Tokens

—

Input

Per 1M tokens

—

Output

Per 1M tokens

—

Benchmarks

How it scores.

  • GPQA Diamond92.6%

    7th of 186

  • SWE-bench Verified75.6%

    12th of 32

  • FrontierMath59.7%

    34th of 81

  • ARC-AGI-277.1%

    14th of 79

  • Humanity's Last Exam43.7%

    3rd of 40

  • SimpleQA Verified73.5%

    3rd of 77

  • Mock AIME95.6%

    31st of 176

  • Terminal-Bench80.2%

    3rd of 35

Everything else Epoch has for it

  • ARC-AGI98.0%
  • DTBench95.1%
  • METR Time Horizons77.0%
  • SimpleBench75.5%
  • WeirdML72.1%
  • LMCA63.3%
  • Balrog57.0%
  • Chess Puzzles52.6%
  • DeepResearch Bench47.8%
  • APEX-Agents35.3%
  • Mystery Game Puzzles27.3%
  • FrontierMath-Tier-4-v2-Private26.8%
  • ExploitBench26.1%
  • ProofBench26.0%
  • GSO-Bench22.6%
  • PostTrainBench22.0%
  • CL-bench20.8%
  • CL-bench Life16.9%
  • EBR-bench14.3%
  • DeepSWE11.7%
  • MirrorCode8.9%
  • Furniture Assembly0.0%

On the index

Around it.

  1. #31Qwen3.8 Max (0902)Alibaba · $2.00 in155.1
  2. #32DeepSeek V4.1 FlashDeepSeek · $0.30 in154.9
  3. #33Muse Spark 1.2Meta · $1.25 in154.9
  4. #34Gemini 3.1 ProGoogle154.8
  5. #35DeepSeek V4 Flash 0731DeepSeek · $0.006 in154.5
  6. #36Gemini 3.5 FlashGoogle · $1.50 in154.5
  7. #37Gemini 3.6 FlashGoogle · $0.75 in154.3

Sources