in the loop

Gemini 2.0 Flash Thinking (Jan 2025)
#140 of 274 scored.

GoogleJan 21, 2025Closed

Capabilities Index

Range 130.3–138.1

135.4

Context

Tokens

—

Input

Per 1M tokens

—

Output

Per 1M tokens

—

Benchmarks

How it scores.

  • GPQA Diamond42.8%

    120th of 186

  • Humanity's Last Exam1.8%

    32nd of 40

  • Mock AIME57.7%

    107th of 176

Everything else Epoch has for it

  • Lech Mazur Writing73.8%
  • Fiction.LiveBench52.8%
  • Aider polyglot18.2%
  • SimpleBench16.8%

On the index

Around it.

  1. #137DeepSeek-V3 (Mar 2025)DeepSeek135.9
  2. #138o1-miniOpenAI135.8
  3. #139DeepSeek-R1-Distill-Qwen-14BDeepSeek135.4
  4. #140Gemini 2.0 Flash Thinking (Jan 2025)Google135.4
  5. #141Gemini 2.0 ProGoogle135.1
  6. #142GPT-4.1 miniOpenAI · $0.40 in135.0
  7. #143o1-previewOpenAI134.8

Sources