in the loop

Gemini 1.5 Pro (Sept 2024)
#154 of 274 scored.

GoogleSep 24, 2024Closed

Capabilities Index

Range 127.2–133.1

131.7

Context

Tokens

—

Input

Per 1M tokens

—

Output

Per 1M tokens

—

Benchmarks

How it scores.

  • GPQA Diamond43.0%

    119th of 186

  • ARC-AGI-20.8%

    68th of 79

  • Humanity's Last Exam0.0%

    36th of 40

  • Mock AIME23.0%

    126th of 176

Everything else Epoch has for it

  • MMLU82.5%
  • MATH level 570.4%
  • CadEval34.0%
  • DTBench31.6%
  • WeirdML22.2%
  • Balrog21.0%
  • SimpleBench12.5%
  • The Agent Company3.4%

On the index

Around it.

  1. #151DeepSeek-V3DeepSeek · $0.32 in132.3
  2. #152Llama 4 MaverickMeta · $0.19 in132.2
  3. #153Mistral Small 3.2Mistral131.7
  4. #154Gemini 1.5 Pro (Sept 2024)Google131.7
  5. #155Magistral Small 1.2Mistral131.4
  6. #156Grok-2 (Dec 2024)xAI130.5
  7. #157Phi-4Microsoft · $0.07 in130.4

Sources