in the loop

Grok-2 (Dec 2024)
#156 of 274 scored.

xAIDec 12, 2024Closed

Capabilities Index

Range 125.8–132.2

130.5

Context

Tokens

—

Input

Per 1M tokens

—

Output

Per 1M tokens

—

Benchmarks

How it scores.

  • GPQA Diamond38.4%

    129th of 186

  • Mock AIME11.4%

    134th of 176

Everything else Epoch has for it

  • Lech Mazur Writing63.6%
  • MATH level 563.5%
  • DTBench42.0%
  • WeirdML22.2%
  • SimpleBench7.2%
  • FrontierMath-2025-02-28-Private1.2%

On the index

Around it.

  1. #153Mistral Small 3.2Mistral131.7
  2. #154Gemini 1.5 Pro (Sept 2024)Google131.7
  3. #155Magistral Small 1.2Mistral131.4
  4. #156Grok-2 (Dec 2024)xAI130.5
  5. #157Phi-4Microsoft · $0.07 in130.4
  6. #158Gemma 3 27BGoogle · $0.08 in130.0
  7. #159Claude 3.5 SonnetAnthropic130.0

Sources