in the loop

o1-mini
#138 of 274 scored.

OpenAISep 12, 2024Closed

Capabilities Index

Range 131.7–137.3

135.8

Context

Tokens

—

Input

Per 1M tokens

—

Output

Per 1M tokens

—

Benchmarks

How it scores.

  • GPQA Diamond49.8%

    115th of 186

  • ARC-AGI-20.8%

    67th of 79

  • Mock AIME46.9%

    114th of 176

Everything else Epoch has for it

  • MATH level 589.2%
  • Lech Mazur Writing64.9%
  • WeirdML36.3%
  • Aider polyglot32.9%
  • ARC-AGI14.0%
  • Cybench10.0%
  • FrontierMath-2025-02-28-Private3.0%
  • SimpleBench1.7%

On the index

Around it.

  1. #135Qwen3-30B-A3BAlibaba · $0.12 in136.2
  2. #136Qwen3-8BAlibaba136.2
  3. #137DeepSeek-V3 (Mar 2025)DeepSeek135.9
  4. #138o1-miniOpenAI135.8
  5. #139DeepSeek-R1-Distill-Qwen-14BDeepSeek135.4
  6. #140Gemini 2.0 Flash Thinking (Jan 2025)Google135.4
  7. #141Gemini 2.0 ProGoogle135.1

Sources