in the loop

gpt-oss-120b
#118 of 274 scored.

OpenAIAug 5, 2025Open weights

Capabilities Index

Range 135.2–142.8

139.9

Context

Takes text

131K

Input

Per 1M tokens

$0.037

Output

Per 1M tokens

$0.17

Benchmarks

How it scores.

  • GPQA Diamond67.7%

    92nd of 186

  • Mock AIME88.9%

    54th of 176

  • Terminal-Bench18.7%

    31st of 35

Everything else Epoch has for it

  • Lech Mazur Writing77.1%
  • DTBench60.5%
  • METR Time Horizons56.6%
  • WeirdML48.2%
  • Fiction.LiveBench44.4%
  • Aider polyglot41.8%
  • LMCA26.1%
  • Surface Evolver Bench25.0%
  • Chess Puzzles15.8%
  • SimpleBench6.5%
  • APEX-Agents4.4%
  • Mystery Game Puzzles0.0%

On the index

Around it.

  1. #115o3-miniOpenAI · $1.10 in140.3
  2. #116Kimi K2 (Jul 2025)Moonshot140.1
  3. #117Gemini 2.5 Flash (Apr 2025)Google140.0
  4. #118gpt-oss-120bOpenAI · $0.037 in139.9
  5. #119DeepSeek-V3.1DeepSeek · $0.25 in139.9
  6. #120Qwen3-30B-A3B-Thinking (Jul 2025)Alibaba139.6
  7. #121Qwen3.5-9BAlibaba · $0.10 in139.5

Sources