in the loop

Qwen3-235B-A22B-Thinking (Jul 2025)
#95 of 274 scored.

AlibabaJul 25, 2025Open weights

Capabilities Index

Range 141.9–145.3

143.8

Context

Tokens

—

Input

Per 1M tokens

—

Output

Per 1M tokens

—

Benchmarks

How it scores.

  • GPQA Diamond73.4%

    80th of 186

  • SimpleQA Verified40.4%

    45th of 77

  • Mock AIME86.7%

    61st of 176

Everything else Epoch has for it

  • Lech Mazur Writing82.4%
  • Fiction.LiveBench75.0%
  • DTBench67.1%
  • WeirdML41.0%
  • LMCA34.5%
  • FrontierMath-2025-02-28-Private14.9%
  • Chess Puzzles7.4%
  • FrontierMath-Tier-4-2025-07-01-Private0.0%
  • Mystery Game Puzzles0.0%

On the index

Around it.

  1. #92Claude Opus 4.1Anthropic · $15 in144.1
  2. #93Qwen 3.5 Flash (hosted 35B-A3B)Alibaba144.0
  3. #94Qwen 3.6 35B-A3BAlibaba · $0.15 in143.9
  4. #95Qwen3-235B-A22B-Thinking (Jul 2025)Alibaba143.8
  5. #96GLM-4.7Z.ai · $0.60 in143.5
  6. #97Qwen 3.6 FlashAlibaba · $0.19 in143.3
  7. #98Gemini 2.5 Flash (Sep 2025)Google143.0

Sources