in the loop

Qwen 3.8 Max
#22 of 274 scored.

AlibabaAug 2, 2026Closed

Capabilities Index

Range 154.3–158.6

156.4

Context

Tokens

—

Input

Per 1M tokens

—

Output

Per 1M tokens

—

Benchmarks

How it scores.

  • GPQA Diamond90.2%

    21st of 186

  • FrontierMath74.7%

    18th of 81

  • SimpleQA Verified45.8%

    36th of 77

  • Mock AIME99.4%

    13th of 176

Everything else Epoch has for it

  • DTBench86.7%
  • ProofBench58.0%
  • DeepSWE57.5%
  • LMCA54.4%
  • FrontierMath-Tier-4-v2-Private46.3%
  • Mystery Game Puzzles31.7%
  • Chess Puzzles25.3%
  • FrontierSWE15.8%

On the index

Around it.

  1. #19Muse Spark 1.3Meta · $1.25 in156.8
  2. #20Gemini 3.8 FlashGoogle · $0.75 in156.7
  3. #21Grok 4.6xAI · $2.00 in156.4
  4. #22Qwen 3.8 MaxAlibaba156.4
  5. #23GPT-5.6 LunaOpenAI · $0.20 in156.4
  6. #24GPT-6 LunaOpenAI · $0.10 in156.3
  7. #25Claude Opus 4.7Anthropic · $5.00 in156.3

Sources