in the loop

Qwen2-72B
#183 of 274 scored.

AlibabaJun 7, 2024Open weights

Capabilities Index

Range 119.0–126.9

125.3

Context

Tokens

—

Input

Per 1M tokens

—

Output

Per 1M tokens

—

Benchmarks

How it scores.

  • GPQA Diamond21.0%

    155th of 186

Everything else Epoch has for it

  • MMLU76.5%
  • MATH level 539.1%
  • METR Time Horizons29.9%
  • WeirdML11.3%
  • The Agent Company1.1%

On the index

Around it.

  1. #180Llama 3.1-70BMeta125.9
  2. #181GPT-4 (Mar 2023)OpenAI125.9
  3. #182Llama 3.2 90BMeta125.5
  4. #183Qwen2-72BAlibaba125.3
  5. #184DeepSeek-V2 (MoE-236B, May 2024)DeepSeek124.8
  6. #185Amazon Nova ProAmazon123.8
  7. #186Gemma 3 12BGoogle · $0.05 in123.5

Sources