in the loop

Qwen2.5-7B
#204 of 274 scored.

AlibabaSep 19, 2024Open weights

Capabilities Index

Range 110.9–121.2

118.5

Context

Tokens

—

Input

Per 1M tokens

—

Output

Per 1M tokens

—

Benchmarks

How it scores.

  • GPQA Diamond14.0%

    166th of 186

  • Mock AIME2.4%

    159th of 176

Everything else Epoch has for it

  • MMLU63.9%
  • DTBench12.9%
  • Balrog7.8%
  • LMCA7.5%
  • Chess Puzzles0.0%

On the index

Around it.

  1. #201Claude 2.1Anthropic119.3
  2. #202Mistral NeMoMistral · $0.029 in118.7
  3. #203GPT-3.5 Turbo (Nov 2023)OpenAI118.5
  4. #204Qwen2.5-7BAlibaba118.5
  5. #205Mixtral 8x7BMistral118.5
  6. #206Claude 3 HaikuAnthropic118.3
  7. #207Ministral 3BMistral118.1

Sources