in the loop

Llama 3.1-8B
#212 of 274 scored.

MetaJul 23, 2024Open weights

Capabilities Index

Range 106.2–121.1

116.6

Context

Tokens

—

Input

Per 1M tokens

—

Output

Per 1M tokens

—

Benchmarks

How it scores.

  • GPQA Diamond2.6%

    178th of 186

  • Mock AIME1.6%

    167th of 176

Everything else Epoch has for it

  • GSM8K82.4%
  • PIQA62.4%
  • MMLU41.5%
  • MATH level 522.9%
  • DTBench18.2%
  • Balrog15.1%
  • LMCA6.3%
  • WeirdML1.7%
  • Chess Puzzles0.0%

On the index

Around it.

  1. #209phi-3-mini 3.8BMicrosoft117.4
  2. #210Stable Beluga 2Stability AI117.1
  3. #211Gemini 1.0 ProGoogle117.0
  4. #212Llama 3.1-8BMeta116.6
  5. #213Llama 3-8BMeta116.5
  6. #214Qwen2.5-Coder-14BAlibaba116.4
  7. #215Gemma 3 4BGoogle · $0.05 in116.0

Sources