in the loop

Llama 3-8B
#213 of 274 scored.

MetaApr 18, 2024Open weights

Capabilities Index

Range 106.8–119.3

116.5

Context

Tokens

—

Input

Per 1M tokens

—

Output

Per 1M tokens

—

Benchmarks

How it scores.

  • GPQA Diamond1.4%

    180th of 186

  • Mock AIME1.8%

    163rd of 176

Everything else Epoch has for it

  • ARC AI277.1%
  • OpenBookQA76.8%
  • TriviaQA67.7%
  • MMLU58.4%
  • Winogrande51.4%
  • ANLI35.9%
  • DTBench6.6%
  • MATH level 56.1%
  • Chess Puzzles0.0%

On the index

Around it.

  1. #210Stable Beluga 2Stability AI117.1
  2. #211Gemini 1.0 ProGoogle117.0
  3. #212Llama 3.1-8BMeta116.6
  4. #213Llama 3-8BMeta116.5
  5. #214Qwen2.5-Coder-14BAlibaba116.4
  6. #215Gemma 3 4BGoogle · $0.05 in116.0
  7. #216GPT-3.5 Turbo (Jan 2024)OpenAI115.7

Sources