in the loop

Llama 3.1-405B
#167 of 274 scored.

MetaJul 23, 2024Open weights

Capabilities Index

Range 123.8–130.8

128.8

Context

Tokens

—

Input

Per 1M tokens

—

Output

Per 1M tokens

—

Benchmarks

How it scores.

  • GPQA Diamond34.5%

    132nd of 186

  • Mock AIME9.6%

    135th of 176

Everything else Epoch has for it

  • ARC AI293.7%
  • HellaSwag85.6%
  • TriviaQA82.7%
  • MMLU79.3%
  • Winogrande78.4%
  • BBH77.2%
  • PIQA71.8%
  • MATH level 549.8%
  • DTBench35.6%
  • WeirdML21.4%
  • SimpleBench7.6%
  • Cybench7.5%
  • The Agent Company7.4%

On the index

Around it.

  1. #164GPT-4o (May 2024)OpenAI129.0
  2. #165GPT-4o (Nov 2024)OpenAI128.8
  3. #166GPT-4o (Aug 2024)OpenAI128.8
  4. #167Llama 3.1-405BMeta128.8
  5. #168Mistral Large 2 (Nov 2024)Mistral128.5
  6. #169Qwen2.5-32BAlibaba128.5
  7. #170Mistral Large 2 (Jul 2024)Mistral127.5

Sources