in the loop

Claude 2.1
#201 of 274 scored.

AnthropicNov 21, 2023Closed

Capabilities Index

Range 111.3–122.0

119.3

Context

Tokens

—

Input

Per 1M tokens

—

Output

Per 1M tokens

—

Benchmarks

How it scores.

  • GPQA Diamond10.6%

    170th of 186

  • Mock AIME1.8%

    163rd of 176

Everything else Epoch has for it

  • MMLU64.7%
  • DTBench18.3%
  • WeirdML7.1%

On the index

Around it.

  1. #198Gemma 2 9BGoogle119.8
  2. #199Qwen2.5-Coder-32BAlibaba119.5
  3. #200Command R+Cohere119.3
  4. #201Claude 2.1Anthropic119.3
  5. #202Mistral NeMoMistral · $0.029 in118.7
  6. #203GPT-3.5 Turbo (Nov 2023)OpenAI118.5
  7. #204Qwen2.5-7BAlibaba118.5

Sources