in the loop

Phi-4
#157 of 274 scored.

MicrosoftDec 12, 2024Open weights

Capabilities Index

Range 124.7–132.0

130.4

Context

Takes text

16K

Input

Per 1M tokens

$0.07

Output

Per 1M tokens

$0.14

Benchmarks

How it scores.

  • GPQA Diamond41.4%

    124th of 186

  • Mock AIME13.7%

    133rd of 176

Everything else Epoch has for it

  • MMLU79.7%
  • MATH level 564.9%
  • Lech Mazur Writing62.6%
  • Balrog11.6%
  • Chess Puzzles0.0%

On the index

Around it.

  1. #154Gemini 1.5 Pro (Sept 2024)Google131.7
  2. #155Magistral Small 1.2Mistral131.4
  3. #156Grok-2 (Dec 2024)xAI130.5
  4. #157Phi-4Microsoft · $0.07 in130.4
  5. #158Gemma 3 27BGoogle · $0.08 in130.0
  6. #159Claude 3.5 SonnetAnthropic130.0
  7. #160Llama 4 ScoutMeta · $0.10 in129.6

Sources