Grok 4
#75 of 274 scored.
xAIJul 9, 2025Closed
Capabilities Index
Range 144.7–148.1
146.4
Context
Tokens
—
Input
Per 1M tokens
—
Output
Per 1M tokens
—
Benchmarks
How it scores.
- GPQA Diamond82.7%
56th of 186
- ARC-AGI-216.0%
41st of 79
- Mock AIME84.0%
74th of 176
- Terminal-Bench27.2%
28th of 35
Everything else Epoch has for it
- Fiction.LiveBench94.4%
- Aider polyglot79.6%
- Lech Mazur Writing76.9%
- ARC-AGI66.7%
- METR Time Horizons66.6%
- SimpleBench52.6%
- DeepResearch Bench47.3%
- WeirdML45.7%
- GeoBench45.0%
- Balrog43.6%
- Cybench43.0%
- FrontierMath-2025-02-28-Private34.5%
- Chess Puzzles24.2%
- GDPval21.1%
- FrontierMath-Tier-4-2025-07-01-Private3.5%
On the index
Around it.
Sources