MMLU-Pro leaderboard
MMLU-Pro is the harder successor to the original MMLU benchmark. It tests general knowledge and reasoning across 57 subjects (math, physics, law, medicine, philosophy, etc.) using multiple-choice questions designed to require multi-step reasoning rather than memorization. MMLU-Pro is the standard "is this model smart" benchmark for general-purpose use cases.
Full leaderboard
| # | Model | Provider | Score | Released |
|---|---|---|---|---|
| 1 | GPT-5.5 | OpenAI | 94.2% | 2026-04 |
| 2 | Claude Opus 4.7 | Anthropic | 93.8% | 2026-04 |
| 3 | Claude Opus 4.6 | Anthropic | 92.4% | 2026-03 |
| 4 | Claude Fable 5.1 | Anthropic | 92.38%independent | 2026-09 |
| 5 | o1 | OpenAI | 91.8% | 2024-12 |
| 6 | Claude Opus 5 | Anthropic | 91.59%independent | 2026-07 |
| 7 | DeepSeek V4 Pro | DeepSeek | 91.5% | 2026-04 |
| 8 | Claude Fable 5 | Anthropic | 91.5%independent | 2026-06 |
| 9 | Gemini 2.5 Pro | 91.2% | 2026-01 | |
| 10 | Gemini 3.8 Flash | 90.22%independent | 2026-09 | |
| 11 | Gemini 3.7 Flash | 90.12%independent | 2026-08 | |
| 12 | GPT-4.5 | OpenAI | 90.1% | 2025-12 |
| 13 | Gemini 3.5 Flash | 89.52%independent | 2026-05 | |
| 14 | Grok 4.6 | xAI | 89.4%independent | 2026-08 |
| 15 | Llama 4 Maverick | Meta | 89.3% | 2025-04 |
| 16 | Gemini 3.6 Flash | 89.28%independent | 2026-07 | |
| 17 | Grok 4.5 | xAI | 89.22%independent | 2026-07 |
| 18 | GPT-5.6 Sol | OpenAI | 89.1%independent | 2026-07 |
| 19 | Muse Spark 1.1 | Meta | 88.73%independent | 2026-07 |
| 20 | Claude Sonnet 4.6 | Anthropic | 88.7% | 2026-02 |
| 21 | Qwen3.8-Max | Alibaba | 88.6%independent | 2026-08 |
| 22 | Muse Spark 1.2 | Meta | 88.28%independent | 2026-08 |
| 23 | DeepSeek V3 | DeepSeek | 88.1% | 2025-12 |
| 24 | Kimi K3 | Moonshot AI | 87.97%independent | 2026-07 |
| 25 | Claude Sonnet 5 | Anthropic | 87.55%independent | 2026-06 |
| 26 | GPT-4o | OpenAI | 87.2% | 2024-05 |
| 27 | Mistral Large | Mistral | 86.8% | 2025-11 |
| 28 | GLM-5.3 | Z.ai | 86.77%independent | 2026-08 |
| 29 | GLM-5.2 | Z.ai | 86.71%independent | 2026-06 |
| 30 | GPT-5.6 Terra | OpenAI | 86.66%independent | 2026-07 |
| 31 | Inkling | Thinking Machines | 86.3%independent | 2026-07 |
| 32 | o3-mini | OpenAI | 86.3% | 2025-01 |
| 33 | GLM-5.3 Flash | Z.ai | 86.06%independent | 2026-08 |
| 34 | GPT-5.6 Luna | OpenAI | 86.04%independent | 2026-07 |
| 35 | Llama 4 Scout | Meta | 85.9% | 2025-04 |
| 36 | Gemini 3.5 Flash-Lite | 85.84%independent | 2026-07 | |
| 37 | DeepSeek V4 Flash | DeepSeek | 85.2% | 2026-04 |
| 38 | Gemini 2.0 Flash | 84.5% | 2025-02 | |
| 39 | Qwen3.8 27B | Alibaba | 84.34%independent | 2026-08 |
| 40 | MiniMax M3 | MiniMax | 84.22%independent | 2026-06 |
| 41 | Claude Haiku 4.5 | Anthropic | 82.1% | 2026-01 |
| 42 | Nemotron 3.5 Lightning | NVIDIA | 81.94%vendor | 2026-08 |
| 43 | Mistral Small | Mistral | 78.4% | 2025-01 |
Score interpretation
Scores are reported as % of questions answered correctly. The chance baseline is roughly 25% (4-choice). The 2026 frontier sits above 90%, with the strongest models in the mid-90s. A 5-point gap on MMLU-Pro is meaningful; a 1-point gap is within noise.
Why this matters for AI agents
For general chat assistants, research synthesis, and any workload where the model needs broad knowledge plus reasoning, MMLU-Pro is the best single proxy for capability. Models that lead MMLU-Pro almost always lead other reasoning benchmarks too.
Other benchmarks
Premium API: time-series for MMLU-Pro
The leaderboard above is a snapshot. Want to see how a model's MMLU-Pro score has moved over the last 30-90 days, or set a webhook that fires when a score crosses a threshold? The premium API has both:
/api/premium/history/benchmarks/series?model=&benchmark=mmlu_pro: daily score evolution for one model on this benchmark, 1 credit per call