Skip to content
All systems operational0 AI providers monitored, polled every 2 minutes
Live status
All benchmarks

MMLU-Pro leaderboard

MMLU-Pro is the harder successor to the original MMLU benchmark. It tests general knowledge and reasoning across 57 subjects (math, physics, law, medicine, philosophy, etc.) using multiple-choice questions designed to require multi-step reasoning rather than memorization. MMLU-Pro is the standard "is this model smart" benchmark for general-purpose use cases.

Current leader
GPT-5.5(OpenAI)94.2%

Last refreshed 2026-09-23. 43 models scored on this benchmark.

Full leaderboard

#ModelProviderScoreReleased
1GPT-5.5OpenAI94.2%2026-04
2Claude Opus 4.7Anthropic93.8%2026-04
3Claude Opus 4.6Anthropic92.4%2026-03
4Claude Fable 5.1Anthropic92.38%independent2026-09
5o1OpenAI91.8%2024-12
6Claude Opus 5Anthropic91.59%independent2026-07
7DeepSeek V4 ProDeepSeek91.5%2026-04
8Claude Fable 5Anthropic91.5%independent2026-06
9Gemini 2.5 ProGoogle91.2%2026-01
10Gemini 3.8 FlashGoogle90.22%independent2026-09
11Gemini 3.7 FlashGoogle90.12%independent2026-08
12GPT-4.5OpenAI90.1%2025-12
13Gemini 3.5 FlashGoogle89.52%independent2026-05
14Grok 4.6xAI89.4%independent2026-08
15Llama 4 MaverickMeta89.3%2025-04
16Gemini 3.6 FlashGoogle89.28%independent2026-07
17Grok 4.5xAI89.22%independent2026-07
18GPT-5.6 SolOpenAI89.1%independent2026-07
19Muse Spark 1.1Meta88.73%independent2026-07
20Claude Sonnet 4.6Anthropic88.7%2026-02
21Qwen3.8-MaxAlibaba88.6%independent2026-08
22Muse Spark 1.2Meta88.28%independent2026-08
23DeepSeek V3DeepSeek88.1%2025-12
24Kimi K3Moonshot AI87.97%independent2026-07
25Claude Sonnet 5Anthropic87.55%independent2026-06
26GPT-4oOpenAI87.2%2024-05
27Mistral LargeMistral86.8%2025-11
28GLM-5.3Z.ai86.77%independent2026-08
29GLM-5.2Z.ai86.71%independent2026-06
30GPT-5.6 TerraOpenAI86.66%independent2026-07
31InklingThinking Machines86.3%independent2026-07
32o3-miniOpenAI86.3%2025-01
33GLM-5.3 FlashZ.ai86.06%independent2026-08
34GPT-5.6 LunaOpenAI86.04%independent2026-07
35Llama 4 ScoutMeta85.9%2025-04
36Gemini 3.5 Flash-LiteGoogle85.84%independent2026-07
37DeepSeek V4 FlashDeepSeek85.2%2026-04
38Gemini 2.0 FlashGoogle84.5%2025-02
39Qwen3.8 27BAlibaba84.34%independent2026-08
40MiniMax M3MiniMax84.22%independent2026-06
41Claude Haiku 4.5Anthropic82.1%2026-01
42Nemotron 3.5 LightningNVIDIA81.94%vendor2026-08
43Mistral SmallMistral78.4%2025-01

Score interpretation

Scores are reported as % of questions answered correctly. The chance baseline is roughly 25% (4-choice). The 2026 frontier sits above 90%, with the strongest models in the mid-90s. A 5-point gap on MMLU-Pro is meaningful; a 1-point gap is within noise.

90%+
Frontier reasoning. Comparable to PhD-level human performance.
80-90%
Strong general assistant. Production-ready for most knowledge tasks.
60-80%
Useful for everyday queries, weak on harder reasoning.
< 60%
Below the threshold for reliable knowledge work.

Why this matters for AI agents

For general chat assistants, research synthesis, and any workload where the model needs broad knowledge plus reasoning, MMLU-Pro is the best single proxy for capability. Models that lead MMLU-Pro almost always lead other reasoning benchmarks too.

Other benchmarks

SWE-bench
Live SWE-bench leaderboard for major AI models. Real-world software engineering tasks from GitHub issues. Pricing per model and rankings updated weekly on TensorFeed.
HumanEval
Live HumanEval leaderboard for major AI models. Python code generation and problem solving. Updated weekly with pricing on TensorFeed.
GPQA Diamond
Live GPQA Diamond leaderboard for major AI models. Graduate-level physics, chemistry, and biology questions. Updated on TensorFeed with pricing per model.
MATH
Live MATH benchmark leaderboard for major AI models. Competition-level mathematics problems. Updated weekly with pricing per model on TensorFeed.
OSWorld 2.0
Live OSWorld 2.0 leaderboard ranking AI models on long-horizon computer use across real desktop applications. Binary and partial scoring explained, updated on TensorFeed.
BrowseComp
Live BrowseComp leaderboard for AI agentic web search. 1,266 hard-to-find questions, why every score is self-reported, and how context strategy moves results. On TensorFeed.
FrontierCode v1.1
Live FrontierCode v1.1 leaderboard. Scores whether an AI patch is mergeable, not just test-passing, graded by open-source maintainers against a rubric. On TensorFeed.
Terminal-Bench 4.0
Live Terminal-Bench 4.0 leaderboard on one independent harness. 66 long-horizon terminal tasks across software, science, ML, and security, scored all or nothing.
Humanity's Last Exam (tools)
Live Humanity's Last Exam with-tools leaderboard on TensorFeed. Why the tool-augmented mode is not the same benchmark as closed-book HLE, and how far apart the two run.

Premium API: time-series for MMLU-Pro

The leaderboard above is a snapshot. Want to see how a model's MMLU-Pro score has moved over the last 30-90 days, or set a webhook that fires when a score crosses a threshold? The premium API has both:

MMLU-Pro source ·Last refreshed 2026-09-23·Max score 100