Skip to content
All systems operational0 AI providers monitored, polled every 2 minutes
Live status
All benchmarks

GPQA Diamond leaderboard

GPQA Diamond is the hardest subset of the Graduate-level Physics and Quantum questions benchmark. The 198 questions in the Diamond subset have been verified by domain experts to be difficult even for PhDs in the relevant field. Scoring well on GPQA Diamond requires multi-step scientific reasoning, not just memorization.

Current leader
GPT-6 Astra(OpenAI)96%

Last refreshed 2026-09-23. 54 models scored on this benchmark.

Full leaderboard

#ModelProviderScoreReleased
1GPT-6 AstraOpenAI96%vendor2026-09
2Fugu Ultra v2.0Sakana AI95.5%vendor2026-09
3Fugu MaxSakana AI95.5%vendor2026-09
4Gemini 3.8 FlashGoogle95.3%independent2026-09
5Grok 4.6xAI94.9%independent2026-08
6GPT-5.6 SolOpenAI94.6%vendor2026-07
7Gemini 3.7 FlashGoogle94.5%independent2026-08
8Muse Spark 1.3Meta94.1%independent2026-09
9Claude Opus 4.8Anthropic93.6%2026-05
10Kimi K3Moonshot AI93.5%vendor2026-07
11Gemini 3.6 FlashGoogle93.43%independent2026-07
12Claude Opus 5Anthropic93.43%independent2026-07
13Claude Fable 5.1Anthropic93.43%independent2026-09
14Claude Fable 5Anthropic93.18%independent2026-06
15Grok 4.5xAI92.93%independent2026-07
16GPT-5.6 TerraOpenAI92.9%vendor2026-07
17MiniMax M3MiniMax92.68%independent2026-06
18Gemini 3.5 FlashGoogle92.68%independent2026-05
19Qwen3.8 2.4T-A95BAlibaba92.6%vendor2026-08
20Qwen3.8-MaxAlibaba92.6%vendor2026-08
21DeepSeek V4 ProDeepSeek92.4%vendor2026-04
22GPT-5.6 LunaOpenAI92.3%vendor2026-07
23Hy4 previewTencent92.3%vendor2026-08
24Qwen3.8-FlashAlibaba91.7%vendor2026-08
25GLM-5.2Z.ai91.2%vendor2026-06
26Muse Spark 1.1Meta91.16%independent2026-07
27DeepSeek V4.1 FlashDeepSeek90.9%vendor2026-09
28Qwen3.8 27BAlibaba89.2%vendor2026-08
29LongCat-2.0Meituan88.9%vendor2026-06
30Claude Sonnet 5Anthropic88.89%independent2026-06
31GLM-5.3Z.ai88.13%independent2026-08
32InklingThinking Machines87.2%vendor2026-07
33Gemini 3.1 Flash-LiteGoogle86.9%vendor2026-03
34GLM-5.3 FlashZ.ai86.36%independent2026-08
35Gemini 3.5 Flash-LiteGoogle83.84%independent2026-07
36Muse Glimmer 30BMeta83.5%vendor2026-08
37GPT-5.5OpenAI78.3%2026-04
38Claude Opus 4.7Anthropic76.5%2026-04
39Nemotron 3.5 LightningNVIDIA75.44%vendor2026-08
40Claude Opus 4.6Anthropic74.2%2026-03
41o1OpenAI72.5%2024-12
42Gemini 2.5 ProGoogle71.9%2026-01
43GPT-4.5OpenAI68.7%2025-12
44Claude Sonnet 4.6Anthropic65.8%2026-02
45Llama 4 MaverickMeta64.1%2025-04
46DeepSeek V3DeepSeek63.5%2025-12
47o3-miniOpenAI60.3%2025-01
48GPT-4oOpenAI59.1%2024-05
49DeepSeek V4 FlashDeepSeek58.7%2026-04
50Mistral LargeMistral57.3%2025-11
51Llama 4 ScoutMeta56.2%2025-04
52Gemini 2.0 FlashGoogle54.8%2025-02
53Claude Haiku 4.5Anthropic52.4%2026-01
54Mistral SmallMistral44.6%2025-01

Score interpretation

The chance baseline is 25% (4-choice). Random guessing scores ~25%, expert non-specialists score ~34%, expert specialists score ~65%. The 2026 frontier hits ~80% on the strongest reasoning models. The gap between this and MMLU-Pro is the gap between "knows facts" and "can reason from facts under pressure."

70%+
Above expert-specialist level. Frontier reasoning.
50-70%
Strong scientific reasoning, comparable to expert non-specialists.
30-50%
Better than chance, weak on multi-step inference.
< 30%
At or near chance baseline for the benchmark.

Why this matters for AI agents

GPQA Diamond is the benchmark that separates models that have memorized scientific content from models that can reason scientifically. For research agents, technical writing, and any workload involving multi-step inference over unfamiliar domains, GPQA Diamond predicts capability better than MMLU-Pro.

Other benchmarks

SWE-bench
Live SWE-bench leaderboard for major AI models. Real-world software engineering tasks from GitHub issues. Pricing per model and rankings updated weekly on TensorFeed.
MMLU-Pro
Live MMLU-Pro leaderboard for major AI models. General knowledge and reasoning across 57 subjects. Updated weekly with pricing per model on TensorFeed.
HumanEval
Live HumanEval leaderboard for major AI models. Python code generation and problem solving. Updated weekly with pricing on TensorFeed.
MATH
Live MATH benchmark leaderboard for major AI models. Competition-level mathematics problems. Updated weekly with pricing per model on TensorFeed.
OSWorld 2.0
Live OSWorld 2.0 leaderboard ranking AI models on long-horizon computer use across real desktop applications. Binary and partial scoring explained, updated on TensorFeed.
BrowseComp
Live BrowseComp leaderboard for AI agentic web search. 1,266 hard-to-find questions, why every score is self-reported, and how context strategy moves results. On TensorFeed.
FrontierCode v1.1
Live FrontierCode v1.1 leaderboard. Scores whether an AI patch is mergeable, not just test-passing, graded by open-source maintainers against a rubric. On TensorFeed.
Terminal-Bench 4.0
Live Terminal-Bench 4.0 leaderboard on one independent harness. 66 long-horizon terminal tasks across software, science, ML, and security, scored all or nothing.
Humanity's Last Exam (tools)
Live Humanity's Last Exam with-tools leaderboard on TensorFeed. Why the tool-augmented mode is not the same benchmark as closed-book HLE, and how far apart the two run.

Premium API: time-series for GPQA Diamond

The leaderboard above is a snapshot. Want to see how a model's GPQA Diamond score has moved over the last 30-90 days, or set a webhook that fires when a score crosses a threshold? The premium API has both:

GPQA Diamond source ·Last refreshed 2026-09-23·Max score 100