Skip to content
All systems operational0 AI providers monitored, polled every 2 minutes
Live status
All benchmarks

Humanity's Last Exam (tools) leaderboard

Humanity's Last Exam is a multidisciplinary set of expert-level questions built to resist the saturation that overtook MMLU. This column reports the tool-augmented mode, where the model may search and run tools while answering. That distinction is the whole story: the official maintainer leaderboards are closed-book by construction, evaluated text-only at temperature zero with searchable questions deliberately removed from the set so retrieval cannot substitute for knowledge. There is no separate with-tools dataset, paper, repository, or maintainer. It is a reporting mode, not a distinct benchmark, and the figures come from the labs themselves.

Current leader
Claude Opus 5.5(Anthropic)67.7%

Last refreshed 2026-09-23. 15 models scored on this benchmark.

Full leaderboard

#ModelProviderScoreReleased
1Claude Opus 5.5Anthropic67.7%vendor2026-09
2Claude Fable 5.1Anthropic65%vendor2026-09
3DeepSeek V4.1 FlashDeepSeek63.9%vendor2026-09
4Claude Fable 5Anthropic63.9%2026-06
5Claude Opus 5Anthropic63.6%vendor2026-07
6GLM-5.3Z.ai62.5%vendor2026-08
7Claude Opus 4.8Anthropic57.9%2026-05
8Claude Sonnet 5Anthropic57.4%vendor2026-06
9GPT-6 AstraOpenAI57.2%vendor2026-09
10Qwen3.8 2.4T-A95BAlibaba56.2%vendor2026-08
11Qwen3.8-MaxAlibaba56.2%vendor2026-08
12Kimi K3Moonshot AI56%vendor2026-07
13GLM-5.3 FlashZ.ai55.3%vendor2026-08
14GLM-5.2Z.ai54.7%vendor2026-06
15InklingThinking Machines46%vendor2026-07

Score interpretation

Never compare a with-tools score against a closed-book one. The gap is roughly ten points at the frontier, and it measures the tools, not the model. Because the HLE dataset is public, granting a model search access also lets it retrieve discussion of the questions themselves, which is exactly why the maintainers keep the official board closed-book. Treat every number in this column as vendor-reported and contamination-exposed, and treat the closed-book board as the cleaner measurement of what a model actually knows.

60%+
Frontier with tools. Expert-level answers across most disciplines.
45-60%
Strong. Handles hard questions when retrieval cooperates.
25-45%
Mixed. Reliable on mainstream topics, weak at the edges.
< 25%
Below the useful bar for expert work, even with tools.

Why this matters for AI agents

Tool-augmented reasoning is how agents actually run in production, so this mode is closer to real deployment than the closed-book board. It is simply not a knowledge measurement. Read this column as an upper bound on what a model plus its retrieval stack can do on expert questions, and read closed-book HLE when you want to know what the weights themselves contain.

Other benchmarks

SWE-bench
Live SWE-bench leaderboard for major AI models. Real-world software engineering tasks from GitHub issues. Pricing per model and rankings updated weekly on TensorFeed.
MMLU-Pro
Live MMLU-Pro leaderboard for major AI models. General knowledge and reasoning across 57 subjects. Updated weekly with pricing per model on TensorFeed.
HumanEval
Live HumanEval leaderboard for major AI models. Python code generation and problem solving. Updated weekly with pricing on TensorFeed.
GPQA Diamond
Live GPQA Diamond leaderboard for major AI models. Graduate-level physics, chemistry, and biology questions. Updated on TensorFeed with pricing per model.
MATH
Live MATH benchmark leaderboard for major AI models. Competition-level mathematics problems. Updated weekly with pricing per model on TensorFeed.
OSWorld 2.0
Live OSWorld 2.0 leaderboard ranking AI models on long-horizon computer use across real desktop applications. Binary and partial scoring explained, updated on TensorFeed.
BrowseComp
Live BrowseComp leaderboard for AI agentic web search. 1,266 hard-to-find questions, why every score is self-reported, and how context strategy moves results. On TensorFeed.
FrontierCode v1.1
Live FrontierCode v1.1 leaderboard. Scores whether an AI patch is mergeable, not just test-passing, graded by open-source maintainers against a rubric. On TensorFeed.
Terminal-Bench 4.0
Live Terminal-Bench 4.0 leaderboard on one independent harness. 66 long-horizon terminal tasks across software, science, ML, and security, scored all or nothing.

Premium API: time-series for Humanity's Last Exam (tools)

The leaderboard above is a snapshot. Want to see how a model's Humanity's Last Exam (tools) score has moved over the last 30-90 days, or set a webhook that fires when a score crosses a threshold? The premium API has both:

Humanity's Last Exam (tools) source ·Last refreshed 2026-09-23·Max score 100