Skip to content
All systems operational0 AI providers monitored, polled every 2 minutes
Live status
All benchmarks

MATH leaderboard

The MATH benchmark consists of 12,500 competition-level mathematics problems sourced from AMC, AIME, and Putnam-style competitions. Each problem requires multi-step algebraic, geometric, or combinatorial reasoning, and the answer must match exactly (no partial credit). MATH is one of the toughest standardized math benchmarks for LLMs.

Current leader
GPT-5.5(OpenAI)95.8%

Last refreshed 2026-09-23. 18 models scored on this benchmark.

Full leaderboard

#ModelProviderScoreReleased
1GPT-5.5OpenAI95.8%2026-04
2o1OpenAI94.6%2024-12
3Claude Opus 4.7Anthropic93.1%2026-04
4DeepSeek V4 ProDeepSeek92.4%2026-04
5Claude Opus 4.6Anthropic91.8%2026-03
6Gemini 2.5 ProGoogle90.5%2026-01
7GPT-4.5OpenAI88.2%2025-12
8o3-miniOpenAI87.1%2025-01
9Llama 4 MaverickMeta86.7%2025-04
10DeepSeek V3DeepSeek85.9%2025-12
11Claude Sonnet 4.6Anthropic85.4%2026-02
12DeepSeek V4 FlashDeepSeek82.1%2026-04
13GPT-4oOpenAI81.3%2024-05
14Mistral LargeMistral80.4%2025-11
15Llama 4 ScoutMeta79.8%2025-04
16Gemini 2.0 FlashGoogle77.2%2025-02
17Claude Haiku 4.5Anthropic74.6%2026-01
18Mistral SmallMistral68.9%2025-01

Score interpretation

Scores are exact-match accuracy on the test set. As of 2026 the frontier is in the mid-90s, but the variance between problem categories is high: most models do well on AMC-level algebra and worse on AIME-level combinatorics or proof-style problems.

90%+
Frontier. Solves most competition-level problems.
70-90%
Strong on AMC-level, struggles on AIME/Putnam.
40-70%
Useful for routine math but unreliable on multi-step problems.
< 40%
Weak general math; unreliable for quantitative tasks.

Why this matters for AI agents

MATH performance correlates strongly with multi-step reasoning capability in general. Models that can carry algebraic state through 5-10 steps on MATH problems tend to be the same models that can carry argumentative state through long agent workflows. If your agent does any quantitative work, MATH is a useful proxy.

Other benchmarks

SWE-bench
Live SWE-bench leaderboard for major AI models. Real-world software engineering tasks from GitHub issues. Pricing per model and rankings updated weekly on TensorFeed.
MMLU-Pro
Live MMLU-Pro leaderboard for major AI models. General knowledge and reasoning across 57 subjects. Updated weekly with pricing per model on TensorFeed.
HumanEval
Live HumanEval leaderboard for major AI models. Python code generation and problem solving. Updated weekly with pricing on TensorFeed.
GPQA Diamond
Live GPQA Diamond leaderboard for major AI models. Graduate-level physics, chemistry, and biology questions. Updated on TensorFeed with pricing per model.
OSWorld 2.0
Live OSWorld 2.0 leaderboard ranking AI models on long-horizon computer use across real desktop applications. Binary and partial scoring explained, updated on TensorFeed.
BrowseComp
Live BrowseComp leaderboard for AI agentic web search. 1,266 hard-to-find questions, why every score is self-reported, and how context strategy moves results. On TensorFeed.
FrontierCode v1.1
Live FrontierCode v1.1 leaderboard. Scores whether an AI patch is mergeable, not just test-passing, graded by open-source maintainers against a rubric. On TensorFeed.
Terminal-Bench 4.0
Live Terminal-Bench 4.0 leaderboard on one independent harness. 66 long-horizon terminal tasks across software, science, ML, and security, scored all or nothing.
Humanity's Last Exam (tools)
Live Humanity's Last Exam with-tools leaderboard on TensorFeed. Why the tool-augmented mode is not the same benchmark as closed-book HLE, and how far apart the two run.

Premium API: time-series for MATH

The leaderboard above is a snapshot. Want to see how a model's MATH score has moved over the last 30-90 days, or set a webhook that fires when a score crosses a threshold? The premium API has both:

MATH source ·Last refreshed 2026-09-23·Max score 100