AI Benchmarks
Compare leading AI models across standardized benchmarks. Last updated 2026-09-23.
Compare specific models, side-by-side
Pick any 2 to 5 models to put head-to-head across benchmarks, pricing, and context windows. Example pairs: Claude Opus 5.5 vs GPT-6 Astra, GPT-6 Sol vs Claude Sonnet 5, Gemini 3.8 Flash vs DeepSeek V4.1 Flash, open-source vs frontier.
How do you know if Claude is smarter than GPT? How does an open-weight model like Kimi K3 stack up against Gemini? Benchmarks provide the answer. These standardized tests measure specific AI capabilities across diverse domains and let us compare models objectively. They're imperfect (benchmarks are often gamed), but they're the only shared language we have for understanding AI progress.
MMLU measures broad knowledge across multiple choice questions across chemistry, history, law, and 50+ other domains. A score of 92 percent means the model answers 92 out of 100 random questions correctly across all topics. MMLU is the closest we have to a general intelligence test for AI. HumanEval tests code generation: the model writes functions to solve programming problems that humans created. GPQA (Graduate-Level Google-Proof Questions) is deliberately hard, asking obscure questions that require deep expertise. MATH benchmarks raw mathematical reasoning. SWE-bench tests software engineering tasks: given a failing test and a codebase, can the model write code to fix it?
No single benchmark captures everything. A model that excels at MMLU might struggle with code. Benchmarks have been leaked and learned during training. And real-world performance depends on your specific task, how you prompt, and how you integrate the model into your system. Use this data to narrow the field of candidates. Then test the finalists on your actual workloads. We've also collected this data in our model comparison tool for side-by-side analysis.
Terminal-Bench 4.0: Long-horizon agentic work in a terminal across software, science, ML, operations, hardware, security, and media (66 tasks, all-or-nothing verifiers). Max score: 100.
| Rank | Model | Provider | Score↓ | Released |
|---|---|---|---|---|
| #1 | Claude Opus 5.5 | Anthropic | 61.6/ 100independent | 2026-09 |
| #2 | GPT-6 Astra | OpenAI | 57.1/ 100independent | 2026-09 |
| #3 | Claude Fable 5.1 | Anthropic | 49.5/ 100independent | 2026-09 |
| #4 | Claude Opus 5 | Anthropic | 45.5/ 100independent | 2026-07 |
| #5 | Grok 4.7 | xAI | 28.3/ 100independent | 2026-09 |
| #6 | GPT-5.6 Sol | OpenAI | 27.8/ 100independent | 2026-07 |
| #7 | GPT-5.6 Terra | OpenAI | 26.3/ 100independent | 2026-07 |
| #8 | GLM-5.3 | Z.ai | 25.3/ 100independent | 2026-08 |
| #9 | MiMo-V2.6-Pro | Xiaomi | 24.8/ 100independent | 2026-09 |
| #10 | Qwen3.8-Max | Alibaba | 24.8/ 100independent | 2026-08 |
| #11 | Claude Fable 5 | Anthropic | 22.7/ 100independent | 2026-06 |
| #12 | MiMo-V2.6-Flash | Xiaomi | 21.2/ 100independent | 2026-09 |
| #13 | GLM-5.3 Flash | Z.ai | 19.7/ 100independent | 2026-08 |
| #14 | Grok 4.6 | xAI | 17.2/ 100independent | 2026-08 |
| #15 | Claude Opus 4.8 | Anthropic | 16.2/ 100independent | 2026-05 |
| #16 | Muse Spark 1.3 | Meta | 15.2/ 100independent | 2026-09 |
| #17 | Gemini 3.8 Flash | 13.1/ 100independent | 2026-09 | |
| #18 | Kimi K3 | Moonshot AI | 12.6/ 100independent | 2026-07 |
| #19 | DeepSeek V4.1 Flash | DeepSeek | 11.6/ 100independent | 2026-09 |
| #20 | Claude Sonnet 5 | Anthropic | 8.1/ 100independent | 2026-06 |
| #21 | Gemini 3.7 Flash | 6.1/ 100independent | 2026-08 | |
| #22 | GPT-5.6 Luna | OpenAI | 4.5/ 100independent | 2026-07 |
| Not reported on Terminal-Bench 4.0 | ||||
| - | Claude Haiku 4.5 | Anthropic | not reported | 2026-01 |
| - | Claude Opus 4.6 | Anthropic | not reported | 2026-03 |
| - | Claude Opus 4.7 | Anthropic | not reported | 2026-04 |
| - | Claude Sonnet 4.6 | Anthropic | not reported | 2026-02 |
| - | DeepSeek V3 | DeepSeek | not reported | 2025-12 |
| - | DeepSeek V4 Flash | DeepSeek | not reported | 2026-04 |
| - | DeepSeek V4 Pro | DeepSeek | not reported | 2026-04 |
| - | Fugu Max | Sakana AI | not reported | 2026-09 |
| - | Fugu Ultra v2.0 | Sakana AI | not reported | 2026-09 |
| - | Gemini 2.0 Flash | not reported | 2025-02 | |
| - | Gemini 2.5 Pro | not reported | 2026-01 | |
| - | Gemini 3.1 Flash-Lite | not reported | 2026-03 | |
| - | Gemini 3.5 Flash | not reported | 2026-05 | |
| - | Gemini 3.5 Flash-Lite | not reported | 2026-07 | |
| - | Gemini 3.6 Flash | not reported | 2026-07 | |
| - | GLM-5.2 | Z.ai | not reported | 2026-06 |
| - | GPT-4.5 | OpenAI | not reported | 2025-12 |
| - | GPT-4o | OpenAI | not reported | 2024-05 |
| - | GPT-5.5 | OpenAI | not reported | 2026-04 |
| - | GPT-6 Luna | OpenAI | not reported | 2026-09 |
| - | GPT-6 Sol | OpenAI | not reported | 2026-09 |
| - | Grok 4.5 | xAI | not reported | 2026-07 |
| - | Hy4 preview | Tencent | not reported | 2026-08 |
| - | Inkling | Thinking Machines | not reported | 2026-07 |
| - | Llama 4 Maverick | Meta | not reported | 2025-04 |
| - | Llama 4 Scout | Meta | not reported | 2025-04 |
| - | LongCat-2.0 | Meituan | not reported | 2026-06 |
| - | MAI-Code-1-Flash | Microsoft | not reported | 2026-06 |
| - | MAI-Code-1.1-Flash | Microsoft | not reported | 2026-08 |
| - | MiniMax M3 | MiniMax | not reported | 2026-06 |
| - | Mistral Large | Mistral | not reported | 2025-11 |
| - | Mistral Medium 3.5 | Mistral | not reported | 2026-05 |
| - | Mistral Small | Mistral | not reported | 2025-01 |
| - | Muse Glimmer 30B | Meta | not reported | 2026-08 |
| - | Muse Spark 1.1 | Meta | not reported | 2026-07 |
| - | Muse Spark 1.2 | Meta | not reported | 2026-08 |
| - | Nemotron 3.5 Lightning | NVIDIA | not reported | 2026-08 |
| - | o1 | OpenAI | not reported | 2024-12 |
| - | o3-mini | OpenAI | not reported | 2025-01 |
| - | Qwen3.7-Max | Alibaba | not reported | 2026-05 |
| - | Qwen3.8 2.4T-A95B | Alibaba | not reported | 2026-08 |
| - | Qwen3.8 27B | Alibaba | not reported | 2026-08 |
| - | Qwen3.8-Flash | Alibaba | not reported | 2026-08 |
| - | Ternary Bonsai 2 27B | PrismML | not reported | 2026-09 |
