Skip to content
All systems operational0 AI providers monitored, polled every 2 minutes
Live status

AI Benchmarks

Compare leading AI models across standardized benchmarks. Last updated 2026-09-23.

Compare specific models, side-by-side

Pick any 2 to 5 models to put head-to-head across benchmarks, pricing, and context windows. Example pairs: Claude Opus 5.5 vs GPT-6 Astra, GPT-6 Sol vs Claude Sonnet 5, Gemini 3.8 Flash vs DeepSeek V4.1 Flash, open-source vs frontier.

How do you know if Claude is smarter than GPT? How does an open-weight model like Kimi K3 stack up against Gemini? Benchmarks provide the answer. These standardized tests measure specific AI capabilities across diverse domains and let us compare models objectively. They're imperfect (benchmarks are often gamed), but they're the only shared language we have for understanding AI progress.

MMLU measures broad knowledge across multiple choice questions across chemistry, history, law, and 50+ other domains. A score of 92 percent means the model answers 92 out of 100 random questions correctly across all topics. MMLU is the closest we have to a general intelligence test for AI. HumanEval tests code generation: the model writes functions to solve programming problems that humans created. GPQA (Graduate-Level Google-Proof Questions) is deliberately hard, asking obscure questions that require deep expertise. MATH benchmarks raw mathematical reasoning. SWE-bench tests software engineering tasks: given a failing test and a codebase, can the model write code to fix it?

No single benchmark captures everything. A model that excels at MMLU might struggle with code. Benchmarks have been leaked and learned during training. And real-world performance depends on your specific task, how you prompt, and how you integrate the model into your system. Use this data to narrow the field of candidates. Then test the finalists on your actual workloads. We've also collected this data in our model comparison tool for side-by-side analysis.

Terminal-Bench 4.0: Long-horizon agentic work in a terminal across software, science, ML, operations, hardware, security, and media (66 tasks, all-or-nothing verifiers). Max score: 100.

RankModelProviderScore↓Released
#1Claude Opus 5.5Anthropic61.6/ 100independent2026-09
#2GPT-6 AstraOpenAI57.1/ 100independent2026-09
#3Claude Fable 5.1Anthropic49.5/ 100independent2026-09
#4Claude Opus 5Anthropic45.5/ 100independent2026-07
#5Grok 4.7xAI28.3/ 100independent2026-09
#6GPT-5.6 SolOpenAI27.8/ 100independent2026-07
#7GPT-5.6 TerraOpenAI26.3/ 100independent2026-07
#8GLM-5.3Z.ai25.3/ 100independent2026-08
#9MiMo-V2.6-ProXiaomi24.8/ 100independent2026-09
#10Qwen3.8-MaxAlibaba24.8/ 100independent2026-08
#11Claude Fable 5Anthropic22.7/ 100independent2026-06
#12MiMo-V2.6-FlashXiaomi21.2/ 100independent2026-09
#13GLM-5.3 FlashZ.ai19.7/ 100independent2026-08
#14Grok 4.6xAI17.2/ 100independent2026-08
#15Claude Opus 4.8Anthropic16.2/ 100independent2026-05
#16Muse Spark 1.3Meta15.2/ 100independent2026-09
#17Gemini 3.8 FlashGoogle13.1/ 100independent2026-09
#18Kimi K3Moonshot AI12.6/ 100independent2026-07
#19DeepSeek V4.1 FlashDeepSeek11.6/ 100independent2026-09
#20Claude Sonnet 5Anthropic8.1/ 100independent2026-06
#21Gemini 3.7 FlashGoogle6.1/ 100independent2026-08
#22GPT-5.6 LunaOpenAI4.5/ 100independent2026-07
Not reported on Terminal-Bench 4.0
-Claude Haiku 4.5Anthropicnot reported2026-01
-Claude Opus 4.6Anthropicnot reported2026-03
-Claude Opus 4.7Anthropicnot reported2026-04
-Claude Sonnet 4.6Anthropicnot reported2026-02
-DeepSeek V3DeepSeeknot reported2025-12
-DeepSeek V4 FlashDeepSeeknot reported2026-04
-DeepSeek V4 ProDeepSeeknot reported2026-04
-Fugu MaxSakana AInot reported2026-09
-Fugu Ultra v2.0Sakana AInot reported2026-09
-Gemini 2.0 FlashGooglenot reported2025-02
-Gemini 2.5 ProGooglenot reported2026-01
-Gemini 3.1 Flash-LiteGooglenot reported2026-03
-Gemini 3.5 FlashGooglenot reported2026-05
-Gemini 3.5 Flash-LiteGooglenot reported2026-07
-Gemini 3.6 FlashGooglenot reported2026-07
-GLM-5.2Z.ainot reported2026-06
-GPT-4.5OpenAInot reported2025-12
-GPT-4oOpenAInot reported2024-05
-GPT-5.5OpenAInot reported2026-04
-GPT-6 LunaOpenAInot reported2026-09
-GPT-6 SolOpenAInot reported2026-09
-Grok 4.5xAInot reported2026-07
-Hy4 previewTencentnot reported2026-08
-InklingThinking Machinesnot reported2026-07
-Llama 4 MaverickMetanot reported2025-04
-Llama 4 ScoutMetanot reported2025-04
-LongCat-2.0Meituannot reported2026-06
-MAI-Code-1-FlashMicrosoftnot reported2026-06
-MAI-Code-1.1-FlashMicrosoftnot reported2026-08
-MiniMax M3MiniMaxnot reported2026-06
-Mistral LargeMistralnot reported2025-11
-Mistral Medium 3.5Mistralnot reported2026-05
-Mistral SmallMistralnot reported2025-01
-Muse Glimmer 30BMetanot reported2026-08
-Muse Spark 1.1Metanot reported2026-07
-Muse Spark 1.2Metanot reported2026-08
-Nemotron 3.5 LightningNVIDIAnot reported2026-08
-o1OpenAInot reported2024-12
-o3-miniOpenAInot reported2025-01
-Qwen3.7-MaxAlibabanot reported2026-05
-Qwen3.8 2.4T-A95BAlibabanot reported2026-08
-Qwen3.8 27BAlibabanot reported2026-08
-Qwen3.8-FlashAlibabanot reported2026-08
-Ternary Bonsai 2 27BPrismMLnot reported2026-09
Score provenance:vendorpublished by the lab that makes the modelindependentrun by a third party such as Vals AI or Artificial AnalysisAn unmarked score means provenance is not yet recorded. The two kinds are not interchangeable: gaps on the same model and benchmark reach 10.9 points, so treat a mixed ranking with care. Each badge links to its source. Vals AI archived SWE-bench Verified, MMLU-Pro, and GPQA Diamond as saturated on September 1, 2026, so models released after that date, such as Claude Opus 5.5 and GPT-6 Sol, will not get Vals AI scores in those columns.
Last reviewed: reviewed weeklyNext review:

data/benchmarks.json. Add a row when a flagship lands or when benchmark rankings materially shift. Reminder issue opens Mondays via existing weekly-benchmarks-check workflow. Scores are a mix of vendor-reported and independent-evaluator figures and the schema does not yet record which, so verify provenance before citing a single cell.