Skip to content
All systems operational0 AI providers monitored, polled every 2 minutes
Live status
All benchmarks

SWE-bench leaderboard

SWE-bench evaluates language models on their ability to resolve real GitHub issues from popular Python repositories. The model is given an issue description and the repository state, and must produce a patch that resolves the issue and passes the project's existing test suite. SWE-bench is the benchmark that most closely tracks "useful for autonomous coding agents" because the tasks are not toy problems, the success criteria is the project's actual tests, and the input footprint forces the model to reason over real-world code at scale.

Current leader
Claude Opus 5(Anthropic)97%

Last refreshed 2026-09-23. 46 models scored on this benchmark.

Full leaderboard

#ModelProviderScoreReleased
1Claude Opus 5Anthropic97%independent2026-07
2GPT-5.6 SolOpenAI96.2%independent2026-07
3Grok 4.6xAI95.6%independent2026-08
4GLM-5.3Z.ai95.4%independent2026-08
5GPT-5.6 TerraOpenAI95.4%independent2026-07
6Claude Fable 5Anthropic95%2026-06
7Kimi K3Moonshot AI93.4%independent2026-07
8GPT-5.6 LunaOpenAI93%independent2026-07
9GLM-5.3 FlashZ.ai92%independent2026-08
10Claude Opus 4.8Anthropic88.6%2026-05
11Claude Opus 4.7Anthropic87.6%2026-04
12Grok 4.5xAI86.6%independent2026-07
13Muse Spark 1.2Meta86.6%independent2026-08
14Qwen3.8 27BAlibaba86%independent2026-08
15Qwen3.8-MaxAlibaba85.6%independent2026-08
16Claude Sonnet 5Anthropic85.2%vendor2026-06
17GLM-5.2Z.ai82.8%independent2026-06
18GPT-5.5OpenAI82.6%2026-04
19Muse Spark 1.1Meta82%independent2026-07
20Gemini 3.7 FlashGoogle80.8%independent2026-08
21Claude Opus 4.6Anthropic80.8%2026-03
22DeepSeek V4 ProDeepSeek80.6%2026-04
23MiniMax M3MiniMax80.5%vendor2026-06
24Gemini 3.8 FlashGoogle80%independent2026-09
25Gemini 3.6 FlashGoogle79.6%independent2026-07
26Claude Sonnet 4.6Anthropic79.6%2026-02
27DeepSeek V4 FlashDeepSeek79%2026-04
28Gemini 3.5 FlashGoogle78.8%independent2026-05
29InklingThinking Machines77.6%vendor2026-07
30Mistral Medium 3.5Mistral77.6%2026-05
31Muse Glimmer 30BMeta76%vendor2026-08
32Gemini 3.5 Flash-LiteGoogle75%independent2026-07
33Claude Haiku 4.5Anthropic73.3%2026-01
34MAI-Code-1.1-FlashMicrosoft72.6%vendor2026-08
35MAI-Code-1-FlashMicrosoft71.6%vendor2026-06
36Qwen3.7-MaxAlibaba68.8%independent2026-05
37Gemini 2.5 ProGoogle63.8%2026-01
38Ternary Bonsai 2 27BPrismML60.8%vendor2026-09
39Nemotron 3.5 LightningNVIDIA51.56%vendor2026-08
40o3-miniOpenAI49.3%2025-01
41o1OpenAI48.9%2024-12
42Mistral LargeMistral47.2%2025-11
43DeepSeek V3DeepSeek42%2025-12
44GPT-4.5OpenAI38%2025-12
45GPT-4oOpenAI33.2%2024-05
46Llama 4 MaverickMeta24%2025-04

Score interpretation

Scores are reported as resolution rate (% of issues correctly patched). The headline number on TensorFeed is the SWE-bench Verified subset, the human-validated tasks where the test suite has been confirmed to be a fair signal. Mind the gap between sources: the official swebench.com board tops out around 79% and has taken no new submission since early 2026, while vendor launch materials and independent evaluators now report figures in the low to mid nineties on their own harnesses. Both are real measurements of different setups, and the spread between them is wider than the spread between most models. Anything above 60% is a genuinely useful coding agent; above 90% on a vendor harness, check which harness before you compare. Vals AI archived its SWE-bench Verified board on September 1, 2026 as saturated and no longer runs new models on it, and the September 2026 flagships do not report it, so newer models often have no cell here; see Terminal-Bench 4.0 for a board that still separates the frontier.

70%+
Frontier-class. Genuinely useful coding agent territory.
50-70%
Production-ready for assisted coding workflows.
30-50%
Useful for narrow tasks but not autonomous agents.
< 30%
Plausible-looking code that often does not work.

Why this matters for AI agents

If you are building a coding agent, this is the benchmark that matters most. Models with high SWE-bench scores produce patches that compile, pass tests, and respect existing patterns in the codebase. Models with low SWE-bench scores produce code that looks plausible but breaks the build.

Other benchmarks

MMLU-Pro
Live MMLU-Pro leaderboard for major AI models. General knowledge and reasoning across 57 subjects. Updated weekly with pricing per model on TensorFeed.
HumanEval
Live HumanEval leaderboard for major AI models. Python code generation and problem solving. Updated weekly with pricing on TensorFeed.
GPQA Diamond
Live GPQA Diamond leaderboard for major AI models. Graduate-level physics, chemistry, and biology questions. Updated on TensorFeed with pricing per model.
MATH
Live MATH benchmark leaderboard for major AI models. Competition-level mathematics problems. Updated weekly with pricing per model on TensorFeed.
OSWorld 2.0
Live OSWorld 2.0 leaderboard ranking AI models on long-horizon computer use across real desktop applications. Binary and partial scoring explained, updated on TensorFeed.
BrowseComp
Live BrowseComp leaderboard for AI agentic web search. 1,266 hard-to-find questions, why every score is self-reported, and how context strategy moves results. On TensorFeed.
FrontierCode v1.1
Live FrontierCode v1.1 leaderboard. Scores whether an AI patch is mergeable, not just test-passing, graded by open-source maintainers against a rubric. On TensorFeed.
Terminal-Bench 4.0
Live Terminal-Bench 4.0 leaderboard on one independent harness. 66 long-horizon terminal tasks across software, science, ML, and security, scored all or nothing.
Humanity's Last Exam (tools)
Live Humanity's Last Exam with-tools leaderboard on TensorFeed. Why the tool-augmented mode is not the same benchmark as closed-book HLE, and how far apart the two run.

Premium API: time-series for SWE-bench

The leaderboard above is a snapshot. Want to see how a model's SWE-bench score has moved over the last 30-90 days, or set a webhook that fires when a score crosses a threshold? The premium API has both:

SWE-bench source ·Last refreshed 2026-09-23·Max score 100