Skip to content
All systems operational0 AI providers monitored, polled every 2 minutes
Live status
All benchmarks

Terminal-Bench 4.0 leaderboard

Terminal-Bench 4.0 measures whether an agent can finish real, long-running work in a terminal. It is maintained by the Terminal-Bench team on the Harbor framework and was released at the end of August 2026 as a revision of 3.0: eight tasks were removed (saturated, refusal-prone, publicly solved, or broken), nineteen were fixed, and resource limits were recalibrated around a flat eight-hour agent timeout. The 66 tasks span software, science, machine learning, operations, hardware, security, and media, and each one produces a real deliverable, such as a running service, a proof, a CAD model, a kernel, or a forensic report. About three quarters of the set falls outside traditional software engineering, and the median task carries an expert time estimate of around four hours.

Current leader
Claude Opus 5.5(Anthropic)61.62%

Last refreshed 2026-09-23. 22 models scored on this benchmark.

Full leaderboard

#ModelProviderScoreReleased
1Claude Opus 5.5Anthropic61.62%independent2026-09
2GPT-6 AstraOpenAI57.07%independent2026-09
3Claude Fable 5.1Anthropic49.49%independent2026-09
4Claude Opus 5Anthropic45.45%independent2026-07
5Grok 4.7xAI28.28%independent2026-09
6GPT-5.6 SolOpenAI27.78%independent2026-07
7GPT-5.6 TerraOpenAI26.26%independent2026-07
8GLM-5.3Z.ai25.25%independent2026-08
9MiMo-V2.6-ProXiaomi24.75%independent2026-09
10Qwen3.8-MaxAlibaba24.75%independent2026-08
11Claude Fable 5Anthropic22.73%independent2026-06
12MiMo-V2.6-FlashXiaomi21.21%independent2026-09
13GLM-5.3 FlashZ.ai19.7%independent2026-08
14Grok 4.6xAI17.17%independent2026-08
15Claude Opus 4.8Anthropic16.16%independent2026-05
16Muse Spark 1.3Meta15.15%independent2026-09
17Gemini 3.8 FlashGoogle13.13%independent2026-09
18Kimi K3Moonshot AI12.63%independent2026-07
19DeepSeek V4.1 FlashDeepSeek11.62%independent2026-09
20Claude Sonnet 5Anthropic8.08%independent2026-06
21Gemini 3.7 FlashGoogle6.06%independent2026-08
22GPT-5.6 LunaOpenAI4.55%independent2026-07

Score interpretation

Every task has its own verifier and passes only if every check passes, so there is no partial credit; timeouts and unverifiable submissions score zero, and refusals count as failures. TensorFeed reports the Vals AI run, which uses one harness (Terminus 2) for every model and averages three runs, so the column compares like with like. Two things move the number more than people expect. The harness: the same model can land ten points apart on Vals, Artificial Analysis, and the official board, where each lab runs its own agent. And fallback: when a safeguard reroutes a task to another model, Vals counts the attempt, and the per-cell note says how many attempts that was. Terminal-Bench 2.x and 4.0 share no tasks, so never compare scores across versions.

55%+
Frontier. Finishes most long-horizon terminal tasks unattended.
35-55%
Strong. Completes a meaningful share of multi-hour work.
15-35%
Capable on shorter tasks; expect long runs to stall.
< 15%
Not yet an unattended terminal agent.

Why this matters for AI agents

Most agent failures in production are not a single bad answer; they are a long task that drifts, stalls, or gives up halfway. Terminal-Bench 4.0 is built out of exactly those tasks, with hours-long horizons and binary outcomes, which makes it a better predictor of whether an agent will finish unattended work than any single-turn benchmark. It is also one of the few boards still separating the September 2026 frontier, where the saturated academic sets no longer do.

Other benchmarks

SWE-bench
Live SWE-bench leaderboard for major AI models. Real-world software engineering tasks from GitHub issues. Pricing per model and rankings updated weekly on TensorFeed.
MMLU-Pro
Live MMLU-Pro leaderboard for major AI models. General knowledge and reasoning across 57 subjects. Updated weekly with pricing per model on TensorFeed.
HumanEval
Live HumanEval leaderboard for major AI models. Python code generation and problem solving. Updated weekly with pricing on TensorFeed.
GPQA Diamond
Live GPQA Diamond leaderboard for major AI models. Graduate-level physics, chemistry, and biology questions. Updated on TensorFeed with pricing per model.
MATH
Live MATH benchmark leaderboard for major AI models. Competition-level mathematics problems. Updated weekly with pricing per model on TensorFeed.
OSWorld 2.0
Live OSWorld 2.0 leaderboard ranking AI models on long-horizon computer use across real desktop applications. Binary and partial scoring explained, updated on TensorFeed.
BrowseComp
Live BrowseComp leaderboard for AI agentic web search. 1,266 hard-to-find questions, why every score is self-reported, and how context strategy moves results. On TensorFeed.
FrontierCode v1.1
Live FrontierCode v1.1 leaderboard. Scores whether an AI patch is mergeable, not just test-passing, graded by open-source maintainers against a rubric. On TensorFeed.
Humanity's Last Exam (tools)
Live Humanity's Last Exam with-tools leaderboard on TensorFeed. Why the tool-augmented mode is not the same benchmark as closed-book HLE, and how far apart the two run.

Premium API: time-series for Terminal-Bench 4.0

The leaderboard above is a snapshot. Want to see how a model's Terminal-Bench 4.0 score has moved over the last 30-90 days, or set a webhook that fires when a score crosses a threshold? The premium API has both:

Terminal-Bench 4.0 source ·Last refreshed 2026-09-23·Max score 100