Skip to content
All systems operational0 AI providers monitored, polled every 2 minutes
Live status
All benchmarks

FrontierCode v1.1 leaderboard

FrontierCode asks a harder question than most coding benchmarks: not whether a patch passes tests, but whether a maintainer would merge it. Cognition built it with roughly 36 open-source maintainers from projects including Celery, Budibase, uppy, and Mattermost, who authored tasks against real issues in their own repositories. Grading combines held-out tests with a maintainer-written rubric covering behavioral correctness, regression safety, test quality, scope discipline, build and lint cleanliness, and adherence to project conventions. The Main split is the 100 hardest tasks; an Extended split covers all 150.

Current leader
Claude Opus 5.5(Anthropic)54.6%

Last refreshed 2026-09-23. 28 models scored on this benchmark.

Full leaderboard

#ModelProviderScoreReleased
1Claude Opus 5.5Anthropic54.6%independent2026-09
2Claude Fable 5Anthropic53.5%2026-06
3Claude Opus 5Anthropic53.4%vendor2026-07
4GPT-6 AstraOpenAI53.3%vendor2026-09
5Claude Fable 5.1Anthropic50.9%independent2026-09
6GPT-6 SolOpenAI49.3%independent2026-09
7Grok 4.6xAI48%independent2026-08
8GPT-5.6 SolOpenAI47.5%independent2026-07
9Claude Opus 4.8Anthropic46.5%2026-05
10Kimi K3Moonshot AI44.2%independent2026-07
11Gemini 3.7 FlashGoogle43.6%independent2026-08
12GPT-5.5OpenAI43%independent2026-04
13Claude Sonnet 5Anthropic42.7%independent2026-06
14GPT-6 LunaOpenAI42.4%independent2026-09
15Grok 4.5xAI42.4%independent2026-07
16GPT-5.6 TerraOpenAI41.3%independent2026-07
17Gemini 3.8 FlashGoogle41.2%independent2026-09
18GLM-5.3Z.ai40.1%independent2026-08
19GPT-5.6 LunaOpenAI39.8%independent2026-07
20Claude Opus 4.7Anthropic38.5%independent2026-04
21Gemini 3.6 FlashGoogle34.4%independent2026-07
22GLM-5.3 FlashZ.ai31.8%independent2026-08
23Claude Opus 4.6Anthropic26.6%independent2026-03
24GLM-5.2Z.ai24.5%independent2026-06
25Claude Sonnet 4.6Anthropic24.3%independent2026-02
26MiniMax M3MiniMax14.7%independent2026-06
27InklingThinking Machines14%independent2026-07
28Mistral Medium 3.5Mistral8%independent2026-05

Score interpretation

TensorFeed reports the Main split. Two things to watch. Scores vary substantially by reasoning effort, and the leaderboard publishes a separate figure per effort level, so a headline number is meaningless without knowing which one it came from. And the benchmark is deliberately unpublished, with no public repo, which means aggregator sites quoting FrontierCode figures for models that are not on Cognition's own board are reporting something other than a FrontierCode run. Check the maintainer board before trusting a number.

50%+
Frontier. Produces mergeable work on the hardest tasks about half the time.
35-50%
Strong. Real contributions, but review is still mandatory.
20-35%
Useful drafts. Expect substantial rework before merge.
< 20%
Output is a starting point, not a contribution.

Why this matters for AI agents

Test-passing and mergeable are different bars, and the gap between them is where most coding-agent disappointment lives. A patch that turns tests green while sprawling across unrelated files, ignoring project conventions, or quietly weakening a test is a patch a human has to redo. This is the only benchmark we track that puts a maintainer's judgment in the scoring loop, which makes it the best available signal for whether agent output will survive code review.

Other benchmarks

SWE-bench
Live SWE-bench leaderboard for major AI models. Real-world software engineering tasks from GitHub issues. Pricing per model and rankings updated weekly on TensorFeed.
MMLU-Pro
Live MMLU-Pro leaderboard for major AI models. General knowledge and reasoning across 57 subjects. Updated weekly with pricing per model on TensorFeed.
HumanEval
Live HumanEval leaderboard for major AI models. Python code generation and problem solving. Updated weekly with pricing on TensorFeed.
GPQA Diamond
Live GPQA Diamond leaderboard for major AI models. Graduate-level physics, chemistry, and biology questions. Updated on TensorFeed with pricing per model.
MATH
Live MATH benchmark leaderboard for major AI models. Competition-level mathematics problems. Updated weekly with pricing per model on TensorFeed.
OSWorld 2.0
Live OSWorld 2.0 leaderboard ranking AI models on long-horizon computer use across real desktop applications. Binary and partial scoring explained, updated on TensorFeed.
BrowseComp
Live BrowseComp leaderboard for AI agentic web search. 1,266 hard-to-find questions, why every score is self-reported, and how context strategy moves results. On TensorFeed.
Terminal-Bench 4.0
Live Terminal-Bench 4.0 leaderboard on one independent harness. 66 long-horizon terminal tasks across software, science, ML, and security, scored all or nothing.
Humanity's Last Exam (tools)
Live Humanity's Last Exam with-tools leaderboard on TensorFeed. Why the tool-augmented mode is not the same benchmark as closed-book HLE, and how far apart the two run.

Premium API: time-series for FrontierCode v1.1

The leaderboard above is a snapshot. Want to see how a model's FrontierCode v1.1 score has moved over the last 30-90 days, or set a webhook that fires when a score crosses a threshold? The premium API has both:

FrontierCode v1.1 source ·Last refreshed 2026-09-23·Max score 100