Skip to content
All systems operational0 AI providers monitored, polled every 2 minutes
Live status
All benchmarks

BrowseComp leaderboard

BrowseComp tests persistent web research. Its 1,266 questions have short, verifiable answers that are genuinely difficult to locate, requiring an agent to chase entangled information across many pages rather than retrieve a single fact. OpenAI built it and released it in 2025, and it has become the standard reference for deep-research and agentic-search products.

Current leader
GPT-6 Astra(OpenAI)91.5%

Last refreshed 2026-09-23. 13 models scored on this benchmark.

Full leaderboard

#ModelProviderScoreReleased
1GPT-6 AstraOpenAI91.5%vendor2026-09
2Kimi K3Moonshot AI91.2%vendor2026-07
3Claude Opus 5Anthropic90.8%vendor2026-07
4GPT-5.6 SolOpenAI90.4%vendor2026-07
5GPT-5.6 TerraOpenAI87.5%vendor2026-07
6Claude Fable 5Anthropic87.4%2026-06
7Claude Sonnet 5Anthropic84.7%vendor2026-06
8Claude Opus 4.8Anthropic84.3%2026-05
9MiniMax M3MiniMax83.5%vendor2026-06
10GPT-5.6 LunaOpenAI83.3%vendor2026-07
11LongCat-2.0Meituan79.9%vendor2026-06
12InklingThinking Machines77.1%vendor2026-07
13Nemotron 3.5 LightningNVIDIA36.97%vendor2026-08

Score interpretation

Two caveats that matter more here than on any other benchmark we track. First, there is no official leaderboard and no independent evaluator: every published BrowseComp figure is self-reported by the lab that produced it. Second, scores move sharply with context-management strategy, so a single-agent run, a multi-agent run, and a run using context compaction are not the same measurement even for the same model. What a BrowseComp number really ranks is a full system, the model plus its tools plus its context policy, not a bare model. Frontier systems now cluster within a couple of points of each other, which suggests the benchmark is approaching saturation at the top.

85%+
Frontier research agent. Finds answers most humans would give up on.
70-85%
Strong. Reliable on hard lookups, occasional dead ends.
40-70%
Useful for ordinary search, weak on genuinely buried facts.
< 40%
Not a research agent. Expect confident wrong answers.

Why this matters for AI agents

If you are building research agents, this is the closest published proxy for whether the thing will actually find an obscure answer instead of confidently inventing one. Just do not treat small gaps between top models as real. The measurement noise from differing harnesses and context strategies is larger than the differences between the leaders.

Other benchmarks

SWE-bench
Live SWE-bench leaderboard for major AI models. Real-world software engineering tasks from GitHub issues. Pricing per model and rankings updated weekly on TensorFeed.
MMLU-Pro
Live MMLU-Pro leaderboard for major AI models. General knowledge and reasoning across 57 subjects. Updated weekly with pricing per model on TensorFeed.
HumanEval
Live HumanEval leaderboard for major AI models. Python code generation and problem solving. Updated weekly with pricing on TensorFeed.
GPQA Diamond
Live GPQA Diamond leaderboard for major AI models. Graduate-level physics, chemistry, and biology questions. Updated on TensorFeed with pricing per model.
MATH
Live MATH benchmark leaderboard for major AI models. Competition-level mathematics problems. Updated weekly with pricing per model on TensorFeed.
OSWorld 2.0
Live OSWorld 2.0 leaderboard ranking AI models on long-horizon computer use across real desktop applications. Binary and partial scoring explained, updated on TensorFeed.
FrontierCode v1.1
Live FrontierCode v1.1 leaderboard. Scores whether an AI patch is mergeable, not just test-passing, graded by open-source maintainers against a rubric. On TensorFeed.
Terminal-Bench 4.0
Live Terminal-Bench 4.0 leaderboard on one independent harness. 66 long-horizon terminal tasks across software, science, ML, and security, scored all or nothing.
Humanity's Last Exam (tools)
Live Humanity's Last Exam with-tools leaderboard on TensorFeed. Why the tool-augmented mode is not the same benchmark as closed-book HLE, and how far apart the two run.

Premium API: time-series for BrowseComp

The leaderboard above is a snapshot. Want to see how a model's BrowseComp score has moved over the last 30-90 days, or set a webhook that fires when a score crosses a threshold? The premium API has both:

BrowseComp source ·Last refreshed 2026-09-23·Max score 100