Skip to content
All systems operational0 AI providers monitored, polled every 2 minutes
Live status

Last Updated: September 23, 2026

The Frontier Model Wars

Real-time tracking of the race between Anthropic, OpenAI, Google, and every lab pushing the edge.

71
Active models
660
Benchmark scores
...
Last update

Every frontier AI lab is running the same race. Scale up compute, scale up data, push a model through pre-training, run an increasingly elaborate post-training pipeline, stamp a release candidate, and ship. The top of the pack has never been more crowded. Anthropic, OpenAI, and Google DeepMind trade the overall lead every few weeks. xAI has closed more ground than most people expected. The open weight race now runs through Chinese labs like Moonshot, DeepSeek, Alibaba, and Z.ai, while Meta sells its Muse Spark models through a paid API. DeepSeek keeps proving that efficiency is a category of its own.

This page is where we track the state of that race in public. The leaderboard updates daily from our benchmark pipeline. The recent releases section pulls straight from the pricing database and filters to anything shipped in the last ninety days. The winners by category read off the latest benchmark scores and pricing data. The news feed at the bottom filters the full TensorFeed stream for model release coverage only.

Current Frontier Leaderboard

Ranked by relative score: each published score in our benchmark database is converted to a percentile within its own benchmark (SWE-bench, GPQA Diamond, OSWorld 2.0, Terminal-Bench 4.0, and more), and the percentiles are averaged, so a model is not rewarded or penalized for which benchmarks it reports. Models need at least two published scores to appear. Coverage still varies, so treat close scores with care.

RankModelProviderReleasedRelative ScorePricing (in / out, per 1M)
#1Claude Opus 5.5AnthropicSep 2026100.0$4.00 / $20.00
#2Claude Fable 5.1AnthropicSep 202688.5$10.00 / $50.00
#3Claude Opus 5AnthropicJul 202687.8$5.00 / $25.00
#4GPT-6 AstraOpenAISep 202685.4$10.00 / $50.00
#5Claude Fable 5AnthropicJun 202679.1$10.00 / $50.00
#6GPT-5.6 SolOpenAIJul 202679.0$4.00 / $20.00
#7GPT-5.5OpenAIApr 202675.6$5.00 / $30.00
#8Grok 4.6xAIAug 202674.6$2.00 / $6.00
#9DeepSeek V4 ProDeepSeekApr 202673.2$1.32 / $3.96
#10Claude Opus 4.7AnthropicApr 202669.6$5.00 / $25.00
#11Grok 4.5xAIJul 202665.7$2.00 / $6.00
#12Muse Spark 1.2MetaAug 202662.8$1.25 / $4.25
#13Claude Opus 4.8AnthropicMay 202661.1$5.00 / $25.00
#14Kimi K3Moonshot AIJul 202660.5$3.00 / $15.00
#15Claude Opus 4.6AnthropicMar 202660.4$5.00 / $25.00

Want to drill into a specific benchmark? See the full benchmarks page.

Head to Head Matchups

Recent frontier releases from Anthropic, OpenAI, and Google, at a glance. Each card pulls live scores from the benchmark database. The benchmarks covered differ by model, so compare individual rows; the headline number is the same relative score used in the leaderboard above.

Claude Opus 5.5

Anthropic
100.0
Relative score across 3 published benchmarks
Terminal-Bench 4.061.6
FrontierCode v1.154.6
Humanity's Last Exam (tools)67.7

GPT-6 Astra

OpenAI
85.4
Relative score across 5 published benchmarks
Terminal-Bench 4.057.1
GPQA Diamond96.0
BrowseComp91.5
FrontierCode v1.153.3
Humanity's Last Exam (tools)57.2

Gemini 3.8 Flash

Google
57.3
Relative score across 5 published benchmarks
Terminal-Bench 4.013.1
SWE-bench80.0
MMLU-Pro90.2
GPQA Diamond95.3
FrontierCode v1.141.2

Build your own matchup on the compare tool.

Recent Releases (Last 90 Days)

Every frontier model released in the last quarter, newest first. Auto-updated from the model database.

Sep 2026Claude Opus 5.5
Anthropic
Sep 2026Claude Fable 5.1
Anthropic
Sep 2026Claude Mythos 5.1
Anthropic
Sep 2026GPT-6 Astra
OpenAI
Sep 2026GPT-6 Sol
OpenAI
Sep 2026GPT-6 Luna
OpenAI
Sep 2026Gemini 3.8 Flash
Google
Sep 2026Muse Spark 1.3
Meta
Sep 2026Muse Spark 1.3 Contributor
Meta
Sep 2026DeepSeek V4.1 Flash
DeepSeek
Sep 2026Qwen3.8-Omni-Flash
Alibaba
Sep 2026Grok 4.7
xAI
Sep 2026GLM-5.3-FlashX
Z.ai (Zhipu AI)
Sep 2026Fugu Ultra v2.0
Sakana AI
Sep 2026Fugu Max
Sakana AI
Sep 2026Pareto
Unbiased
Sep 2026MiMo-V2.6-Pro
Xiaomi
Sep 2026MiMo-V2.6-Flash
Xiaomi
Aug 2026Gemini 3.7 Flash
Google
Aug 2026Muse Glimmer 30B
Meta
Aug 2026Muse Spark 1.2
Meta
Aug 2026Muse Spark 1.2 Contributor
Meta
Aug 2026Qwen3.8 27B
Alibaba
Aug 2026Qwen3.8 2.4T-A95B
Alibaba
Aug 2026Qwen3.8-Max
Alibaba
Aug 2026Qwen3.8-Flash
Alibaba
Aug 2026Grok 4.6
xAI
Aug 2026Nemotron 3.5 Lightning
NVIDIA
Aug 2026MAI-Code-1.1-Flash
Microsoft
Aug 2026GLM-5.3
Z.ai (Zhipu AI)
Aug 2026GLM-5.3 Flash
Z.ai (Zhipu AI)
Aug 2026Hy4 preview
Tencent
Jul 2026Claude Opus 5
Anthropic
Jul 2026GPT-5.6 Sol
OpenAI
Jul 2026GPT-5.6 Terra
OpenAI
Jul 2026GPT-5.6 Luna
OpenAI
Jul 2026Gemini 3.6 Flash
Google
Jul 2026Gemini 3.5 Flash-Lite
Google
Jul 2026Muse Spark 1.1
Meta
Jul 2026Grok 4.5
xAI
Jul 2026Kimi K3
Moonshot AI
Jul 2026Laguna S 2.1
poolside
Jun 2026Claude Fable 5
Anthropic
Jun 2026Claude Sonnet 5
Anthropic
Jun 2026MiniMax M3
MiniMax
Jun 2026LongCat-2.0
Meituan
Jun 2026GLM-5.2
Z.ai (Zhipu AI)

Who Is Winning at What?

Category leaders pulled live from the benchmark and pricing databases. Scores update daily.

Best at Coding
Claude Opus 5
Anthropic · SWE-bench
97.0 / 100
Best at Reasoning
GPT-6 Astra
OpenAI · GPQA Diamond
96.0 / 100
Best at Math
GPT-5.5
OpenAI · MATH benchmark
95.8 / 100
Best Long Context
Llama 4 Scout
Meta · Context window
10M tokens
Best Value
Nemotron 3.5 Lightning
NVIDIA · Lowest blended price
$0.08 in / $0.20 out per 1M
Best Open Source
DeepSeek V4 Pro
DeepSeek · Top open weight
Open weights available

Incoming Models

What the rumor mill says is coming next. Treat everything here as unofficial unless linked directly to a lab announcement.

Gemini 3.5 ProGoogle
Not yet released

Gemini 3.1 Pro is still labeled Preview on the Gemini API pricing page, and no Gemini 3.5 Pro is listed. Google has shipped Flash models instead, with Gemini 3.6, 3.7, and 3.8 Flash landing between July and September 2026.

Larger DeepSeek V4.1 modelsDeepSeek
Signaled

DeepSeek describes V4.1 Flash, released September 10, 2026, as the smallest model in its new architecture family, which points to larger siblings. No names or dates have been announced.

Claude Mythos 5.1 (wider access)Anthropic
Limited availability

Mythos 5.1 shipped in September 2026 at the same $10/$50 list price as Claude Fable 5.1, but Anthropic lists it as limited availability through Project Glasswing. Most developers can buy Fable 5.1 today and not Mythos 5.1.

Claude Sonnet 5.5 and Haiku 5.5Anthropic
Announced

Anthropic said on September 22, 2026, alongside the Claude Opus 5.5 launch, that Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks. No dates or prices have been published.

Qwen 4Alibaba
In training

Alibaba announced Qwen 4 in Max, Flash, Plus, and 27B versions as in training at its Apsara conference on September 22, 2026. Nothing has been released yet.

Provider Spotlights

Each frontier lab is running a different strategy. Understanding the strategic posture helps predict the next move.

Anthropic

Anthropic has leaned hard into agentic coding as the wedge. Claude has become the default model for serious developer tools, which creates a revenue base that funds frontier training. The company pairs this with the most aggressive public stance on safety of any major lab. Responsible scaling commitments, detailed model cards, and constitutional AI research are all part of a single story: if you believe the most capable models are coming soon, your commercial strategy should be inseparable from your safety strategy. Claude Opus 5.5, released September 22, 2026 at $4/$20, is now the default model in Claude Code and first on both the Vals Index and the Artificial Analysis Intelligence Index, and Anthropic shipped it with the same class of safeguards as Fable 5.1.

OpenAI

OpenAI still owns the largest consumer surface area in AI. ChatGPT is the default chatbot for a huge fraction of the market, which creates data, revenue, and distribution. The o-series reasoning models were the first public bet on test time compute as a primary capability lever, and the results shifted how every other lab thinks about reasoning. Expect continued emphasis on multimodal, voice, and vertical integration through Codex and ChatGPT Work, with GPT-6 Astra now available on paid ChatGPT plans. On September 22, 2026 OpenAI filled out the GPT-6 generation with Sol at $2/$10 and Luna at $0.10/$0.50, and said Astra remains its best model across the board.

Google DeepMind

Google has structural advantages no one else has. Custom TPU infrastructure, a search index, YouTube, and decades of research depth at DeepMind. Gemini has closed most of the capability gap and owns the long context and multimodal categories. The real leverage is distribution: Gemini is shipping into every Google surface, from Search to Workspace to Android. Every Google user becomes a Gemini user by default.

xAI

xAI has compressed an enormous amount of capability into a short timeline. The Memphis training cluster came online with unusual speed, Grok 3 closed most of the gap to the frontier, and the Grok 4 line has stayed close since: Grok 4.6 posts 94.9 on GPQA Diamond in Artificial Analysis testing. Grok 4.7 moved to a new, larger base model on September 21, 2026 at the same $2/$6 rate and scores 46 on the Artificial Analysis Intelligence Index, twelve points behind the leader. Distribution through X gives xAI a feedback loop that other labs do not have. The open question is how long the pace can be sustained.

Meta

Meta did more than anyone to put open weight models near the frontier. The Llama family forced every other lab to compete on value, not just capability, and Mark Zuckerberg framed open weights as a strategic asset, not a charity move. The strategy has since split: Meta now sells its Muse Spark models through a paid API without releasing their weights, while shipping smaller open models such as Muse Glimmer 30B under Apache 2.0.

Latest Model Wars News

Live stream of model release and frontier capability coverage. Filtered from the full TensorFeed news feed.

Frequently Asked Questions

Which AI model is the best in 2026?

There is no single best model. On the independent aggregate indexes, Claude Opus 5.5, released September 22, 2026, is first on both the Vals Index (69.69 percent) and the Artificial Analysis Intelligence Index (58). In our benchmark data, Claude Opus 5 (97.0) and GPT-5.6 Sol (96.2) top SWE-bench Verified, Claude Opus 5 leads OSWorld 2.0 computer use (70.6), GPT-6 Astra leads GPQA Diamond (96.0) and BrowseComp (91.5), and Claude Opus 5.5 leads Humanity's Last Exam with tools (67.7) and Cognition's FrontierCode (54.6). On price, Gemini 3.8 Flash posts 95.3 on GPQA Diamond at $0.75/$3.75 per 1M tokens. The right answer depends on the workload. See our benchmark leaderboard for category winners.

Who is winning the frontier AI race?

As of September 2026, Anthropic and OpenAI trade the top benchmark spots, with Claude Opus 5.5, Opus 5, and Fable 5.1 on one side and GPT-6 Astra on the other. Claude Opus 5.5, released September 22 at $4/$20, now leads both the Vals Index and the Artificial Analysis Intelligence Index, while OpenAI answered the same day with GPT-6 Sol and Luna at half the price or less of the GPT-5.6 models they follow. Google competes hardest on price with its Flash line, and xAI keeps pace with Grok 4.7. The open weight race is led by Chinese labs, with Moonshot's Kimi K3, DeepSeek V4.1 Flash, Alibaba's Qwen3.8 family, and Xiaomi's MiMo-V2.6, whose Pro model posts the top open-weights score on the Artificial Analysis index at 46. DeepSeek has continued to punch above its weight on efficiency.

How often do new frontier models ship?

Frontier releases now land every four to eight weeks on average. Anthropic, OpenAI, and Google each ship a major update every quarter, interspersed with point releases. Open weight launches from Meta, Mistral, and DeepSeek add to the cadence.

What is a frontier model?

A frontier model is one of the most capable AI systems publicly available at a given moment, typically defined by benchmark performance, compute used during training, and agentic capability. The frontier moves constantly as new releases ship.

Which frontier model is cheapest?

Among the big three labs, GPT-6 Luna ($0.10/$0.50 per 1M tokens), Gemini 3.8 Flash ($0.75/$3.75 through December 31, 2026), and Claude Haiku 4.5 ($1/$5) offer the lowest prices per token. DeepSeek V4.1 Flash costs $0.30/$1.20 at peak on its API, Xiaomi's MIT-licensed MiMo-V2.6-Flash is $0.14/$0.28, and open weights like DeepSeek V4.1 Flash and Qwen3.8-27B are close to free when you self host.

Related Hubs

Keep tracking the frontier from every angle.