AI Timeline
A chronological record of major AI model releases, industry milestones, and pivotal events from 2024 to present.
The pace of AI advancement is accelerating. Every month brings new model releases, price drops, capability breakthroughs, and industry consolidation. Tracking this timeline helps you understand the trajectory: where we came from, where we are now, and what's coming next.
From April 2024 to April 2026, we've seen Claude evolve from Opus to Sonnet to a new generation. OpenAI released GPT-4o with vision and function calling. Google launched Gemini 2.5 with 2 million token context. Mistral and Llama launched open-source models that changed the economics of AI deployment. Pricing dropped dramatically as providers competed for market share. Policy shifted too, with regulation discussions in the EU and debates over AI safety standards.
What patterns are visible in this data? Model releases cluster around major announcements. Pricing announcements typically decrease (rarely increase). Open-source releases create competitive pressure on commercial models. Major companies iterate quarterly. Filter this timeline by category to focus on what matters most to you: model releases, pricing changes, policy shifts, acquisitions, or research breakthroughs.
Anthropic Ships Claude Opus 5.5 at $4 and $20 and Takes First on the Vals Index and Artificial Analysis
AnthropicAnthropic released Claude Opus 5.5 on September 22, 2026 as claude-opus-5-5, generally available the same day on the Claude API, Amazon Bedrock, Google Cloud, and Microsoft Foundry. It is priced at $4 per million input tokens and $20 output, 20 percent under Claude Opus 5, and cache reads drop to $0.20 on a 0.05x multiplier where other Claude models use 0.1x, a 60 percent cut on the line that dominates long agent sessions. Context is 1 million tokens with no long-context surcharge and 128K output. Anthropic says it performs at the level of Fable 5.1 on most work while costing about 40 percent less to run than Opus 5, and the independent boards moved accordingly: first on the Vals Index at 69.69 percent, ahead of Fable 5.1 at 68.83 and GPT-6 Astra at 66.61, first on the Artificial Analysis Intelligence Index at 58 against 53 for Astra and Fable 5.1, and first on Cognition's FrontierCode board at 54.6 percent. Anthropic's own table reports 66.4 percent on Terminal-Bench 4.0 at xhigh effort, which Vals measured at 61.62. Four API changes break Opus 5 code: thinking cannot be disabled, default effort drops to medium, forced tool choice returns a 400, and computer use needs the new computer toolset. Claude Code now defaults to it, Opus 5 stays available and is not deprecated, and Anthropic says Sonnet 5.5 and Haiku 5.5 will follow in the coming weeks.
Model ReleaseOpenAI Fills Out GPT-6 With Sol at $2 and $10 and Luna at 10 Cents, and Skips Terra
OpenAIOpenAI released GPT-6 Sol and GPT-6 Luna on September 22, 2026, the same day Anthropic shipped Claude Opus 5.5, completing the GPT-6 generation below Astra. GPT-6 Sol costs $2 per million input tokens, $0.20 cached, and $10 output, half of GPT-5.6 Sol's promotional $4 and $20, and GPT-6 Luna costs $0.10, $0.01, and $0.50, down from $0.20 and $1.20 on GPT-5.6 Luna. An OpenAI spokesperson confirmed the GPT-6 rates are permanent rather than promotional. Both carry a 1,050,000 token context with 128K output, text and image input, and a reasoning effort ladder from none through max, and above 272,000 input tokens the whole request bills at 2x input and 1.5x output. There is no GPT-6 Terra, and OpenAI says Astra remains its best model across the board. The independent read: Cognition measures FrontierCode v1.1 at 49.3 percent for Sol and 42.4 for Luna, the Vals Index places Sol eighth at 62.57 percent and Luna twentieth at 58.45, and Artificial Analysis scores them 48 and 37. Sol now sits at exactly the $2 and $10 of Claude Sonnet 5, which is the price war worth watching. GPT-5.6 Sol is unchanged at $4 and $20 through at least November 21.
Model ReleasexAI Ships Grok 4.7 on a Larger Base Model and Leaves the Price at $2 and $6
xAIxAI released Grok 4.7 on September 21, 2026 as grok-4.7, calling it its most capable model for coding and knowledge work, and the headline is what did not change: $2 per million input tokens, $0.50 cached, and $6 output below 200,000 prompt tokens, doubling to $4, $1, and $12 at or above that threshold for every token in the request, the same rate card as Grok 4.6. The difference is underneath. Grok 4.6 was a post-training upgrade on the 1.5 trillion parameter V9 foundation; 4.7 uses a new, larger base model, reported at 2.1 trillion parameters, trained with a longer reinforcement learning run weighted toward multi-hour tasks. Context stays at 500,000 tokens with text and image input and reasoning effort up to xhigh. xAI's table shows CursorBench 4.0 at 46.3 against 40.4 for Grok 4.6 and Terminal-Bench 4.0 at 38.0, where Fable 5.1 reaches 57.9. The independent read is more modest, 46 on the Artificial Analysis Intelligence Index and tenth on the Vals Index at 60.22 percent, which makes it a strong price-performance option rather than a frontier leader. A fast variant runs at twice the output speed and twice the price.
Model ReleaseXiaomi Open-Sources MiMo-V2.6 Pro and Flash Under MIT, and Pro Takes the Top Open-Weights Score
XiaomiXiaomi released the MiMo-V2.6 series on September 21, 2026 with MIT-licensed weights on Hugging Face and ModelScope: MiMo-V2.6-Pro, a 1.02 trillion parameter mixture-of-experts with 42 billion active, and MiMo-V2.6-Flash at 309 billion with 15 billion active. Both take text, image, video, and audio input natively with a 1 million token context. API prices held at V2.5 levels, $0.435 and $0.87 per million tokens for Pro and $0.14 and $0.28 for Flash, with a Pro-UltraSpeed tier at ten times the Pro rate. Artificial Analysis scores Pro at 46 on its Intelligence Index, its highest open-weights result, and the Vals Index puts Flash at 59.58 percent and Pro at 59.47, both ahead of DeepSeek V4.1 Flash. Xiaomi's own table reports 71.9 on DeepSWE v1.1 and 89.9 on Terminal-Bench 2.1 for Pro, with Terminal-Bench 4.0 at 34.9 marking the distance to the closed frontier. The release also ships the technical report, training environments, and reinforcement learning code, and Xiaomi says the Pro RL run cost about $2.62 million over under six days.
Open SourceAlibaba Ships Qwen3.8-Omni-Flash and Cuts the Price of an Hour of Audio by 98 Percent
AlibabaAlibaba released Qwen3.8-Omni-Flash on September 18, 2026, its first omni-modal model built around agentic capability rather than chat. It accepts text, image, audio, and video on one OpenAI-compatible endpoint, is built natively on the Qwen3.8-Flash foundation instead of bolting perception encoders onto a text model, and carries a 1 million token context window with up to 131K output tokens. QwenCloud lists it at $0.15 per million input tokens and $0.47 output with implicit cache reads at $0.016, which is the same rate card Qwen3.8-Flash already carries. The rate card is not where the change lands. Measured per hour of media, Alibaba says audio input costs 98 percent less than on Qwen3.5-Omni-Plus and combined audio and video more than 93 percent less, which moves meeting archives, call logs, and long-video corpora from pilot budgets into standing jobs. It handles up to an hour of continuous audio or audio-video per call and recognizes speech in 74 languages, and Alibaba reports an average gain above 25 percent across 29 evaluations against Qwen3.5-Omni-Plus, all vendor-reported. One limitation is worth reading before designing around the name: despite a predecessor that generated speech, this model outputs text only, so anything voice-facing still needs a separate synthesis stage. Weights have not been published. Available through Qwen Chat, QwenCloud, and the Model Studio API.
Model ReleaseZ.ai Prices GLM-5.3-FlashX at 2.5x the Base Model for Speed and Nothing Else
Z.aiZ.ai added GLM-5.3-FlashX to its price list on September 18, 2026 at $0.37 per million input tokens, $0.075 cached, and $1.25 output, with cache storage free for a limited time. That is roughly two and a half times what GLM-5.3 Flash charges at $0.15 and $0.50, and the unusual part is what the premium does not buy. FlashX serves the same weights: the same 320 billion parameter mixture-of-experts activating about 18 billion per token, the same hybrid sparse and linear attention, the same 1 million token context with up to 131,072 output, and the same text, image, video, and file input. What changes is the serving stack. Z.ai advertises up to 200 output tokens per second, and independent measurement against its API has landed nearer 98, so the vendor figure is a peak rather than a throughput you should budget. This makes the choice unusually clean, because capability is not a variable: if a person is waiting on the stream, the multiple is defensible, and if a batch runs overnight, it is not. FlashX is a hosted tier with no separate weight release, so the MIT-licensed checkpoint on Hugging Face remains GLM-5.3-Flash, and self-hosting it gives you the model without the stack FlashX is actually selling.
Model ReleaseGoogle Limits Gemini 2.5 to Existing Users Without Deprecating It
GoogleGoogle changed access to its Gemini 2.5 models on September 18, 2026, writing in the Gemini API release notes that to ensure reliable performance for everyone it is limiting access to the 2.5 models to users who have actively used them in the past. Google says the 2.5 models are not deprecated and will continue to be served until further notice, and it points new projects at Gemini 3.5 Flash-Lite or Gemini 3.8 Flash. The practical effect is that Gemini 2.5 Pro, still priced at $1.25 and $10 per million tokens, is no longer something a new team can adopt, even though nothing has a shutdown date. It is a quieter form of retirement than a deprecation notice, and it leaves the newest Pro-tier model on Google's price list as Gemini 3.1 Pro Preview at $2 and $12.
Model ReleaseUnion Alpha Turns Out to Be Pareto, and Unbiased Starts Charging $2.50 and $7.50
UnbiasedOn September 16, 2026 a listing called Union Alpha appeared on OpenRouter under the provider name Stealth: free, multimodal, 262K context, no lab attached. Within a day it was taking enough traffic to degrade its own latency, capacity was tripled overnight with AWS help, and it still ran short. At 23:24 UTC on September 17 the stealth listing lost its endpoints and the model went live under its real name, Pareto 26.9 from Unbiased, the platform built by Circuit & Chisel, at $2.50 per million input tokens, $0.25 cached, and $7.50 output. The architecture is the reason it is worth a timeline entry. Pareto is not a single model: Unbiased describes it as a blended model in which several frontier and open-source models run against each request and the best answer is kept, behind one model string and one bill, and it does not switch models mid-conversation so prompt caching keeps working. That is a structural hedge against any one vendor moving price or pulling access, and it is also why the normal questions have no answer here, because Unbiased publishes no parameter count, no architecture, and no list of what is in the pool. Specs are 262,144 tokens of context with up to 131,072 output, text and image input, text output, and tool calling with structured output. On its own table Pareto scores 88 on ArXivMath, 78 on MMMU-Pro, and ties GPT-6 Astra and DeepSeek V4.1 Flash at 74 on DeepSWE; no independent evaluator has published a score and Artificial Analysis has no entry, so the vendor table is the only table. Access runs through the Unbiased platform, OpenRouter, Cloudflare AI Gateway, Kilo Gateway, and NanoGPT at the same rate.
Model ReleaseAnthropic Says Claude Now Leads 26 Percent of Its Own AI R&D, Up From Under 1 Percent in February
AnthropicAnthropic published figures on September 17, 2026 quantifying how much of its own model research and development Claude runs. As of August 2026, Claude led roughly 26 percent of AI R&D tasks, a category Anthropic distinguishes from collaboration, against under 1 percent in February 2026. More than 90 percent of its AI R&D now has Claude involved at least as a collaborator, and about 30,000 agents run concurrently on its most-used internal agent platform. The oversight numbers are the ones worth keeping: across more than one billion agent decisions in August, the internal monitor blocked about one in 47,000. Anthropic says roughly 6 percent of total AI R&D compute went to safety research, rising to about 12 percent when counting only work where AI is doing the AI R&D. This is self-reported and self-defined, and the definition of leading is doing a lot of work here, but it puts a number on how much of a frontier lab's own research its model now drives, which few labs have published at all. For anyone forecasting release cadence, it is a better leading indicator than headcount.
ResearchOpenAI Ships Astra for Law, a GPT-6 Astra Configuration With a 230 Million URL Case Law Index
OpenAIOpenAI released Astra for Law on September 17, 2026, GPT-6 Astra wrapped in a legal search index and a set of legal-analysis instructions, sold to law firms and to the software companies that sell to them. The index covers US case law, statutes, regulations, court rules, and administrative decisions across more than 230 million URLs, with sources added daily. On 200 questions drawn from the private validation set of Vals AI's Legal Research Bench, Astra for Law passed the overall correctness check on 54 percent, against 38.7 percent for stock GPT-6 Astra with web search. That is a real lift and still a coin flip, which is the number to quote to anyone treating this as a research substitute rather than a research assistant. Harvey and Legora can build on it, and it integrates with Relativity and Clio. Latham and Watkins, Ropes and Gray, Cooley, and Sullivan and Cromwell are named as early adopters. This is a configuration and a retrieval layer, not a new model, so the underlying GPT-6 Astra rate card still applies.
Model ReleasePrismML Compresses Qwen3.8 27B to 5.93 GB and Keeps 98.2 Percent of It
PrismMLPrismML released Ternary Bonsai 2 27B on September 17, 2026, a ternary-weight build of Qwen3.8 27B under Apache 2.0. Every weight is -1, 0, or +1, with one FP16 scale per group of 128, which lands at 1.76 effective bits per weight in the dense PTQ1_0 packing and a 5.93 GB file against 53.80 GB in FP16, about 9.1 times smaller. Only 26.2 million parameters, 0.0976 percent, stay in higher precision. Weights are stored in a rotated basis using a blockwise Hadamard transform with block size 1,024, citing SpinQuant. Context is 262K, input is text and images, and it runs today on a 16 GB laptop or a single 24 GB GPU, though stock llama.cpp rejects the file types: PrismML's fork or its MLX runtime is required. The retention headline is 98.2 percent of the parent's average across 20 benchmarks, 83.9 against 85.4, and the category breakdown is uneven: math 99.5 percent, coding 99.3 percent, instruction following 101.7 percent, but knowledge and reasoning 96.9 and vision 96.3. The honest caveat is in the two benchmarks that sit outside that 20-benchmark average. Terminal-Bench 2.1 drops to 52.8 from 69.7 and SWE-bench Verified to 60.8 from 80.6, roughly 75 percent retention, so long-horizon agent work is where the compression actually costs you. Against conventional quantization the gap is wide: an IQ2_XXS build of the same model averages 75.2 at 7.3 GB. Throughput is 142.5 tokens per second on an RTX 5090 and 46.8 on an M5 Max at batch size 1. All results are PrismML's own and have not been independently reproduced.
Open SourceTypeSafe AI Ships Jev, a Model That Returns Typed Decisions Instead of Text at $0.042 per Million In
TypeSafe AITypeSafe AI announced Jev on September 15, 2026 and called it the first of a new class it names System One models. Instead of generating tokens one at a time, Jev takes unstructured program state plus a set of typed questions and answers all of them in a single parallel pass, returning typed decisions with calibrated probabilities rather than a string. TypeSafe claims 70 to 500 milliseconds per call, 40 to 200 times faster than frontier LLMs on comparable tasks, and says the output shape makes hallucination and type errors structurally impossible rather than merely unlikely. The training method is its own, called Reinforcement Learning for Calibrated Decisions. The OpenRouter listing (typesafe/jev-1.13, listed September 18) shows $0.042 per million input tokens with output free and a 32,000 token context, under a decisions output modality. Access is early-access by waitlist and there is no independent benchmark yet. TensorFeed is keeping this out of the pricing tables for now: a model that does not emit text does not belong in a rate-card comparison against models that do, and a $0.042 row with free output would top every cheapest-model sort on the site for the wrong reason. It sits on the timeline until either the category or the gating settles.
Model ReleaseZ.ai Raises About $5 Billion in Hong Kong Five Days After a US Advisory Named It
Z.aiZ.ai, the Hong Kong-listed lab formerly known as Zhipu AI, disclosed on September 13, 2026 that it is raising roughly $5 billion in two pieces, according to its Hong Kong exchange filing and Reuters. The first is a placement of up to 21,965,000 new H shares at HK$714, a 9.96 percent discount to the September 11 close, for gross proceeds of about HK$15.68 billion, roughly $2 billion. The second is RMB 20.14 billion of US dollar settled zero-coupon convertible bonds due 2027, about $3 billion. The timing is the story. Five days earlier a joint CISA, NSA, and FBI advisory named Z.ai among six China-based labs accused of industrial-scale distillation of US frontier models, and investors still absorbed a multibillion-dollar raise at a single-digit discount. For buyers of GLM-5.3 and GLM-5.3 Flash, the practical read is that the lab behind one of the cheapest frontier-adjacent coding rate cards on this site is not short of capital, whatever Washington thinks of how it trained.
Sakana Ships Fugu Max at $2 and $6 and Fugu Ultra v2 Without Fable or Astra in the Pool
Sakana AISakana AI released Fugu Max v1.0 and Fugu Ultra v2.0 on September 11, 2026, both orchestrators rather than models: a request goes to one API and Fugu decides which models in its pool do the work. Fugu Max is $2 per million input tokens and $6 output with $0.25 cached input at any context length, which Sakana says undercuts Sonnet 5, GPT-5.6 Terra, and Kimi K3 output rates by 40 to 60 percent, and it reports best overall score on six benchmarks including Terminal Bench 2.1, GPQAD, AA-LCR, GDP.pdf, AutomationBench, and the internal SWEFish. Fugu Ultra v2 sits above it at $5 and $30 with $0.50 cached input, rising to $10, $45, and $1.00 once context passes 272K tokens, and reports best or joint-best on five of eight benchmarks and top two on seven of eight, with Chartography at 48.3 against Opus 5 at 27.3 and Fable 5 at 29.5, and DeepSWE at 74.3. The detail that carries the most weight is a footnote on Sakana's own chart: Fable 5, Fable 5.1, and GPT-6 Astra are not in Fugu Ultra v2's pool, and its training cutoff is August 28, 2026, so the scores come from orchestrating open and specialized models rather than from renting the frontier. Context is 1 million tokens on both. Billing is the thing to read closely. On Fugu Ultra, orchestration tokens are returned inside token_details but bill at full input and output rates, so the invoice covers work the user never sees; on Fugu Max, web_search and web_fetch bill at $0.007 per call and one query can take several. Fugu is not yet offered in the EU or EEA. All benchmark figures are self-reported, and SWEFish is Sakana's own.
Model ReleaseDeepSeek Cancels the V4 Pro Retirement Three Days Before It Was Due
DeepSeekDeepSeek announced on September 11, 2026 that it will keep providing API service for DeepSeek V4 Pro after September 14, 2026, with the billing method unchanged, reversing the phase-out it had published the day before alongside V4.1 Flash. The original plan was to route every deepseek-v4-pro request to V4.1 Flash from September 14 and bill it at Flash rates. DeepSeek says it changed course in response to user demand, and its own changelog now carries the continuation notice. The whole cycle, announced, scheduled, objected to, and withdrawn, ran in five days, and the practical consequence is that there is nothing to migrate: a call naming V4 Pro is still answered by V4 Pro at V4 Pro rates. Those rates are the part worth keeping in view, because they are not close to the Flash tier. V4 Pro bills $1.32 per million cache-miss input tokens and $3.96 output at peak, halved off-peak to $0.66 and $1.98, with cache hits at $0.044 and $0.022, against $0.30 and $1.20 peak on V4.1 Flash. DeepSeek's own table has V4.1 Flash ahead on agentic work, so the narrow case for staying is knowledge and science reasoning, where V4 Pro still leads GPQA Diamond 92.4 to 90.9 and Humanity's Last Exam without tools. It is a useful reminder for anyone planning around a published sunset date: a retirement notice is a statement of intent, not a fact about the future.
Model ReleaseMoonshot Rolls Out Kimi K2.8 Preview With 1M Context on Every Tier and the Same API Id as K2.7
Moonshot AIMoonshot AI rolled out Kimi K2.8 Preview on September 11, 2026 across Kimi Code and Kimi Work, a mid-tier model sitting between the coding-focused Kimi K2.7 Code and the K3 flagship. It adds text, image, and video input, lifts every Kimi Code membership tier to a 1 million token context from the 262,144 on K2.7 Code, and brings the three thinking effort levels K3 introduced (low, high, max), running at max by default. Moonshot describes the coding and agent gains as close to K3 and says thinking efficiency improved substantially over K2.7 Code. The operational detail that matters is that K2.8 Preview keeps the same API model id as K2.7 Code, so existing integrations and third-party tools picked up a different model with no configuration change and no version pin to hold them back. Moonshot has not published a first-party rate card for it; the third-party gateway TokenRa lists roughly $0.80 per million input, $3.35 output, and $0.14 cached read, which is why TensorFeed is not carrying a pricing row for it yet.
Model ReleaseShanghai AI Lab Drops Atria Dawn Preview, a 744B MIT-Licensed Agentic MoE, With the Paper Three Days Later
Shanghai AI LaboratoryShanghai AI Laboratory published Atria Dawn Preview on September 11, 2026 as MIT-licensed FP8 weights on Hugging Face and ModelScope with no blog post and no paper, a 744 billion parameter mixture-of-experts agentic model built on the GLM-5.2 foundation with a 256K context. The technical report followed on arXiv on September 14 with more than 140 authors, and it is the first full account of what the model is and how it was trained. The target is long-horizon research agents: scenarios that need continuous environmental understanding, tool use, and multistep completion, with the stated goal of carrying scientific work from a method in the literature through to executable experiments, reproducible metrics, and a report someone else can inspect. Documentation and evaluation results are at atria-asi.ai. MIT weights on a 744B agentic model is a notable license at that scale, and the weights-first, paper-later sequencing is becoming the Chinese lab default.
Open SourceDeepSeek Ships V4.1 Flash at 552B and Retires V4 Flash
DeepSeekDeepSeek released V4.1 Flash on September 10, 2026, a 552 billion parameter mixture-of-experts trained from scratch on 45 trillion multimodal tokens, built on an asymmetric Causal Encoder-Decoder that activates about 8 billion parameters per token on input and 16 billion on output, with a separate 196 billion parameter Engram memory table. Context is 1,048,576 tokens with 384K output and native image input, and the weights are MIT licensed on Hugging Face. The release reorganized the whole API: V4 Flash and V4 Flash Vision Exp were retired on launch day with their model names routed to V4.1 Flash, and DeepSeek said that from September 14 every deepseek-v4-pro request would follow. It withdrew that second step on September 11 after user objections, so V4 Pro remains on the API at its own rates. DeepSeek argued for the move with a self-reported table that favors V4.1 Flash on agentic work, Terminal-Bench 2.1 at 90.6 against V4 Pro at 87.9, DeepSWE v1.1 at 74.2 against 62.7, CyberGym at 88.1 against 83.3, NL2Repo-Bench at 64.0 against 61.5, and a 3,471 Codeforces rating, while V4 Pro still leads GPQA Diamond at 92.4 against 90.9. Pricing splits by the clock: $0.30 per million input tokens and $1.20 output at peak, halved to $0.15 and $0.60 off-peak, with cache hits at $0.006 and $0.003, and peak defined as Monday through Friday, 01:00 to 04:00 and 06:00 to 10:00 UTC. Even at peak that undercuts the $0.44 and $1.32 V4 Flash charged. The cache-hit rate is the headline most coverage led with, and it is real, but it only applies to tokens the model has already seen.
Model ReleaseMistral Raises €3 Billion at a Valuation Above €21 Billion in a Samsung-Led Series D
MistralMistral announced on September 8, 2026 a €3 billion Series D at a post-money valuation of more than €21 billion, which it calls the largest equity fundraising round ever completed by a European technology company. Samsung Electronics led, with Scaleup Europe Fund (managed by EQT) and existing investor PSG Equity as co-leads. Advent, funds managed by BlackRock, and the Grand Duchy of Luxembourg joined as new investors, and existing shareholders including NVIDIA, ASML, a16z, Lightspeed, and General Catalyst participated. The industrial names matter more than the number: a Series C led by ASML and a Series D led by Samsung put two of the most important non-US companies in the chip supply chain on Mistral's cap table, and Samsung separately announced a strategic partnership with Mistral on AI for semiconductor infrastructure. Mistral says the money goes to frontier research, compute capacity, and infrastructure for a full stack built on open weights, and the lineup already points that way: Mistral Large 3 and Mistral Small 4 ship under Apache 2.0, and Medium 3.5 is its frontier-class agentic and coding model.
CISA, NSA, and FBI Accuse Six Chinese AI Labs of Industrial-Scale Distillation of US Models
US GovernmentCISA, the NSA, and the FBI released a joint cybersecurity advisory on September 8, 2026 accusing six China-based AI companies, DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI, of extracting billions of tokens across millions of requests from US frontier models, including variants of Claude, GPT, Gemini, and Grok, since at least late 2024 and likely with Chinese government awareness. The agencies describe the campaigns as the core of China's AI development strategy and say they violate US providers' terms of service. It is guidance, not a rule. The advisory asks US labs to detect anomalous accounts and usage, such as immediate maximum usage from new accounts, to subtly alter responses to suspected distillation traffic, and to share intelligence across model providers, cloud platforms, and API aggregators. That last recommendation has the most reach, because aggregators sit between most buyers and most models. Five of the six named labs have models in the TensorFeed catalog, so expect the advisory to surface in procurement reviews well before it shows up in regulation.
PolicyNvidia Confirms a $12.93 Billion Deal for Hugging Face and Owns the Shelf Its Chips Stock
NvidiaNvidia confirmed on September 3, 2026 that it will acquire Hugging Face for roughly $12.93 billion, its largest acquisition on record, with about $11.9 billion going to shareholders and up to $1 billion held in equity to retain staff joining Nvidia. CNBC first reported the agreement on August 27 and Nvidia filed an 8-K dated September 2. The transaction is expected to close in the first half of 2027 subject to regulatory approval, so nothing about the platform changes yet. Hugging Face was founded in 2016 by Clement Delangue, Julien Chaumond, and Thomas Wolf, and Nvidia puts the platform at more than 18 million developers sharing over 3 million models, 500,000 datasets, and 1 million applications. Jensen Huang says Hugging Face will continue supporting open source and open-weight models. The structural point is that Nvidia now owns both the silicon that trains models and the default registry where they are discovered and pulled, which is a distribution chokepoint rather than a capability purchase, and it is the part regulators are most likely to examine. Every open-weight release TensorFeed tracks, including this week's MIT-licensed DeepSeek V4.1 Flash weights, ships through a repository that is now committed to a chip vendor.
AcquisitionOpenAI Ships GPT-6 Astra at $10 and $50 With a 1.05M Context and a 272K Billing Cliff
OpenAIOpenAI released GPT-6 Astra on September 3, 2026, first as a limited preview for trusted partners and then to paid users the following day in a restricted build that refuses certain cybersecurity prompts. The API id is gpt-6-astra at $10 per million input tokens and $50 output, with cached input at $1, a $12.50 cache write, batch at half price, and a Fast mode at 2x. That is two and a half times the $4 and $20 that GPT-5.6 Sol has carried since the August 21 cut, so Astra is a price increase at the flagship, not a continuation. The context window is 1,050,000 tokens with 128K max output, text and image input, and an April 30, 2026 knowledge cutoff, but prompts above 272,000 input tokens bill at 2x input and cache rates and 1.5x output, which makes the headline window and the affordable window two different numbers. OpenAI reports GPQA Diamond at 96.0, FrontierMath Tier 4 at 97.6, ARC-AGI-3 at 99.9 under its own provider adapter harness, 100 percent on ExploitBench, and MRCR v2 8-needle at 100 percent in the 256K to 512K band and 96.3 in the 512K to 1M band against Sol's 91.5 and 73.8. SRE-Bench incident response is 88.0 solved on the first attempt against 55.9. No SWE-bench Pro figure was published. All self-reported.
Model ReleaseMeta Ships Muse Spark 1.3 With Fewer Tool Calls and an Open-Weight Release Still Only Promised
MetaMeta made Muse Spark 1.3 available in Muse Code and the Meta Model API on September 2, 2026, holding Standard pricing at $1.25 per million input tokens, $4.25 output, and $0.15 cached input, identical to Muse Spark 1.2. The Contributor tier also holds at $0.10 and $0.20 in exchange for letting Meta train on submitted prompts and completions. The pitch is efficiency rather than raw score: roughly 20 percent fewer tool calls and 25 percent fewer tokens than 1.2 on the same agentic work, which lands on the bill rather than on a leaderboard. Meta reports DeepSWE 1.1 at 75.4, SWEAtlas CodeBase QnA at 59.4, Terminal-Bench tied with GPT-5.6 Sol at 88.8, and MRCR at 98.5 in the 256K to 512K band and 98.1 from 512K to 1M against Sol's 91.5 and 73.8. One caveat travels with those numbers: the 1.3 column runs at max reasoning while the 1.2 column runs at xhigh, so part of the version-over-version jump is a tier gap. Context is 1 million tokens. Despite the Muse Spark name carrying open-weight expectations, 1.3 is not open weights; Meta says a Muse Spark open-weight release is coming with no date, variant, or license named.
Model ReleaseGoogle Releases Gemini 3.8 Flash and Prints the Date Its Introductory Price Ends
GoogleGoogle released Gemini 3.8 Flash on September 2, 2026 as gemini-3.8-flash, holding the $0.75 per million input and $3.75 output that Gemini 3.7 Flash introduced, but with an expiry now stated outright: the rate runs through December 31, 2026, and on January 1, 2027 both sides double to $1.50 and $7.50. Batch and Flex are half those rates and Priority is 1.8x. Anyone modeling 2027 spend on a Flash-tier workload should budget the post-January number. The model takes text, image, audio, video, and PDF input with a 1,048,576 token context window and 65,536 output tokens. Artificial Analysis measures GPQA Diamond at 95.3 for the high reasoning tier, and Google reports SWE-bench Pro at 61.6 with sharp gains on coding and terminal tasks over 3.7 Flash. A separate Cyber variant shipped alongside it behind Fairwind gating.
Model ReleaseAnthropic Ships Claude Fable 5.1 and Mythos 5.1 and Cuts Cache Reads 75 Percent
AnthropicAnthropic released Claude Fable 5.1 and Claude Mythos 5.1 on September 1, 2026. List pricing does not move: both are $10 per million input tokens and $50 output, matching Fable 5. The change is underneath, where cache reads drop 75 percent to $0.25 per million. Five-minute cache writes stay at $12.50, one-hour writes are $20, and the Batch API halves both sides to $5 and $25. Anthropic puts the practical effect at roughly 25 percent off typical workloads and up to 45 percent off heavily agentic ones, which means the saving is real but entirely dependent on cache hit rate, so a workload that does not reuse context sees no discount at all. Both models carry 1 million token context with 128K output. Fable 5.1 adds a beta for adjusting effort mid-conversation without restarting the session, and Anthropic reports GPQA Diamond at 92.6 and SWE-bench Pro at 81.2, leading with agentic rows rather than SWE-bench Verified, for which no 5.1 figure was published; the 95 percent number circulating on third-party leaderboards is Fable 5 from June. Availability differs sharply between the two. Fable 5.1 is broadly available on the Claude API and cloud platforms, while Mythos 5.1 stays gated behind the Cyber Verification and Life Sciences Verification programs run with the US government, so its published price is a list rate most developers cannot transact at. The Python SDK moves from httpx to httpx2, requires Python 3.10 or later, and drops the legacy Text Completions API.
Model ReleaseTencent Open Sources Hy4 Preview, a 770B Apache 2.0 Model at $0.834 and $2.501
TencentTencent released and open sourced Hy4 preview on August 28, 2026: a mixture-of-experts with 770 billion total parameters and 49 billion active per token, a 1M token context window, and weights under Apache 2.0 on Hugging Face, ModelScope, GitCode, and CNB. API pricing through Tencent Cloud TokenHub is $0.834 per million input tokens, $2.501 output, and $0.042 for cache hits, and the model is also on OpenRouter and inside Tencent products including WorkBuddy, CodeBuddy, and Yuanbao. The architecture openly borrows from DeepSeek and GLM, with gated DeepSeek Sparse Attention and IndexCache, plus a 10 billion parameter multi-token prediction layer for speculative decoding. The evidence is thinner than the spec sheet: Tencent's headline result is an internal blind test in which 163 of its own experts scored Hy4 preview 2.99 out of 4 on 203 engineering tasks, against 2.94 for Kimi K3 and 2.92 for GLM-5.3, and its benchmark appendix is published as images. Tencent also says the model helped optimize its own training and lifted inference throughput 31.8 percent, which it frames as an early recursive self-improvement loop. The preview label is literal, and Tencent says more Hy4 models are coming soon.
Open SourceZ.ai Publishes GLM-5.3 Open Weights After a Two-Week Safety Review
Z.aiZ.ai published the GLM-5.3 weights to Hugging Face at zai-org/GLM-5.3 on August 28, 2026, two weeks after the August 14 API-first launch and exactly on the date the gated repository had been counting down to. Z.ai cited a completed safety review as the reason for the delay and describes the release as its most capable model for agentic coding and cyber defense. The gap between announcement and availability is the part worth noting: for fourteen days GLM-5.3 was an open-weight model in every press description and a closed API in practice, reachable only through the Z.ai API and the GLM Coding Plan at $1.40 per million input tokens and $4.40 output. Read the license before deploying. GLM-5.2 and GLM-5.3 Flash ship under plain MIT, but the flagship weights, roughly 753 billion parameters, carry a custom GLM-5.3 License: permissive for almost everyone, except that any Model-as-a-Service operator whose group revenue tops $10 billion over 12 months must pass a Z.ai security review before commercial use.
Open SourceZ.ai Ships GLM-5.3 Flash, the First Natively Multimodal GLM-5, at $0.15 and $0.50
Z.aiZ.ai released GLM-5.3 Flash on August 26, 2026, the first natively multimodal model in the GLM-5 family and a separate model from the text-only GLM-5.3 flagship. It is a 320 billion parameter mixture-of-experts activating roughly 18 billion per token across 45 layers, with hybrid linear and sparse attention, a 1 million token context window, and text, image, and video input in one stack. Weights are MIT licensed on Hugging Face at zai-org/GLM-5.3-Flash. List pricing is $0.15 per million input tokens, $0.03 cached, and $0.50 output, roughly a tenth of the $1.40 input rate on text GLM-5.3, with a 50 percent launch promotion running through September 9, 2026 at 16:00 UTC. Budget the list rate, not the promo. Z.ai reports DeepSWE at 63.4 against GLM-5.2's 46.2 and AutomationBench at 48.8 against 26.2, with Toolathlon at 78.4 and Terminal Bench 2.1 at 84.3; all self-reported. The clear miss is BabyVision at 53.4 against Gemini 3.7 Flash's 70.9. The model ran on OpenCode and OpenRouter as the stealth id ox-alpha before the launch, and thinking cannot be disabled.
Model ReleaseAlibaba Releases Qwen3.8-Flash-Next Open Weights as a Qwen4 Architecture Preview
AlibabaAlibaba published Qwen3.8-Flash-Next on August 26, 2026, an open-weight experimental preview of the architecture that will underpin Qwen4. It is 125 billion parameters with only 6 billion activated per token, plus a 51 billion parameter n-gram embedding table indexing 20 million bigrams and trigrams that can sit in system RAM rather than GPU memory, and a 4 billion parameter multi-token prediction layer. Native context is 262,144 tokens, extensible to 1 million with YaRN, and it takes text, image, and video in. Alibaba reports better results than Qwen3.7-Plus at roughly one ninth the training cost, with DeepSWE 1.1 at 58.7 and SWE-bench Pro at 62.5 against DeepSeek-V4-Flash-0731's 54.4 and 56.0, GPQA Diamond at 91.7, and LiveCodeBench v6 at 91.9. NL2Repo-Bench is the miss at 48.1 against 54.2. All figures are self-reported on the Hugging Face card. Two things are easy to conflate here: the open weights ship under a qwen-community-1.0 license with no hosted list price, while the managed Qwen3.8-Flash on QwenCloud is a separate product that launched at $0.16 per million input tokens (QwenCloud now lists $0.15) and $0.47 output with a 1 million token default context.
Open SourceOpenAI Cuts GPT-5.6 Sol API Pricing by 20 to 33 Percent Through November 21
OpenAIOpenAI cut developer pricing on its flagship GPT-5.6 Sol on August 21, 2026, three weeks after leaving Sol untouched in the July 30 cut that dropped Luna 80 percent and Terra 20 percent. Input falls from $5 to $4 per million tokens (down 20 percent), output from $30 to $20 (down 33 percent), and cached input from $0.50 to $0.40. The rate is explicitly promotional and runs through at least November 21, 2026. It applies to the pay-as-you-go API, Codex credits, and eligible ChatGPT Work plans; Pro, Plus, and Business subscription prices are unchanged. This is the first time OpenAI has moved the flagship rate since GPT-5.5 set $5/$30 in April, and it lands with Claude Opus 5, Grok 4.6 at $2/$6, and a widening open-weight floor all pressing on the same tier.
PricingNvidia Pays Poolside $6 Billion to License Model Factory and Hire 109 Staff
NvidiaNvidia agreed to pay roughly $6 billion for a non-exclusive license to Poolside's Model Factory, the training platform behind Poolside's Laguna coding model, and will extend offers to 109 Poolside employees who worked on it. Nvidia is separately investing $1 billion in Poolside at a $12 billion pre-money valuation, and Poolside's three co-founders stay with the company. Both sides say it is neither an acquisition nor an acquihire, and the non-exclusive license leaves Poolside free to license the same software elsewhere. The structure is the story: roughly $7 billion of value transferred, the core training infrastructure and the team that built it moved, and no merger filing triggered. Reporting frames Model Factory as running 10,000 to 20,000 training experiments a month with streaming data feeding training directly.
AcquisitionBroadcom Seeks $70 Billion to $80 Billion of Debt for a Second AI XPV Tranche
BroadcomBloomberg reported on August 20, 2026 that Broadcom was in talks with lenders for more than $60 billion of debt to fund custom AI chips for Anthropic and other labs, and CNBC put the target at $70 billion to $80 billion the next day, with a senior secured tranche around $60 billion to $70 billion and a junior tranche near $30 billion. Some reporting stretches the total toward $100 billion. This is the second transaction on the AI XPV platform Broadcom formed with Apollo and Blackstone in June, whose $35 billion opening deal financed roughly a gigawatt of Anthropic capacity. The structure keeps the accelerators off Anthropic's balance sheet: the vehicle borrows, buys the chips, and leases them back on a five-year term. The partnership targets more than 20 gigawatts delivered to frontier labs by 2028. BofA cut Broadcom credit from Overweight to Marketweight on August 11 citing XPV exposure.
MilestoneAnthropic Reports Claude Designed Working Protein Binders for 14 of 15 Targets
AnthropicAnthropic published lab-validated results on August 20, 2026 showing Claude autonomously designing de novo protein binders, with hits against 14 of 15 targets. Across a multi-arm campaign, Mythos Preview hit 26.7 percent and Opus 4.8 hit 22.6 percent over 48-hour runs, rising to 35.1 percent when Mythos Preview worked a single target across 24 hours, against a 10 to 15 percent rate typical of current protein design campaigns. The campaign produced 354 confirmed binders from 1,320 designs. Adaptyv Bio and Twist Bioscience synthesized and tested the designs independently, and on the RBX1 target Mythos Preview reached a 40 percent hit rate against 3.7 percent among human competition entrants. Anthropic is explicit that protein binders are not drugs and that a high-affinity binder is the first step, not the finish. It also says life-science tasks remain restricted in its most capable model and that an access program for scientists is being prepared.
ResearchOpenAI Launches ChatGPT for Teens With an Age Assurance Layer
OpenAIOpenAI introduced ChatGPT for Teens on August 18, 2026, a separate experience for 13 to 17 year olds with tightened content policies and a learning-oriented default, alongside a partnership with CodeAI aimed at classroom use. The product is best read as a compliance layer rather than a feature: age assurance obligations are landing across US state law, the UK Online Safety Act, and the EU, and shipping a distinct minor-facing tier lets OpenAI apply age-gated policy without reshaping the adult product. The open question is verification strength, since a self-declared birthday and an inference-based age estimate carry very different regulatory weight.
PolicyOpenAI Publishes the Cost of Frontier Containment at Roughly 20 Percent of Inference Compute
OpenAIOpenAI published "Pacing model development in an era of cyber-critical capabilities" on August 18, 2026 and disclosed a number no frontier lab had put in public before: monitoring overhead runs at roughly 20 percent of the inference compute being monitored, with wide variation across workloads. The post also confirmed two weeks of reinforcement learning halted on models intended for deployment, with the largest planned frontier RL run still on hold. Shipped controls include activation classifiers at every sampled token escalating to higher-compute automated investigators, a 30 minute target to surface concerning activity plus a further 30 minutes to clear it, mandatory coverage for all RL training and tool-using evaluations at Sol capability or higher, extended since August 7 to all Astra inference with tools, plus workload and network isolation for model-generated code. OpenAI also conceded that most of its current Preparedness Framework text dates to 2023 and is being rewritten. It is the first public unit price on frontier containment.
PolicyStripe Agrees to Acquire OpenRouter for More Than $7 Billion
StripeStripe agreed to acquire OpenRouter, the model gateway that routes requests across more than 400 models from over 80 providers, in a deal Bloomberg first reported on August 16 and Stripe confirmed on August 18, 2026. Reported prices range from roughly $7 billion to $8 billion depending on the source and on how cash and stock components are counted, with the New York Times putting it near $7.5 billion including $1.5 billion to the founders. Official terms are undisclosed. The strategic read is that model routing and payments are converging: the layer that decides which model serves a request is also the layer that meters and bills it, and Stripe is buying the metering point rather than building one. For anyone whose pricing or usage data depends on OpenRouter, this is a dependency worth watching, because the public rankings and cross-provider catalog that a large slice of the ecosystem treats as neutral reference data now sit inside a payments company.
AcquisitionNvidia Backs $105 Billion in Financing for an OpenAI Data Center in Ohio
NvidiaA securities filing on August 17, 2026 disclosed that Nvidia will backstop up to $105 billion in financing for the PORTS-Pike Technology Campus in Pike County, Ohio, built by SoftBank's SB Energy on a former uranium enrichment site and fully leased to OpenAI. Nvidia is named exclusive AI compute infrastructure provider and its commitment covers the first 4.25 gigawatts; the campus is planned to reach 10 gigawatts, which would make it the largest data center in the world. The disclosed number came in roughly $145 billion below earlier reporting, which drew immediate questions about how much of the announced pipeline is firm and how much is a chip vendor financing demand for its own silicon.
MilestoneAnthropic Annualized Revenue Tops $65 Billion Ahead of Its IPO
AnthropicBloomberg reported on August 17, 2026 that Anthropic's annualized revenue run rate passed $65 billion at the end of July, up more than sevenfold from the end of 2025 and ahead of OpenAI's reported $40 billion. Q2 2026 revenue came in above $11.5 billion, roughly 14 times the same quarter last year. Anthropic filed confidentially in June at a $965 billion Series H valuation and is expected to submit a public registration statement as soon as the end of August, with investors targeting an October listing at $2 trillion or more. Backers project $100 billion to $120 billion annualized before the year closes and roughly $190 billion to $200 billion of annual revenue by 2028, which is the number the valuation actually rests on.
MilestoneAlibaba Releases Qwen3.8-27B Open Weights Under Apache 2.0
AlibabaAlibaba published the Qwen3.8-27B weights on Hugging Face and ModelScope on August 14, 2026, eleven days after the closed Qwen3.8-Max flagship. The checkpoint is 27.78 billion dense parameters under Apache 2.0, natively multimodal across text, image, and video, with a 262,144 token context window. Qwen's model card reports Terminal-Bench 2.1 rising from 63.4 to 73.0 against Qwen3.6-27B, DeepSWE 1.1 from 13.3 to 42.2, OSWorld-Verified from 63.9 to 84.3, and SWE-MM from 25.7 to 38.6. Scores are vendor-reported. The practical read is that the strongest locally deployable multimodal model near 30B parameters is now permissively licensed, which pushes the open-weight floor up again.
Open SourceGoogle Ships Gemini 3.7 Flash at Half the Price of 3.6 While Gemini 3.5 Pro Stays Delayed
GoogleGoogle made Gemini 3.7 Flash generally available on August 13, 2026, three weeks after Gemini 3.6 Flash and with Gemini 3.5 Pro still unreleased after slipping past its July 17 and August 6 targets. Introductory pricing is $0.75 per 1M input tokens and $3.75 per 1M output through December 31, 2026, exactly half the 3.6 Flash rate, reverting to $1.50/$7.50 after that. The model carries a 1,048,576 token input window and 65,536 max output. Google reports DeepSWE v1.1 climbing from 48.6 to 65.3, FrontierCode 1.1 Main from 34.4 to 43.6, Terminal-Bench 2.1 at 85.8, and GDM-MRCR v2 at 97.0 percent for 128K retrieval. Google says it did not retrain from scratch and instead rebuilt on algorithmic gains and user feedback. Available in the Gemini API, AI Studio, Android Studio, Antigravity, the Gemini Enterprise Agent Platform, and the Gemini app.
Model ReleaseAnthropic in Talks to Buy Decart for $6 Billion as Databricks Closes $5 Billion Round
AnthropicBloomberg reported on August 13, 2026 that Anthropic is negotiating to acquire Israeli startup Decart for roughly $6 billion, which would be its largest known acquisition. Decart builds chip efficiency software that cuts AI training and inference cost, plus world models for simulated environments. The company raised $300 million in May at close to a $4 billion valuation, so $6 billion is a premium near 50 percent. Talks are not final. The same day Databricks closed a $5 billion Strategic Growth Round with Coatue, Blackstone, MGX, T. Rowe Price, and Sixth Street Growth, converting the term sheet reported in mid July. Both deals point the same direction: capital is flowing to whoever can lower the cost per token rather than raise the ceiling on capability.
AcquisitionOpenAI Splits Daybreak Into Blue and Red Tiers and Ships GPT-5.6-Cyber
OpenAIOpenAI restructured its Daybreak cyber defense program into two access tiers on August 10, 2026. Blue gives approved defenders a de-guardrailed GPT-5.6 Sol for vulnerability discovery, secure code review, malware analysis, incident response, and patch validation. Red gives vetted security teams the purpose-trained cyber models, including the new GPT-5.6-Cyber, built on Sol but tuned to refuse less on dual-use work and to perform better at exploit development and zero-day discovery. OpenAI reports a 95.0 percent completion rate on its internal Advanced Cybersecurity Completion Rate evaluation against 1.5 percent for stock Sol, and says the model surfaced two previously unknown V8 vulnerabilities, one patched by Google as CVE-2026-15903. This follows Anthropic's cyber-focused Mythos and marks both frontier labs shipping deliberately less restricted models to gated security customers.
xAI Launches Grok 4.6, Betting on Post-Training Over Parameter Count
xAIxAI shipped Grok 4.6 on August 7, 2026 on the same 1.5 trillion parameter V9 foundation as Grok 4.5, with the entire upgrade coming from improved supervised fine-tuning and reinforcement learning. Musk positioned it to challenge Kimi K3 (roughly 2.8T parameters) and Claude Opus 4.8 while keeping Grok 4.5's throughput and cost envelope. xAI has not published a separate rate card, official benchmarks, or a confirmed context window at launch; Grok 4.5's $2 input / $6 output per 1M tokens is the working baseline. A larger 2.1T Grok 4.7 is slated to follow within weeks.
Model ReleaseOpenAI Retunes GPT-5.6 Sol and Makes Luna the Free-Tier Default in ChatGPT
OpenAIOpenAI rolled a retuned GPT-5.6 Sol into ChatGPT and made GPT-5.6 Luna the default model for free users, alongside a move toward unlimited text chat on the free tier. Combined with the July 30 API cuts (Luna to $0.20/$1.20, Terra to $2/$12 per 1M tokens), the free-tier switch puts a recently 80 percent discounted model in front of ChatGPT's largest audience and squeezes budget-tier rivals from the consumer side.
Meta Ships Muse Code and Muse Spark 1.2 With a $0.10/$0.20 Contributor Tier
MetaMeta released Muse Code, a terminal coding agent for macOS and Linux, together with Muse Spark 1.2, a coding-focused update co-trained inside the Muse Code harness. Standard API pricing holds at $1.25 input / $4.25 output per 1M tokens with a 1M context window, but the headline is muse-spark-1.2-contributor: the same weights at $0.10/$0.20 if you let Meta train on your data, undercutting DeepSeek V4 Flash ($0.14/$0.28) as the cheapest frontier-adjacent coding path. Meta reports 82.9 on Terminal-Bench 2.1 and 59.3 on DeepSWE 1.1, second to Claude Opus 5 on both; scores are vendor-reported and not yet on verified leaderboards.
Model ReleaseEU AI Act Becomes Generally Applicable as High-Risk Obligations Take Effect
EUAugust 2, 2026 is the EU AI Act's main application milestone: the bulk of the regulation, including the obligations for high-risk AI systems listed in Annex III and the Article 50 transparency rules for AI-generated content, is now generally applicable across member states. General-purpose model obligations began in August 2025; this date extends enforcement to deployers and providers of high-risk systems, with penalties of up to 35 million euros or 7 percent of global turnover for prohibited practices. Providers selling into the EU now need conformity assessments, technical documentation, and post-market monitoring in place rather than on a roadmap.
PolicyCalifornia SB 942 Goes Operative, the First Binding US Provenance Regime for Generative AI
CaliforniaCalifornia SB 942, the AI Transparency Act as amended by AB 853, became operative on August 2, 2026. Any generative AI provider with more than one million monthly users accessible in California must now offer a free public AI detection tool that answers whether content came from that provider's own system, a visible disclosure option, and a machine-readable C2PA-compatible latent disclosure embedded in AI-generated images, video, and audio. Enforcement runs at $5,000 per violation with each day counted separately, by the Attorney General, city attorneys, and county counsels. AB 853 moved the operative date from January 1 to August 2 to synchronize with the EU AI Act's Article 50 clock, and hosting-platform obligations follow on January 1, 2027.
PolicyOpenAI Cuts GPT-5.6 Luna by 80 Percent and Terra by 20 Percent Three Weeks After Launch
OpenAIOpenAI cut API pricing for two of the three GPT-5.6 models on July 30, three weeks after the family went generally available on July 9. GPT-5.6 Luna dropped 80 percent, from $1 per million input tokens and $6 per million output to $0.20 and $1.20, with cached input reads falling to $0.02. GPT-5.6 Terra dropped 20 percent, from $2.50/$15 to $2/$12, with cached input at $0.20. Flagship Sol stays at $5/$30 but gains a fast mode at twice the price for roughly 2.5 times the speed. ChatGPT and Codex subscription prices and quota budgets are unchanged, though Terra and Luna usage now draws down fewer credits. The cut lands Luna below DeepSeek V4 Flash on output price and reads as a direct response to open-weight price pressure from Kimi K3 and GLM 5.2.
PricingMeta and BlackRock Form $14 Billion, 1 GW El Paso Data Center JV on the 80/20 Template
MetaMeta and BlackRock announced a $14 billion joint venture to develop a 1 gigawatt data center campus in El Paso, Texas, targeting a 2028 launch with Meta as sole tenant. BlackRock funds own 80 percent and Meta keeps 20, with Meta contributing $2.3 billion in land, receiving a $1 billion one-time payment, BlackRock writing $4.9 billion of cash, and $12.5 billion of bonds on top of the capital stack. It is Meta's second 80/20 structure in nine months after the Blue Owl Hyperion JV in Richland Parish, which has since expanded to 5 gigawatts and $50 billion: an asset manager holds title to the gigawatt while Meta signs a compute offtake and keeps the tokens off its CapEx line.
1,178 Frontier Lab Employees Sign Pacing the Frontier; OpenAI and Anthropic Endorse, Meta and Google Do Not
Multiple1,178 employees at OpenAI, Anthropic, Google DeepMind, Meta, and Thinking Machines signed Pacing the Frontier, an open letter asking Washington to fund the technical and governance tooling needed for a verifiable slowdown if recursive self-improvement outruns oversight. Within about six hours OpenAI and Anthropic endorsed the letter at the corporate level. Meta declined to comment, Google did not respond, and Mark Zuckerberg published a Wall Street Journal column the same afternoon arguing that broadly distributed weights are the pacing mechanism. The split mirrors the federal launch-bar fight: the two labs that helped write the pre-release gate endorsed a second regulatory instrument, and the two that route around a closed perimeter said no.
PolicyNVIDIA Invests $5 Billion in Safe Superintelligence for Vera Rubin Compute Access
Safe SuperintelligenceNVIDIA announced a $5 billion investment in Safe Superintelligence, the AI startup co-founded by Ilya Sutskever, paired with a long-term partnership that gives SSI access to NVIDIA's next-generation Vera Rubin CPU-GPU platform and increases its compute by an order of magnitude. SSI has said it does not plan to commercialize products in the near term; it previously raised roughly $3 billion from investors including Andreessen Horowitz, Sequoia, and DST at a reported $32 billion valuation. For NVIDIA the deal buys privileged visibility into frontier research workloads at a moment when hyperscaler customers are shifting spend toward custom silicon.
Moonshot Releases Kimi K3 Open Weights, the Largest Open-Weight Model Ever Shipped
Moonshot AIMoonshot AI published the weights for Kimi K3 on July 27, eleven days after the API launch. The mixture-of-experts model totals 2.8 trillion parameters with 104 billion active per token, takes text, image, and video input, and carries a 1 million token context window. The download weighs roughly 1.4 terabytes in MXFP4 quantization, making it the largest open-weight release to date. K3 sits fourth on the Artificial Analysis Intelligence Index at 57, first among open-weight models, and holds the number one spot on LMArena's Frontend Code Arena. The release came with a technical report and per-benchmark footnotes naming the harness behind every published score, and API pricing stays at $3 per million input tokens and $15 per million output.
Open SourceAnthropic Ships Claude Opus 5 at Opus 4.8 Pricing, Half the Cost of Fable 5
AnthropicAnthropic released Claude Opus 5 (API id claude-opus-5) at $5 per million input tokens and $25 per million output, unchanged from Claude Opus 4.8 and exactly half of Claude Fable 5 at $10/$50. It carries a 1 million token context window as both default and maximum, up to 128K output tokens, and text plus vision input. Anthropic's launch table shows it leading Fable 5 on most published rows despite the lower price: 70.6 percent on OSWorld 2.0 versus 66.1, 90.8 percent on BrowseComp versus 87.4, 64.7 percent on Humanity's Last Exam with tools versus 63.9, and a GDPval-AA v2 score of 1861 versus 1747. It also posts 30.2 percent on ARC-AGI-3, where Opus 4.8 scored 1.5. Fable 5 keeps narrow leads on FrontierCode and DeepSWE and remains Anthropic's stated top capability tier. All figures are vendor-reported. Adaptive thinking is now on by default when the thinking parameter is omitted, disabling thinking is rejected above high effort, and the minimum cacheable prompt drops to 512 tokens from 1024. Opus 5 uses a separate rate-limit pool from the Opus 4.x models and is excluded from Priority Tier. Available on the Claude API, Amazon Bedrock, Google Cloud, and Microsoft Foundry.
Model ReleaseAnt Group Releases Ling-3.0-flash and Black Forest Labs Announces FLUX 3
Ant GroupAnt Group released Ling-3.0-flash, a 124 billion parameter mixture-of-experts activating roughly 5.1 billion per token, claiming it matches the company's own 1 trillion parameter flagship on a twelfth of the active parameters. Ant states an Apache 2.0 license, but the weights had not been posted at launch, so the open-weight status is stated rather than confirmed. The model is free on OpenRouter through August 3, 2026, which makes it a zero-cost evaluation window. All performance figures are vendor-claimed with no independent verification published. The same day Black Forest Labs announced FLUX 3, its first multimodal frontier model, jointly training image, video, audio, and action. Only the video variant is in gated early access, the image model is weeks out, the open-weight Dev backbone has no date, and the only evidence published is Black Forest Labs' own human-preference evals (a claimed 77 percent preference over Runway Gen-4.5 and 93 percent over Luma Ray 3.2). Neither release is in the TensorFeed pricing index yet: Ling has no list price and FLUX 3 has no pricing at all.
Model ReleaseCursor Ships Router as Model Routing Becomes the Answer to Release Velocity
CursorCursor released Cursor Router, per-request automatic model routing that classifies each task and sends it to the cheapest capable model, claiming 30 to 50 percent cost savings in early access with no quality drop. The claim is vendor-stated with press corroboration. The context is what makes it notable: two days earlier Cursor published an agent-swarm SQLite rebuild in which an equivalent-quality task cost $9,373 with a single frontier model doing everything and $411 with a planner and worker pairing, a roughly 23x spread driven entirely by model choice rather than by doing less work. Cline's same-week head-to-head on a real bug fix landed the same point from the other direction: Claude Fable 5 finished 3.4x faster (3.5 minutes, 18 tool calls) while Kimi K3 came out 2.3x cheaper overall ($0.92 against $2.13) despite consuming 1.7x more tokens. Speed, cost, and token count went to different models on a single task. With seven models shipping in seven days, routing infrastructure is becoming the layer that absorbs release velocity so application code does not have to.
Google Ships Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber, Confirms Gemini 4 Pre-Training
GoogleGoogle released three models in a single announcement. Gemini 3.6 Flash (model id gemini-3.6-flash) is the new workhorse at $1.50 per million input tokens and $7.50 per million output, a cut from the $9.00 output rate on Gemini 3.5 Flash, with cached input at $0.15, a 1,048,576 token input window, 64K output, and text, image, audio, video, and PDF input. Google reports DeepSWE at 49 percent against 37 for 3.5 Flash, MLE-Bench at 63.9 against 49.7, OSWorld-Verified at 83.0 against 78.4, GDPval-AA v2 at 1421 Elo against 1349, and 17 percent fewer output tokens on the Artificial Analysis Index. Artificial Analysis independently scored the Intelligence Index flat at 50 for both models while measuring average time per task falling from 2.7 minutes to 1.3 and cost per task from $0.59 to $0.50, which frames this as a per-task economics upgrade rather than a capability jump. Gemini 3.5 Flash-Lite landed at $0.30 and $2.50 running 350 output tokens per second, beating Gemini 3 Flash on SWE-Bench Pro (54.2 against 49.6) and OSWorld-Verified (74.0 against 65.1). Gemini 3.5 Flash Cyber, built on 3.5 Flash and fine-tuned to find and patch vulnerabilities, is restricted to governments and trusted partners through the CodeMender agent under a limited-access pilot, with no public pricing. Google also confirmed that Gemini 3.5 Pro is still only testing with partners after missing its July 17 target, and that pre-training has begun on Gemini 4, which it calls its most ambitious run yet. Computer use is now a built-in client-side tool in the Gemini API.
Model Releasepoolside Releases Laguna S 2.1, an 8B-Active Open-Weight Coding Model at $0.10/$0.20
poolsidepoolside, the San Francisco lab founded by former GitHub CTO Jason Warner and Eiso Kant, released Laguna S 2.1, its first major public model. It is a 118 billion parameter mixture-of-experts activating roughly 8 billion per token, with a 1,048,576 token context window and up to 131,072 output tokens. Weights hit Hugging Face on day one under the OpenMDW-1.1 license, and hosted pricing on OpenRouter is $0.10 per million input tokens, $0.20 per million output, and $0.01 for cache reads. poolside reports 70.2 percent on Terminal-Bench 2.1 in thinking mode, 78.5 percent on SWE-Bench Multilingual, and 40.4 percent on DeepSWE, claiming it outranks models roughly ten times its size. The more durable detail is the eval hygiene: poolside disclosed that its Terminal-Bench numbers come from its own harness and published the full unedited trajectory of every benchmark trial, in a week when three separate launches shipped with nothing an independent party could rerun. At 8 billion active parameters it is also one of the few genuinely self-hostable open-weight coding models, in contrast to the trillion-parameter open releases that still need a rack to load.
Open SourceAlibaba Compresses Three Qwen Releases Into 72 Hours With Almost No Published Evidence
AlibabaAlibaba shipped three Qwen models in three days. Qwen3.8-Max-Preview debuted July 19 at WAIC in Shanghai via a single X post claiming it ranks second only to Claude Fable 5, with no benchmark table, model card, or independent leaderboard listing, and it is closed and API-only. Qwen-Audio-3.0-TTS followed July 20 as a hosted-only voice model priced at roughly a third of ElevenLabs and MiniMax rates, and it is the one release of the three with independent verification: first place on the Artificial Analysis TTS arena at about 1,236 Elo, a statistical tie with second. Qwen-Image-3.0 arrived July 21 through an invite-only API with no benchmarks, weights, license, or technical report, a reversal from Qwen-Image 1.0, which shipped Apache 2.0 the same day it launched. The pattern across the three is a pivot away from the open-weight posture that built Qwen's developer base. None are in the TensorFeed pricing index: the two LLM-adjacent releases publish no usable list pricing or verifiable specs, and the TTS model does not fit a per-token schema.
Model ReleaseCuspAI Raises $450 Million Series B for AI-Designed Materials
CuspAICambridge, UK based CuspAI raised a $450 million Series B led by Kleiner Perkins and NEA, with participation from Bezos Expeditions, the UK government, AMD Ventures, Lux Capital, and Glade Brook Capital Partners, bringing total funding past $650 million. The company applies AI to materials discovery. The same week London-based AI and robotics firm Humanoid raised $152 million at a $1.35 billion post-money valuation. The pattern across the week's rounds is a shift away from paying for generic AI exposure and toward control points: the largest checks went to companies sitting close to budget owners, regulated workflows, or procurement systems that are difficult to displace once embedded.
Gemini 3.5 Pro Misses Its July 17 Target as Google Weighs a Stopgap Flash Release
GoogleGoogle DeepMind let the July 17, 2026 target for Gemini 3.5 Pro pass without a launch, the third missed window after the original Google I/O timeframe in May and a June slip. Reporting attributes the delay to persistent hallucinations, inconsistent real-world outputs, and structural failures in recursive tool calling and SVG generation in the rebuilt architecture, plus a failure to match GPT-5.6 on key benchmarks. Google has registered names including Gemini 3.6 Flash and Gemini 3.5 Flash Light, a sign it may ship a cheaper stopgap Flash variant to manage the competitive gap while Pro development continues. There is still no public API, model card, or pricing for Gemini 3.5 Pro, and every spec, including the rumored 2 million token context window, remains unconfirmed by Google.
Model ReleaseHelsing Raises $1.8 Billion in Europe's Largest Venture Round
HelsingMunich defense AI startup Helsing raised $1.8 billion in a Series E at an $18 billion post-money valuation on July 14, 2026, co-led by Lightspeed and General Catalyst, the largest venture round in European history. It capped another concentrated week for AI capital: Fireworks closed a $1.505 billion Series D at a $17.5 billion valuation led by Atreides, Index Ventures, and TCV, Databricks signed a term sheet for a strategic round at a $188 billion valuation led by Coatue, and Chai Discovery raised $400 million for AI drug discovery. PitchBook put first half 2026 US venture funding at $412.7 billion, with AI deals dominating.
DeepSeek Sets a July 24 Deadline to Retire deepseek-chat and deepseek-reasoner
DeepSeekDeepSeek notified developers that the legacy deepseek-chat and deepseek-reasoner API endpoints stop working on July 24, 2026 at 15:59 UTC, forcing a migration to the V4 model names that have served those aliases since the V4 preview launched on April 24, 2026. The fix is a one-line change of the model parameter to deepseek-v4-pro or deepseek-v4-flash on the same base URL and key, but the mapping has a trap: deepseek-reasoner routes to V4 Flash in thinking mode, not to the stronger V4 Pro, so teams running heavy reasoning on the old alias can quietly drop to Flash tier capability unless they switch explicitly. V4 uses a 1.6 trillion parameter mixture of experts design that activates 49 billion parameters per token, with V4 Pro priced near one fifth of GPT-5.5.
Model ReleaseAnthropic Moves Claude Fable 5 to Metered Usage Credits on July 20
AnthropicAnthropic set July 20, 2026 as the switch to metered usage credits for Claude Fable 5, ending a run of free included access for subscribers that it had extended twice in a week, first to July 12 and then to July 19. From July 20, Fable 5 bills at $10 per million input tokens and $50 per million output on Pro, Max, Team, and select Enterprise plans, with the existing prompt caching discounts (cache reads at $1.00 per million, 5 minute cache writes at $12.50, 1 hour writes at $20.00). Fable 5 is Anthropic's Mythos class flagship for autonomous knowledge work and coding, with a 1 million token context window and up to 128,000 tokens of output. The rate matches the standing list price, so this transition confirms the current number rather than changing it.
PricingOpenAI Makes GPT-5.6 Sol, Terra, and Luna Generally Available
OpenAIOpenAI moved the GPT-5.6 family to general availability on July 9, 2026, two weeks after the June 26 limited preview that had been gated to roughly 20 pre-approved organizations. The three tiers are Sol (flagship, $5 input / $30 output per 1M tokens), Terra (balanced, $2.50 / $15, which OpenAI says matches GPT-5.5 at roughly half the cost), and Luna (fast and cheap, $1 / $6), with cached reads keeping a 90% discount. GPT-5.6 becomes the new default in ChatGPT: Free and Go users get Terra in Work mode, while paid users default to Sol. OpenAI cited gains across coding, biology, and cybersecurity. Aggregate scores put Sol Ultra at 91.9, Sol at 88.8, Terra at 84.3, and Luna at 82.5 against GPT-5.5 at 83.4.
Model ReleaseMeta Opens Muse Spark 1.1, Its First Paid API Model, at $1.25/$4.25
MetaMeta released Muse Spark 1.1 on July 9, 2026, its first paid API model and a break from the fully open-source Llama strategy. Priced at $1.25 per million input tokens and $4.25 per million output, roughly a quarter of OpenAI and Anthropic list rates, with $20 in free credits per new account and a launch preview limited to US developers. Muse Spark 1.1 is a multimodal reasoning model built for agentic and coding work, with a self-managed 1 million token context window, native primary-agent and subagent orchestration, and MCP plus custom-skill support. Meta AI chief Alexandr Wang called it the company's strongest model for agentic and coding work yet. It tops tool-use suites (88.1 on MCP Atlas, 54.7 on JobBench, 62.1 on Humanity's Last Exam with tools) but trails the leaders on pure coding (61.5 on SWE-Bench Pro, 80.0 on Terminal-Bench 2.1). Mark Zuckerberg returned to X after three years to promote the launch.
Model ReleasexAI Releases Grok 4.5, a Coding Model on the V9 Foundation at $2/$6
xAIxAI released Grok 4.5 on July 8, 2026, its first model built specifically for coding and agentic work, ten days after the private-beta announcement. Built on the 1.5 trillion parameter V9 foundation (up from Grok 4.3's V8) and trained on real coding-agent data, it lands fourth on the Artificial Analysis Intelligence Index, above every open-weight model and every Gemini model, at a price more than 60% below Claude Opus 4.8 or GPT-5.5. API pricing is $2 per million input tokens and $6 per million output, with cached input at $0.50 (a 75% discount) and a higher-context surcharge above 200K tokens. xAI narrowed the context window to 500K tokens to focus on coding. Grok 4.5 leads Opus 4.8 on the provider-harness DeepSWE 1.0 score and on Terminal-Bench 2.1 (83.3) but trails it on the neutral DeepSWE 1.1 run and on SWE-Bench Pro (64.7). Built-in tools bill separately at $2.50 to $10 per 1,000 calls.
Model ReleaseGoogle Delays Gemini 3.5 Pro to July 17 for a Full Architectural Rebuild
GoogleGoogle DeepMind delayed Gemini 3.5 Pro to July 17, 2026, scrapping the existing 2.5 Pro architecture for a complete rebuild. The overhaul targets mathematical reasoning, SVG scene generation, and image quality to compete with OpenAI's GPT-5.6 and Anthropic's Fable 5. The model was originally expected around Google I/O in May and slipped through June. As of the delay there is no public API availability, and the launch now trails GPT-5.6 general availability and the Grok 4.5 and Muse Spark 1.1 releases from the same week.
Model ReleaseBaseten Raises $1.5 Billion Series F as AI Inference Infrastructure Consolidates
BasetenAI inference platform Baseten raised $1.5 billion in a Series F on July 6, 2026, its fourth fundraise in 18 months, as capital continued to concentrate in the infrastructure layer that serves open-weight and proprietary models. The round landed in a week where roughly four of every five venture dollars went to AI infrastructure, alongside Groq closing $650 million led by Infinitum and Disruptive and Taktile raising $110 million for an agentic decision platform for banks and insurers. PitchBook reported US venture funding hit $412.7 billion in the first half of 2026, with AI deals dominating.
Anthropic Releases Claude Sonnet 5 at $2/$10 Introductory Pricing
AnthropicAnthropic shipped Claude Sonnet 5 on June 30, 2026 as the most agentic Sonnet-class model yet, positioned to narrow the gap to Opus 4.8 on reasoning, tool use, coding, computer use, and knowledge work while staying priced below the flagship. Introductory API rates were $2 per million input tokens and $10 per million output tokens through August 31, 2026. Anthropic has since confirmed that $2 and $10 is the standard price and that the scheduled increase to $3 and $15 on September 1, 2026 will not happen. Sonnet 5 ships a 1 million token context window with context compaction and adaptive thinking with selectable effort levels up to xhigh. Anthropic's own numbers show 85.2 on SWE-bench Verified, 63.2 on SWE-bench Pro, 78.3 on SWE-bench Multilingual, 81.2 on OSWorld-Verified, 84.7 on BrowseComp agentic search, and 80.4 on Terminal-Bench 2.1 (which beats Opus 4.8 at 74.6 on that specific benchmark). Available on the Claude API, Amazon Bedrock, Google Vertex, and Microsoft Foundry at launch.
Model ReleaseMeituan Open Sources LongCat-2.0, a 1.6T Agentic Coding Model Trained on Chinese Chips
MeituanMeituan open sourced LongCat-2.0 on June 30, 2026, a 1.6 trillion parameter mixture-of-experts coding model published to GitHub and Hugging Face under an MIT license. Dynamic activation of 33 to 56 billion parameters per token, native 1 million token context, and a 30 trillion token pretraining mix spanning Chinese, English, multilingual, and code data. Meituan reports 59.5 on SWE-Bench Pro (self-reported, ahead of Gemini 3.1 Pro, GPT-5.5, and Claude Opus 4.6 by their measure). The model is the first trillion-parameter release to complete full training and inference entirely on a 50,000-card domestic Chinese compute cluster, an important signal for China's ability to build frontier AI without leading-edge Western chips. A preview version has been quietly running on OpenRouter and longcat.ai for weeks, ranking among the top three models globally by call volume during that stealth window.
Open SourceElon Musk Puts Grok 4.5 in Private Beta at SpaceX and Tesla
xAIElon Musk announced on June 28, 2026 that xAI's Grok 4.5 is in private beta at SpaceX and Tesla, its first deployment before any wider release. Grok 4.5 is built on xAI's 1.5 trillion parameter V9 foundation model, which finished training on May 26, 2026 and used data from the Cursor coding environment in supplemental training. Musk claimed early evaluations put performance close to, and potentially above, Claude Opus, but xAI has not published benchmarks or a system card, so capability claims remain unverified. Reinforcement learning is ongoing, and xAI says it plans to release new models trained from scratch through SpaceX every month through the end of 2026. There is no public availability window and no API price for Grok 4.5 as of the announcement.
Model ReleaseOpenAI Previews GPT-5.6 Sol, Terra, and Luna in Limited Release
OpenAIOpenAI unveiled GPT-5.6 as a three-model family (Sol, Terra, Luna) on June 26, 2026, then immediately gated it. Per a US Government request, GPT-5.6 went live only to roughly 20 pre-approved organizations in a limited preview, with general availability planned in the coming weeks. Pricing per 1M tokens is $5 input / $30 output for Sol (frontier reasoning and agentic work), $2.50 / $15 for Terra (a balanced model OpenAI says delivers GPT-5.5-level performance at roughly half the cost), and $1 / $6 for Luna (the fastest and cheapest variant). OpenAI published partial benchmarks (coding, biology, cybersecurity), with Sol Ultra reaching 91.9% on Terminal-Bench 2.1 and base Sol at 88.8%. SWE-bench, MMLU-Pro, GDPval, and FrontierMath numbers were held back until GA. A Cerebras deployment targeting roughly 750 tokens per second on Sol is planned for July.
Model ReleaseOpenAI Retires GPT-4.5 From ChatGPT
OpenAIGPT-4.5 was removed from ChatGPT on June 26, 2026. Existing GPT-4.5 conversations route forward to GPT-5.5, and earlier GPT-5.2 variants (Instant, Thinking, Pro) had already been pulled from ChatGPT on June 12. The API listing for gpt-4.5 remains for now, but the deprecation tightens the active OpenAI lineup to GPT-5.5 (with GPT-5.6 still in limited preview), o1, and o3-mini.
Model ReleaseAnthropic Moves Claude Fable 5 From Plan-Included to Usage Credits
AnthropicAnthropic ended the introductory window for Claude Fable 5 on June 23, 2026. From June 9 through June 22 the model was included at no extra cost on Pro, Max, Team, and seat-based Enterprise plans. Starting June 23 Fable 5 use is billed against usage credits on those plans; API pricing stays at $10 per 1M input and $50 per 1M output, with safety-classifier reroutes to Opus 4.8 billed at Opus rates for the rerouted portion. Opus 4.8 remains the default plan-included model.
PricingAnthropic Releases Claude Fable 5, a New Frontier Tier Above Opus
AnthropicAnthropic launched Claude Fable 5 (API id claude-fable-5), a new frontier tier positioned above Claude Opus 4.8, which remains the default model. Pricing is $10 per million input tokens and $50 per million output with no long-context surcharge, on a 1 million token default context window with up to 128K output tokens and text plus vision input. Adaptive thinking is always on and cannot be disabled, with effort levels spanning low, medium, high, xhigh, and max. Anthropic's launch table reports 80.3 on SWE-bench Pro, 29.3 on FrontierCode Diamond, 85.0 on OSWorld-Verified, and a GDPval-AA ELO of 1932, all vendor-reported. Always-on safety classifiers can reroute flagged requests to Opus 4.8, billed at Opus rates for the rerouted portion. The model is available on the Claude API, Amazon Bedrock, Google Vertex, and Microsoft Foundry at launch.
Model ReleaseMicrosoft Launches MAI-Code-1-Flash and MAI-Thinking-1 at Build 2026
MicrosoftMicrosoft announced its first in-house coding and reasoning models at Build 2026 in San Francisco. MAI-Code-1-Flash rolls out across all GitHub Copilot tiers with a 256K context window, priced at $0.75 per million input tokens ($0.075 cached) and $4.50 per million output, with Microsoft claiming better price to performance than Claude Haiku 4.5 and 60% fewer tokens used on hard tasks. MAI-Thinking-1, a 35B active parameter MoE reasoning model with a 256K context window, posted 97% on AIME 25 and 53% on SWE-Bench Pro, which Microsoft says matches Claude Opus 4.6. It enters private preview on Microsoft Foundry with distribution planned through Fireworks AI, Baseten, and OpenRouter. The launches mark a deliberate step away from Microsoft's dependence on OpenAI models inside Copilot.
Model ReleaseAnthropic Confidentially Files S-1 for IPO
AnthropicAnthropic confidentially submitted a draft S-1 registration statement to the SEC for a proposed initial public offering, less than a week after closing a $65 billion Series H at a $965 billion post-money valuation. Annualized revenue run-rate recently crossed $47 billion on enterprise adoption of Claude for coding and agentic workflows. Share count and pricing are not yet determined. The move puts Anthropic ahead of OpenAI, which is reportedly preparing its own confidential filing, in the race to a landmark AI listing.
MiniMax Releases M3 Open-Weight Coding Model
MiniMaxMiniMax launched M3, an open-weight coding and agentic model built on MiniMax Sparse Attention, which replaces full attention with KV-block selection and cuts per-token compute at 1M context to roughly one twentieth of the previous generation. M3 accepts text, image, and video input across a 1,048,576 token context window with up to 512K output tokens, priced at $0.30 per million input and $1.20 per million output. MiniMax reports 59% on SWE-Bench Pro and 83.5 on BrowseComp, though headline runs used its own infrastructure with agent scaffolding, so independent verification is pending. Weights and a technical report are due on Hugging Face within about ten days.
Open SourceAnthropic Raises $65B Series H at $965B Valuation
AnthropicAnthropic closed a $65 billion Series H funding round at a $965 billion post-money valuation, surpassing OpenAI's reported $852 billion to become the most valuable private AI company. The round landed the same day as the Claude Opus 4.8 release and set the stage for the confidential IPO filing that followed on June 1.
Anthropic Releases Claude Opus 4.8
AnthropicAnthropic shipped Claude Opus 4.8 just six weeks after Opus 4.7, keeping the 1 million token context window. Anthropic's own measures put agentic coding at 69.2% (up from 64.3%) and knowledge work at 1890 (up from 1753), with the model roughly 4x less likely than 4.7 to let flaws in its own code slip through while using about 35% fewer tokens per task. A new fast mode runs roughly 2.5x quicker and about three times cheaper than before, and Dynamic Workflows adds large-scale parallel subagent support.
Model ReleaseTrump Postpones AI Executive Order Hours Before Signing
US GovernmentPresident Trump pulled a landmark AI executive order from the schedule hours before its planned signing ceremony on May 21. The order would have established a voluntary 90-day pre-launch review framework for frontier AI models with NSA involvement and tasked federal agencies with using AI models to harden government network defenses. Trump cited concerns the framework could slow US competitiveness against China. Former AI czar David Sacks, Mark Zuckerberg, and Elon Musk are reported to have lobbied against it directly. No new signing date announced.
PolicyCoupa Acquires Tonkean for Agentic Procurement
CoupaCoupa acquired Tonkean, a Palo Alto-based no-code workflow orchestration platform with 250+ native connectors and multi-agent coordination. Financial terms were not disclosed. The deal is Coupa's second AI-focused acquisition in two weeks, following the May 12 Rossum (intelligent document processing) deal, and feeds Coupa's emerging agentic trade network across procurement, invoicing, and supplier transactions.
AcquisitionCohere Releases Command A+ Open Source Under Apache 2.0
CohereCohere released Command A+, a 218B-total / 25B-active mixture-of-experts model with a 128K input context and 64K max generation, under an Apache 2.0 license. Multimodal text and image input, tool use, 48-language coverage, available on Hugging Face in BF16, FP8, and W4A4 quantizations. Runs on a single NVIDIA Blackwell B200 or two H100s at W4A4. Scores 37 on the Artificial Analysis Intelligence Index, 75.1% on MMMU, and 80.6% on MathVista. Cohere's first fully Apache 2.0 enterprise model and first multimodal reasoning model.
Open SourceAlibaba Unveils Qwen3.7-Max at Cloud Summit
AlibabaAlibaba unveils Qwen3.7-Max at the 2026 Alibaba Cloud Summit, the flagship proprietary model in the Qwen3.7 family. 1 million token context window, extended thinking mode, claims of autonomous operation up to 35 hours on long-horizon agentic tasks. Scored 57 on the Artificial Analysis Intelligence Index (first place on the public leaderboard) and roughly 1,475 Elo on the LM Arena text leaderboard. Landed on the Alibaba API platform on May 19, formally announced May 20. Priced at $2.50 per million input and $7.50 per million output tokens via OpenRouter.
Model ReleaseGoogle Launches Gemini 3.5 Flash at I/O 2026
GoogleGoogle ships Gemini 3.5 Flash to general availability at I/O 2026. Priced at $1.50 per million input and $9.00 per million output tokens, with a 1,048,576 token context window. The first Flash-tier release that outscores the previous Pro flagship (Gemini 3.1 Pro) on agentic coding suites including Terminal-Bench 2.1 (76.2%) and MCP Atlas (83.6%), at roughly 4x the throughput.
Model ReleaseGoogle Announces Gemini Spark Agent at I/O 2026
GoogleGoogle introduces Gemini Spark, a general-purpose AI agent inside the Gemini app that can reason across connected apps and take actions on the user's behalf. Available first to trusted testers and Google AI Ultra subscribers in the week following the keynote.
OpenAI Launches GPT-5.5-Cyber Preview
OpenAIOpenAI rolls out GPT-5.5-Cyber, a variant of GPT-5.5 tuned for cybersecurity workflows, in limited preview to vetted defenders. Launch partners include Cisco, CrowdStrike, Palo Alto Networks, Cloudflare, Intel, Snyk, and SentinelOne. The model offers reduced classifier-based refusals for authorized red team and vulnerability research.
Model ReleaseOpenAI Releases GPT-5.5 Instant for ChatGPT
OpenAIOpenAI ships GPT-5.5 Instant as the new default ChatGPT model, replacing GPT-5.4 Instant. The update tightens accuracy, cuts gratuitous emoji output, and rolls out to free, Plus, Pro, Business, and Edu tiers.
Google Ships Gemini 3.1 Flash-Lite Preview
GoogleGoogle releases Gemini 3.1 Flash-Lite in preview through the Gemini API and Vertex AI. Priced at $0.25 per million input and $1.50 per million output, with a 1M context window, the model targets high-volume workloads at half the cost of Gemini 3 Flash.
Model ReleaseNVIDIA Releases Nemotron 3 Nano Omni
NVIDIANVIDIA ships Nemotron 3 Nano Omni 30B-A3B-Reasoning, an open-weight multimodal model that processes text, image, video, and audio in a unified sequence. Hybrid Mamba-Transformer-MoE backbone (30B total, 3B active per token), 256K context window, native audio handling up to 20 minutes per clip. Tops six leaderboards for document intelligence and video/audio understanding including OCRBenchV2-En (65.8), MMLongBench-Doc (57.5), OSWorld (47.4), Video-MME (72.2), and VoiceBench (89.4). Available on Hugging Face in BF16, FP8, and NVFP4 quantizations.
Open SourceAmazon Commits $25B to Anthropic
AnthropicAmazon invests up to $25 billion in additional funding to Anthropic, bringing its total commitment to $33 billion. The deal includes over $100 billion in AWS infrastructure spending over the next decade.
DeepSeek V4 Pro & Flash Released
DeepSeekDeepSeek releases V4 Pro (1.6T parameters, 49B active) and V4 Flash (284B total, 13B active) under the MIT license with native 1M context windows. V4 Pro scores 80.6% on SWE-bench Verified.
Open SourceGPT-5.5 Released
OpenAIOpenAI launches GPT-5.5, its first fully retrained base model since GPT-4.5. Features 1M context, native omnimodal capabilities, and benchmark leadership at $5/$30 per million tokens.
Model ReleaseClaude Design Launched
AnthropicAnthropic launches Claude Design as an Anthropic Labs research preview. It reads codebases, generates design systems, and hands off finished prototypes to Claude Code.
SpaceX Acquires xAI for $250B
xAISpaceX acquires Elon Musk's xAI in a $250 billion deal, the largest acquisition in AI history. The merger brings Grok and the Colossus GPU cluster under SpaceX.
AcquisitionClaude Opus 4.7 Released
AnthropicAnthropic releases Claude Opus 4.7 with a 1 million token context window at flagship pricing and incremental gains across reasoning, code, and SWE-bench.
Model ReleaseClaude Sonnet 4.6 Released
AnthropicAnthropic releases Claude Sonnet 4.6, a faster and more affordable model in the 4.6 family with strong coding and reasoning performance.
Model ReleaseClaude Opus 4.6 Released
AnthropicAnthropic releases Claude Opus 4.6 with extended thinking and a 1M context window option, setting new benchmarks across reasoning tasks.
Model ReleaseGemini 2.5 Pro Released
GoogleGoogle launches Gemini 2.5 Pro with a native 1M token context window and improved multimodal reasoning capabilities.
Model ReleaseLlama 4 Scout & Maverick Released
MetaMeta releases Llama 4 with Scout (10M context) and Maverick (1M context) variants, pushing open-source model capabilities forward.
Open SourceGPT-4.5 Released
OpenAIOpenAI launches GPT-4.5 with a 256K context window and significantly improved creative writing and emotional intelligence.
Model ReleaseGrok 3 Released
xAIxAI releases Grok 3 with strong reasoning performance, trained on the Colossus 200K GPU cluster.
Model ReleaseDeepSeek V3 Released
DeepSeekDeepSeek releases V3, a powerful open-source model trained with a fraction of the compute used by competitors, sparking industry debate about efficiency.
Open SourceClaude 3.5 Haiku Released
AnthropicAnthropic releases Claude 3.5 Haiku, a fast and affordable model that outperforms the original Claude 3 Opus on many benchmarks.
Model ReleaseClaude 3.5 Sonnet v2 Released
AnthropicAnthropic ships an upgraded Claude 3.5 Sonnet with computer use capabilities and significant improvements in coding and tool use.
Model ReleaseEU AI Act Enforcement Begins
EUThe European Union begins enforcing the AI Act, requiring risk assessments and transparency for high-risk AI systems deployed in EU markets.
PolicyOpenAI o1 Released
OpenAIOpenAI launches o1, a reasoning-focused model that uses chain-of-thought at inference time to solve complex math and coding problems.
Model ReleaseLlama 3.1 405B Released
MetaMeta releases Llama 3.1 with a 405B parameter variant, the largest open-source model at the time, matching frontier closed models.
Open SourceGPT-4o Mini Released
OpenAIOpenAI launches GPT-4o Mini at $0.15 per million input tokens, dramatically undercutting existing pricing and pressuring the entire market.
PricingClaude 3.5 Sonnet Released
AnthropicAnthropic releases Claude 3.5 Sonnet, surpassing GPT-4o on most benchmarks while maintaining fast response times and lower costs.
Model ReleaseGPT-4o Released
OpenAIOpenAI releases GPT-4o with native multimodal capabilities including real-time voice and vision, available free to all ChatGPT users.
Model ReleaseLlama 3 Released
MetaMeta releases Llama 3 in 8B and 70B variants, establishing a new bar for open-source language model performance.
Open SourceClaude 3 Family Released
AnthropicAnthropic launches the Claude 3 model family (Haiku, Sonnet, Opus) with vision capabilities and a 200K context window.
Model ReleaseGemini 1.5 Pro Released
GoogleGoogle launches Gemini 1.5 Pro with a 1M token context window, a major leap in long-context processing for production use.
Model ReleaseGemini 2.0 Flash Released
GoogleGoogle releases Gemini 2.0 Flash with native tool use and multimodal output, optimized for agentic workflows.
Model ReleaseMistral Large 2 Released
MistralMistral releases Large 2 with a 256K context window, strong multilingual support, and competitive pricing for enterprise use.
Model ReleaseOpenAI Valued at $157B
OpenAIOpenAI closes a $6.6B funding round at a $157B valuation, making it the most valuable private AI company in the world.
AcquisitionEU AI Act Enters Into Force
EUThe EU AI Act officially enters into force, establishing the world's first comprehensive legal framework for artificial intelligence regulation.
PolicyAnthropic Raises $2.75B from Amazon
AnthropicAmazon completes its $2.75B additional investment in Anthropic, bringing its total commitment to $4B and deepening the AWS partnership.
Acquisition