Frontier AI Labs, Daily
Two model launches, one price point, two theories about what the frontier becomes next.
Date Tuesday, September 29, 2026 Window Sep 28 to Sep 29 Anchor Anthropic Also OpenAI, Meta AI, xAI
The Open
The frame for today
Two model launches. One price. Opposite bets on what capable means.

Anthropic's Sonnet 5.5 shipped Tuesday morning. The Terminal-Bench 4.0 number is 70.6%, against Sonnet 5's 10.3%. A gap that wide does not describe an improvement; it describes a threshold crossed. The tasks that required Opus, Sonnet can now handle. The things you were paying the expensive model to do, the mid-tier model can do now.

OpenAI was at Fort Mason for DevDay 2026. GPT-6.1 Sol launched at the same price point: $2 per million tokens in, $10 out. Dots launched as always-on agents with their own cloud computers and 4,000-plus app integrations. An ultrafast inference tier launched at 6x standard rates for latency-sensitive workloads. Three announcements, one event, one carefully chosen Tuesday in San Francisco.

Meta launched an enterprise platform and expanded its Muse agent to small businesses the same day. The frontier on September 29 is not a quiet Tuesday. It is the mid-tier repricing itself, and two labs arriving at identical price points with completely different arguments for why you should choose them.

Lead Story 01
Anthropic, Model Launch, September 28, 2026

Sonnet
5.5

Anthropic closes the agentic coding gap. A Terminal-Bench score of 70.6% changes what the mid-tier can do.
Source anthropic.com/claude-sonnet-5-5    Area Model Launch    Price $2 / $10 per million tokens
By the Numbers 70.6% Terminal-Bench 4.0
(vs. 10.3% for Sonnet 5)

30%+ faster outputs

Up to 30% lower per-task cost

$2 / $10 per million tokens

Available on API, AWS, Google Cloud, Azure Foundry
Sonnet 5.5, Lead Story

The number is 70.6%.

Sonnet 5's Terminal-Bench 4.0 score was 10.3%. That is not a gap you close with incremental tuning. It describes two different tools wearing the same name. The Sonnet tier handled the fast, clean tasks. Opus handled the agentic coding work, the multi-step loops, the long-horizon reasoning, the things that required the expensive model to not make a mess. Sonnet and Opus were not interchangeable; they were complementary.

Sonnet 5.5 crosses that line. At 70.6% on Terminal-Bench 4.0, the capability that drove teams to Opus 5.5 now lives at Sonnet price. The mechanism is efficiency: Sonnet 5.5 generates outputs 30% faster and costs up to 30% less per task, at identical per-token pricing ($2/$10), because it completes the same work in fewer tokens. For agentic loops running hundreds of completions a day, 30% fewer tokens per task compounds. A team running 10,000 completions a day is looking at a different monthly bill than they were yesterday.

The second number in the launch is quieter and more consequential for enterprise buyers. Sonnet 5.5 is the first Sonnet model deployed with the cyber safeguards previously reserved for Opus 5.5. For compliance teams that needed Opus's safety pedigree, this changes a procurement question: the model they now need is not the most expensive one. That conversation opened for the first time today.

The cross-lab contrast is hard to ignore. GPT-6.1 Sol launched the same day, same price point, same pitch: near-flagship capability at mid-tier price. OpenAI's framing is "near-GPT-6 Astra intelligence for one-fifth Astra's cost." Anthropic's framing is a Terminal-Bench number and a system card. Both are talking to the same developer, the one running production agentic loops who wants to stop paying Opus or Astra prices. One came with a conference in San Francisco. The other came with a benchmark. The builder who has to choose between them this week will read the benchmark and watch the keynote replay at 2x speed. Then they will run evals, because this is not a decision you make from the press release.

The pricing parity itself deserves a moment. Two major labs launched competing models at identical price points on the same calendar day. Whether this is market coordination or parallel convergence, $2/$10 per million tokens is now the reference price for capable mid-tier models in late 2026. Launching above that price requires a flagship justification. Launching below it means racing on margin. The labs that have not yet landed at $2/$10 will. The ones that have are now competing entirely on argument.

Builder's move: update to claude-sonnet-5-5. Available now on API, AWS Bedrock, Google Cloud Vertex, and Microsoft Azure Foundry. Run your existing Opus 5.5 evals against it. The 30%-per-task efficiency claim is the one to verify against your own workload. Aggregate token counts across a full agentic run, not just latency on a single call.

Also Shipped
Four more releases that moved the frontier
OpenAI, DevDay 2026, September 29
Dots: Always-On Agents, and the Always-On Question

The model story at DevDay 2026 was GPT-6.1 Sol: near-Astra capability at $2/$10 per million tokens, available as gpt-6.1-sol via API and in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users. It is a direct market position against Sonnet 5.5, priced identically, aimed at the same developer. But the more interesting launch was Dots.

Each dot is an always-on agent: its own cloud computer, its own browser, a goal you assign, a connection to 4,000-plus apps, and a feedback loop that teaches it over time. No explicit invocation required. The dot runs while you sleep. Available to Pro, Business Premium, and Enterprise users at no extra charge, one dot per eligible seat. Reachable via Slack and Teams for background operation.

The mechanism is persistence. Today's agents are stateless between invocations. Dots are stateful over days. That distinction has real implications for both capability and trust. A stateful agent that ran toward a goal for 72 hours and made hundreds of small decisions along the way is harder to audit than one that was invoked, answered, and dismissed. OpenAI's framing is about the learning feedback loop; the transparency question is the harder one, and the DevDay materials leave it aspirational.

The contrast with Anthropic: Sonnet 5.5 is a capability upgrade for models a developer explicitly invokes. Dots is a paradigm for agents that run without being asked. They are not competing products. They may well be the same builder's next two purchases, in that order.

OpenAI, API, September 29
Ultrafast: 300 Tokens Per Second at 6x the Price

OpenAI's new premium inference tier delivers 300 tokens per second in Codex (8x the standard rate) and up to 6x standard in the API. Price: $60 per million input tokens, $300 per million output. Launches for GPT-6 Astra immediately; Sol support to follow.

The target is latency-sensitive applications: real-time voice, interactive code generation, streaming UIs where waiting 200 milliseconds is a user experience problem. At 6x the price, this is a specialty product, not a volume tier. For the applications where response latency is the differentiating experience factor, the math works differently than it does for batch processing.

The cross-lab read: Anthropic's Sonnet 5.5 efficiency gains (30% faster outputs) come from doing the same work in fewer tokens, not from premium compute. OpenAI's ultrafast tier comes from premium compute, at premium prices. These are different product philosophies about how to compete on speed. One is cheaper by design; the other is faster by configuration. Neither is wrong.

Anthropic, Claude Code, September 28
v2.1.283: Enterprise Instrumentation for Gateway Deployments

Three new managed settings for teams running Claude Code at enterprise scale. The x-claude-code-prompt-id gateway hint header groups LLM requests by user prompt, enabling cost attribution and debugging at the session level rather than the individual API-call level. Opt in via CLAUDE_CODE_GATEWAY_HINT_HEADERS=1.

Two new model governance settings ship alongside: availableModelsMatch: "exact" blocks undeclared model versions from slipping through approved lists; deniedModels blacklists specific models even when availableModels would otherwise permit them. The combination gives enterprise operators fine-grained, auditable control over which model version is running in production.

The pattern: this is the third consecutive Claude Code release expanding enterprise control surface, following SCIM and managed settings work from earlier in September. Teams that have been patching around model governance in Claude Code deployments now have first-class tooling for it. The blast radius is any organization running Claude Code at scale where model version compliance is a real requirement, which at enterprise level is most of them.

Meta AI, September 28 to 29
Meta Enterprise Platform and Muse for Small Business: The Two-Sided Play

Meta shipped on both ends of the enterprise market in one day. At the top: the Meta Enterprise Platform, combining Muse, Meta Business Agent, the Muse API, and Muse Code under a single umbrella. Chirantan "CJ" Desai joined as Chief Enterprise Platform Officer. The underlying Muse Spark 1.3 model uses roughly 20% fewer tool calls and 25% fewer tokens on longer tasks than its predecessor.

The Desai hire is the signal. Meta has been building AI products; it is now building an AI business unit. Chief Enterprise Platform Officer is a new title at Meta. It signals that Muse is no longer a product team inside Facebook; it has its own P&L ambition. Whether that ambition materializes depends on whether business owners will pay for AI that runs their operations, or continue patching together free tiers. Meta's distribution gives it a starting point few competitors have: the install base is already there.

At the other end: Muse expanded to small businesses, connecting to Canva, QuickBooks, Shopify, Slack, Stripe, Instagram professional accounts, Facebook Pages, and Meta ad accounts. Zuckerberg described it as a "digital operations layer." Free tier with usage limits; paid subscription for higher usage. The beachhead is different from OpenAI's Dots: OpenAI's entry point is the knowledge worker with a ChatGPT subscription, Meta's is the business owner with an ad account. Different starting lines, the same claim on the background-agent market. Neither has proven the failure mode yet.

Anthropic, SDK Releases, September 28
Between-Tools Thinking and Sonnet 5.5 Across All SDKs

Same-day as the Sonnet 5.5 launch, Anthropic shipped coordinated SDK updates across five packages: Python SDK v1.9.0, TypeScript SDK v0.129.0, and matching version bumps to the Vertex, Bedrock, and Foundry SDKs. All five add claude-sonnet-5-5 to their model enum. All five add between_tools as a new thinking type in extended thinking configurations.

The between_tools thinking type is the mechanism worth noting. Extended thinking now supports a reasoning mode that fires between tool calls in an agentic sequence, not just at the start of a response. This means an agent can reason mid-flow, after seeing the result of one tool, before deciding the next action. The blast radius: every agentic loop using extended thinking gets smarter decision-making at tool boundaries without changing the API call structure. Update the SDK and the capability is there.

Signal
Quiet
on the Wire

Google DeepMind was quiet in this window. Gemini 3.8 Flash, AlphaGenome Atlas, and WeatherNext 3 all shipped earlier in September. The lab appears to be between launches; expect the next Google model move before end of quarter.

Mistral closed the Pimento acquisition on September 23 (the Paris adtech startup, 12.7 million euros). No product news in this window.

xAI published a Grok Bot deployment case study covering SpaceXAI's internal use: 20,000 daily feedback points synthesized into engineering themes, automated issue routing via Linear and Datadog, video reproduction of reported bugs, and coaching feedback on both human and bot agent interactions. Grok 4.7 shipped September 21; today's piece is deployment documentation, not a new capability.

The two-lab pricing convergence is the structural story the quiet labs underscore. When two of the six labs independently set a price on the same calendar day, the others face a positioning question. None of them have announced a competing price point since Sonnet 5.5 and Sol landed this morning. That will change, probably soon.

Worth watching across labs: the always-on agent paradigm is arriving from two directions simultaneously. OpenAI's Dots and Meta's Muse both claim the background-agent market. Neither has published meaningful failure-mode data. That gap will be the story of Q4.

The Close
Anthropic shipped a benchmark.
OpenAI threw a party.
The price was the same. Pick the model that survived your evals.
●
Reference

Release
Log

Every confirmed item from the Sep 28 to Sep 29 window, grouped by category.
Models
2 releases
Two mid-tier flagship launches on the same day at the same price.
MODEL
Claude Sonnet 5.5 (Anthropic)
Second model in the Claude 5.5 family. Scores 70.6% on Terminal-Bench 4.0 (vs. Sonnet 5's 10.3%), generates outputs 30%+ faster, and costs up to 30% less per task at identical per-token pricing. First Sonnet model deployed with cyber safeguards at the level previously reserved for Opus 5.5. Available as claude-sonnet-5-5 on API, AWS Bedrock, Google Cloud Vertex, and Azure Foundry. System card published same day.
How to use Set model: "claude-sonnet-5-5". Pricing: $2/1M input, $10/1M output (same as prior Sonnet). Run existing Opus 5.5 evals; audit aggregate token counts on long agentic runs to verify the per-task efficiency gain.
Why it matters The agentic coding capability that drove teams to Opus 5.5 now lives at Sonnet price, with Opus-level safety. Two separate procurement decisions just collapsed into one.
MODEL
GPT-6.1 Sol (OpenAI)
New OpenAI model delivering near-GPT-6 Astra capability at one-fifth Astra's token prices ($2/1M input, $10/1M output vs. Astra's $10/$50). Available as gpt-6.1-sol via API and in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users.
How to use Set model: "gpt-6.1-sol" in OpenAI API calls. Available in ChatGPT Work and Codex tiers immediately.
API and Platform
1 release
API
OpenAI Ultrafast Speed Tier
Premium inference tier delivering 300 tokens/second in Codex (8x standard) and up to 6x standard in the API. Priced at $60/1M input, $300/1M output (6x standard rates). Launches for GPT-6 Astra immediately; Sol support to follow. Designed for latency-sensitive applications: real-time voice, interactive code generation, streaming UIs.
How to use Request the ultrafast tier via the OpenAI API or Codex interface. Available for GPT-6 Astra models at launch; GPT-6.1 Sol support to follow.
Apps
2 releases
APPS
OpenAI Dots: Always-On Agents
New class of persistent AI agents inside ChatGPT. Each dot runs on GPT-6 Astra, gets its own cloud computer and browser, works toward user-assigned goals continuously, learns from feedback over time, and connects to 4,000+ apps. Available at no extra cost to Pro, Business Premium, and Enterprise users (one dot per eligible user). Reachable via Slack and Teams.
How to use Available in ChatGPT for Pro, Business Premium, and Enterprise subscribers. Assign a goal; connect to Slack or Teams for background operation.
APPS
Meta Muse for Small Business
Muse AI agent expanded to small businesses. Connects to Canva, QuickBooks, Shopify, Slack, Stripe, Instagram professional accounts, Facebook Pages, and Meta ad accounts. Free tier with usage limits; paid subscription for higher usage. Positioned as a "digital operations layer" for small business owners.
Claude Code
1 release
CODE
Claude Code v2.1.283
Three new enterprise managed settings. Gateway hint header x-claude-code-prompt-id groups LLM requests per user prompt for cost attribution and debugging (opt-in via CLAUDE_CODE_GATEWAY_HINT_HEADERS=1). New availableModelsMatch: "exact" option blocks undeclared model versions. New deniedModels setting blacklists specific models even when availableModels would allow them.
How to use Set CLAUDE_CODE_GATEWAY_HINT_HEADERS=1 to enable prompt-ID grouping. Configure availableModelsMatch and deniedModels in managed settings for enterprise model governance.
Agent SDKs
2 releases
SDK-PY
Anthropic Python SDK v1.9.0
Added between_tools thinking type support and claude-sonnet-5-5 to the model enum.
How to use pip install anthropic==1.9.0. Use thinking_type="between_tools" in extended thinking configurations; use "claude-sonnet-5-5" as the model identifier.
SDK-TS
Anthropic TypeScript SDK v0.129.0, plus platform SDKs
Added between_tools thinking type and claude-sonnet-5-5 to the model type union. Same-day additions to Vertex SDK (v0.20.0), Bedrock SDK (v0.34.0), and Foundry SDK (v0.5.0).
How to use npm install @anthropic-ai/sdk@0.129.0. Update platform SDKs to matching versions for Vertex, Bedrock, and Foundry. claude-sonnet-5-5 available in the model type union immediately.
News and Partnerships
2 items
NEWS
Meta Enterprise Platform Launch
Meta launched the Meta Enterprise Platform combining Muse, Meta Business Agent, the Muse API, and Muse Code. Chirantan "CJ" Desai joined as Chief Enterprise Platform Officer. Underlying Muse Spark 1.3 model uses roughly 20% fewer tool calls and 25% fewer tokens on longer tasks than the prior version.
NEWS
SpaceXAI: Grok Bot in Customer Support (xAI case study)
xAI published a case study on SpaceXAI's internal Grok Bot deployment. The system synthesizes 20,000+ daily feedback points into engineering themes, creates and routes tickets via Linear and Datadog, reproduces reported issues with video, and provides coaching feedback on human and bot agent interactions, without adding headcount.
Stay on the frontier

Get Shipped. in your inbox.

Daily digest at 9 PM ET. Weekly magazine every Friday morning. Six labs, one feed. No spam, one-click unsubscribe.