Anthropic's Sonnet 5.5 shipped Tuesday morning. The Terminal-Bench 4.0 number is 70.6%, against Sonnet 5's 10.3%. A gap that wide does not describe an improvement; it describes a threshold crossed. The tasks that required Opus, Sonnet can now handle. The things you were paying the expensive model to do, the mid-tier model can do now.
OpenAI was at Fort Mason for DevDay 2026. GPT-6.1 Sol launched at the same price point: $2 per million tokens in, $10 out. Dots launched as always-on agents with their own cloud computers and 4,000-plus app integrations. An ultrafast inference tier launched at 6x standard rates for latency-sensitive workloads. Three announcements, one event, one carefully chosen Tuesday in San Francisco.
Meta launched an enterprise platform and expanded its Muse agent to small businesses the same day. The frontier on September 29 is not a quiet Tuesday. It is the mid-tier repricing itself, and two labs arriving at identical price points with completely different arguments for why you should choose them.
The number is 70.6%.
Sonnet 5's Terminal-Bench 4.0 score was 10.3%. That is not a gap you close with incremental tuning. It describes two different tools wearing the same name. The Sonnet tier handled the fast, clean tasks. Opus handled the agentic coding work, the multi-step loops, the long-horizon reasoning, the things that required the expensive model to not make a mess. Sonnet and Opus were not interchangeable; they were complementary.
Sonnet 5.5 crosses that line. At 70.6% on Terminal-Bench 4.0, the capability that drove teams to Opus 5.5 now lives at Sonnet price. The mechanism is efficiency: Sonnet 5.5 generates outputs 30% faster and costs up to 30% less per task, at identical per-token pricing ($2/$10), because it completes the same work in fewer tokens. For agentic loops running hundreds of completions a day, 30% fewer tokens per task compounds. A team running 10,000 completions a day is looking at a different monthly bill than they were yesterday.
The second number in the launch is quieter and more consequential for enterprise buyers. Sonnet 5.5 is the first Sonnet model deployed with the cyber safeguards previously reserved for Opus 5.5. For compliance teams that needed Opus's safety pedigree, this changes a procurement question: the model they now need is not the most expensive one. That conversation opened for the first time today.
The cross-lab contrast is hard to ignore. GPT-6.1 Sol launched the same day, same price point, same pitch: near-flagship capability at mid-tier price. OpenAI's framing is "near-GPT-6 Astra intelligence for one-fifth Astra's cost." Anthropic's framing is a Terminal-Bench number and a system card. Both are talking to the same developer, the one running production agentic loops who wants to stop paying Opus or Astra prices. One came with a conference in San Francisco. The other came with a benchmark. The builder who has to choose between them this week will read the benchmark and watch the keynote replay at 2x speed. Then they will run evals, because this is not a decision you make from the press release.
The pricing parity itself deserves a moment. Two major labs launched competing models at identical price points on the same calendar day. Whether this is market coordination or parallel convergence, $2/$10 per million tokens is now the reference price for capable mid-tier models in late 2026. Launching above that price requires a flagship justification. Launching below it means racing on margin. The labs that have not yet landed at $2/$10 will. The ones that have are now competing entirely on argument.
Builder's move: update to claude-sonnet-5-5. Available now on API, AWS Bedrock, Google Cloud Vertex, and Microsoft Azure Foundry. Run your existing Opus 5.5 evals against it. The 30%-per-task efficiency claim is the one to verify against your own workload. Aggregate token counts across a full agentic run, not just latency on a single call.
The model story at DevDay 2026 was GPT-6.1 Sol: near-Astra capability at $2/$10 per million tokens, available as gpt-6.1-sol via API and in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users. It is a direct market position against Sonnet 5.5, priced identically, aimed at the same developer. But the more interesting launch was Dots.
Each dot is an always-on agent: its own cloud computer, its own browser, a goal you assign, a connection to 4,000-plus apps, and a feedback loop that teaches it over time. No explicit invocation required. The dot runs while you sleep. Available to Pro, Business Premium, and Enterprise users at no extra charge, one dot per eligible seat. Reachable via Slack and Teams for background operation.
The mechanism is persistence. Today's agents are stateless between invocations. Dots are stateful over days. That distinction has real implications for both capability and trust. A stateful agent that ran toward a goal for 72 hours and made hundreds of small decisions along the way is harder to audit than one that was invoked, answered, and dismissed. OpenAI's framing is about the learning feedback loop; the transparency question is the harder one, and the DevDay materials leave it aspirational.
The contrast with Anthropic: Sonnet 5.5 is a capability upgrade for models a developer explicitly invokes. Dots is a paradigm for agents that run without being asked. They are not competing products. They may well be the same builder's next two purchases, in that order.
OpenAI's new premium inference tier delivers 300 tokens per second in Codex (8x the standard rate) and up to 6x standard in the API. Price: $60 per million input tokens, $300 per million output. Launches for GPT-6 Astra immediately; Sol support to follow.
The target is latency-sensitive applications: real-time voice, interactive code generation, streaming UIs where waiting 200 milliseconds is a user experience problem. At 6x the price, this is a specialty product, not a volume tier. For the applications where response latency is the differentiating experience factor, the math works differently than it does for batch processing.
The cross-lab read: Anthropic's Sonnet 5.5 efficiency gains (30% faster outputs) come from doing the same work in fewer tokens, not from premium compute. OpenAI's ultrafast tier comes from premium compute, at premium prices. These are different product philosophies about how to compete on speed. One is cheaper by design; the other is faster by configuration. Neither is wrong.
Three new managed settings for teams running Claude Code at enterprise scale. The x-claude-code-prompt-id gateway hint header groups LLM requests by user prompt, enabling cost attribution and debugging at the session level rather than the individual API-call level. Opt in via CLAUDE_CODE_GATEWAY_HINT_HEADERS=1.
Two new model governance settings ship alongside: availableModelsMatch: "exact" blocks undeclared model versions from slipping through approved lists; deniedModels blacklists specific models even when availableModels would otherwise permit them. The combination gives enterprise operators fine-grained, auditable control over which model version is running in production.
The pattern: this is the third consecutive Claude Code release expanding enterprise control surface, following SCIM and managed settings work from earlier in September. Teams that have been patching around model governance in Claude Code deployments now have first-class tooling for it. The blast radius is any organization running Claude Code at scale where model version compliance is a real requirement, which at enterprise level is most of them.
Meta shipped on both ends of the enterprise market in one day. At the top: the Meta Enterprise Platform, combining Muse, Meta Business Agent, the Muse API, and Muse Code under a single umbrella. Chirantan "CJ" Desai joined as Chief Enterprise Platform Officer. The underlying Muse Spark 1.3 model uses roughly 20% fewer tool calls and 25% fewer tokens on longer tasks than its predecessor.
The Desai hire is the signal. Meta has been building AI products; it is now building an AI business unit. Chief Enterprise Platform Officer is a new title at Meta. It signals that Muse is no longer a product team inside Facebook; it has its own P&L ambition. Whether that ambition materializes depends on whether business owners will pay for AI that runs their operations, or continue patching together free tiers. Meta's distribution gives it a starting point few competitors have: the install base is already there.
At the other end: Muse expanded to small businesses, connecting to Canva, QuickBooks, Shopify, Slack, Stripe, Instagram professional accounts, Facebook Pages, and Meta ad accounts. Zuckerberg described it as a "digital operations layer." Free tier with usage limits; paid subscription for higher usage. The beachhead is different from OpenAI's Dots: OpenAI's entry point is the knowledge worker with a ChatGPT subscription, Meta's is the business owner with an ad account. Different starting lines, the same claim on the background-agent market. Neither has proven the failure mode yet.
Same-day as the Sonnet 5.5 launch, Anthropic shipped coordinated SDK updates across five packages: Python SDK v1.9.0, TypeScript SDK v0.129.0, and matching version bumps to the Vertex, Bedrock, and Foundry SDKs. All five add claude-sonnet-5-5 to their model enum. All five add between_tools as a new thinking type in extended thinking configurations.
The between_tools thinking type is the mechanism worth noting. Extended thinking now supports a reasoning mode that fires between tool calls in an agentic sequence, not just at the start of a response. This means an agent can reason mid-flow, after seeing the result of one tool, before deciding the next action. The blast radius: every agentic loop using extended thinking gets smarter decision-making at tool boundaries without changing the API call structure. Update the SDK and the capability is there.
Google DeepMind was quiet in this window. Gemini 3.8 Flash, AlphaGenome Atlas, and WeatherNext 3 all shipped earlier in September. The lab appears to be between launches; expect the next Google model move before end of quarter.
Mistral closed the Pimento acquisition on September 23 (the Paris adtech startup, 12.7 million euros). No product news in this window.
xAI published a Grok Bot deployment case study covering SpaceXAI's internal use: 20,000 daily feedback points synthesized into engineering themes, automated issue routing via Linear and Datadog, video reproduction of reported bugs, and coaching feedback on both human and bot agent interactions. Grok 4.7 shipped September 21; today's piece is deployment documentation, not a new capability.
The two-lab pricing convergence is the structural story the quiet labs underscore. When two of the six labs independently set a price on the same calendar day, the others face a positioning question. None of them have announced a competing price point since Sonnet 5.5 and Sol landed this morning. That will change, probably soon.
Worth watching across labs: the always-on agent paradigm is arriving from two directions simultaneously. OpenAI's Dots and Meta's Muse both claim the background-agent market. Neither has published meaningful failure-mode data. That gap will be the story of Q4.
claude-sonnet-5-5 on API, AWS Bedrock, Google Cloud Vertex, and Azure Foundry. System card published same day.model: "claude-sonnet-5-5". Pricing: $2/1M input, $10/1M output (same as prior Sonnet). Run existing Opus 5.5 evals; audit aggregate token counts on long agentic runs to verify the per-task efficiency gain.gpt-6.1-sol via API and in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users.model: "gpt-6.1-sol" in OpenAI API calls. Available in ChatGPT Work and Codex tiers immediately.x-claude-code-prompt-id groups LLM requests per user prompt for cost attribution and debugging (opt-in via CLAUDE_CODE_GATEWAY_HINT_HEADERS=1). New availableModelsMatch: "exact" option blocks undeclared model versions. New deniedModels setting blacklists specific models even when availableModels would allow them.CLAUDE_CODE_GATEWAY_HINT_HEADERS=1 to enable prompt-ID grouping. Configure availableModelsMatch and deniedModels in managed settings for enterprise model governance.between_tools thinking type support and claude-sonnet-5-5 to the model enum.pip install anthropic==1.9.0. Use thinking_type="between_tools" in extended thinking configurations; use "claude-sonnet-5-5" as the model identifier.between_tools thinking type and claude-sonnet-5-5 to the model type union. Same-day additions to Vertex SDK (v0.20.0), Bedrock SDK (v0.34.0), and Foundry SDK (v0.5.0).npm install @anthropic-ai/sdk@0.129.0. Update platform SDKs to matching versions for Vertex, Bedrock, and Foundry. claude-sonnet-5-5 available in the model type union immediately.Daily digest at 9 PM ET. Weekly magazine every Friday morning. Six labs, one feed. No spam, one-click unsubscribe.