Shipped. Daily, Six Frontier Labs
Shipped.
One lab paused the frontier; the other published the prompts.
Edition Daily, Frontier Date Wednesday, August 19, 2026 Window Aug 18 to Aug 19 Beat Six frontier labs
The Open
six labs, one window
The framework fired. The train stopped.

The Preparedness Framework was always criticized as convenient. A table of thresholds OpenAI built to justify moving fast, because listing the things that would slow them down made speed look responsible. Two years of benchmarks, evaluations, and escalating capability tiers. Zero times the Critical row had been checked. Until Tuesday.

An unreleased model named Astra, evaluated with safety restrictions deliberately switched off to gauge raw offensive cybersecurity capability, showed performance that OpenAI could not rule out as Critical under its own rubric. The ability to find and exploit zero-days in hardened real-world systems without human assistance. Reinforcement-learning training on deployment-bound models stopped. Not the lab's discretion, the framework's logic.

Four time zones south on the Embarcadero, also on Tuesday, Anthropic published a paper: Claude Opus 4.8 and Mythos Preview, given a single protocol prompt and zero additional guidance, ran 24-to-48-hour autonomous campaigns against 16 protein targets. The researchers granted access and watched from a distance. The paper includes the prompts. The capability is dual-use. Both things are true at once, and Anthropic said so.

Lead01
OpenAI

The Train
Stopped

OpenAI's Astra model tripped the Critical cybersecurity threshold in its own Preparedness Framework, pausing frontier training and raising the question every lab has been avoiding.
Lab OpenAIArea News, PolicySource: openai.com/index/pacing-model-development
By the numbers Threshold: Critical
Training pause: 2 weeks (RL)
Model: Astra (unreleased)
Frontier RL run: on hold
Framework: Preparedness v2
OpenAI, Aug 18

The test that triggered the pause was deliberately adversarial. OpenAI researchers switched off safety restrictions to measure Astra's raw offensive cybersecurity capability, the same way you assess a car's top speed on a closed track. What came back was a model that OpenAI "cannot rule out" has reached the Critical threshold: the ability to identify and develop functional zero-day exploits against hardened real-world critical systems, at scale, without human assistance. That is the top of the Preparedness Framework's four-tier cyber ladder. Everything below it was acceptable to ship. Astra may sit above it.

The response was fast. Reinforcement-learning training on deployment-bound models stopped on August 7, two weeks before Tuesday's public disclosure. As of the announcement, the largest planned frontier RL run remains on hold. OpenAI is now running smaller-scale evaluations, assessing model behavior, and expanding monitoring infrastructure before deciding whether to resume. The company is also rewriting the Preparedness Framework itself, adding earlier safeguards and stronger monitoring checkpoints before the next scaling run.

The mechanism matters here. The pause was not a discretionary call by leadership deciding to slow down. It was the framework's own logic, triggered by the framework's own evaluation criteria. OpenAI built a document with a row labeled Critical and a consequence labeled Stop. Tuesday was the first time that row got checked in two years of operation. Whether that is a story about the framework working exactly as designed, or about capability advancing faster than the framework anticipated, depends on which version of this you find reassuring.

OpenAI's Preparedness Framework is no longer a marketing document. That took two years and a model named Astra. The framework did what it was built to do. That is either the best possible outcome of governance-by-document, or the first time the document admitted the gap between its ambitions and the pace of the work.

The builder's move is unglamorous. If you have delivery dates tied to upcoming OpenAI model releases, your timeline now carries an unresolved variable. OpenAI has not said when the frontier RL run resumes, what Astra will look like when it does, or when any of this will translate into API availability. "Smaller-scale evaluations" does not have a date attached. Meanwhile, xAI's Grok 4.6 arrived on Amazon Bedrock the next morning. The table is not paused.

Dig02
Anthropic Research

The Researcher
Walked Away

Claude Opus 4.8 and Mythos Preview ran autonomous binder design campaigns against 16 protein targets. No human input after the prompt. The paper ships the methodology.
Lab AnthropicArea ResearchSource: anthropic.com/research/Claude-accelerates-protein-design
Anthropic, Aug 18

The procedure was simple. Write a single protocol prompt specifying no epitope, no scaffold, no sequence for any target. Hand it to Claude Opus 4.8 and Mythos Preview. Walk away. The models ran 24-to-48-hour campaigns against 16 protein targets, making all design decisions autonomously, from target selection to final binder candidate. The researchers' only involvement after initiating the run was granting access approvals and monitoring infrastructure. They came back to results.

Protein binder design is a discipline where a skilled human operator typically needs weeks to work a single target. The paper documents campaigns running in parallel across 16 targets without supervision. This is the blast radius: the time cost of a capability that has historically constrained drug discovery and biomedical research. It also expands the accessible attack surface for anyone who wants to use the same system for harm. Anthropic says this clearly in the paper. The capability is dual-use. The dual-use framing does not change the capability.

Releasing the prompts is a choice. Anthropic's position is that scientific transparency requires publishing the methodology, that withholding the prompts would obscure the risk rather than reduce it, and that the research community needs the full picture to understand what is now achievable. That is a defensible position. It is also a very different answer than the one OpenAI gave on the same Tuesday, in a different domain, to a structurally similar question about a capability that exceeded expectations.

If you are working in drug discovery or protein engineering, the methodology is in the paper. The campaign ran on Claude Science infrastructure, using the long-running Claude framework for scientific computing that Anthropic has been building this year. If you are working on AI biosecurity, August 18 is a date to mark.

Also Shipped
three more moves from the window
xAI, Aug 19
Grok 4.6 lands on Amazon Bedrock

Grok 4.6 is now available in all AWS Regions through Amazon Bedrock. The model carries a 500,000-token context window and configurable reasoning levels (low, medium, high, xhigh). Pricing: $2 per million input tokens, $6 per million output tokens. Enterprise security, monitoring, cross-region inference included by default.

This is the third xAI enterprise move in nine days: Grok Bot (AI teammates) launched August 11, Grok 4.6 as a standalone model August 12, Bedrock availability August 19. The pattern is distribution-first, the same playbook Anthropic and OpenAI have both run. Launch the model, land the partners, deploy through clouds. Bedrock gives xAI access to enterprise teams already running AWS infrastructure without asking them to change a procurement line.

The $2/$6 pricing is aggressive. It undercuts GPT-5.6 Sol at comparable context lengths and positions close to Claude Opus tiers. With OpenAI's Astra deployment timeline now uncertain, xAI is occupying ground. The timing is clean. Builder's move: Available now via Bedrock Converse API or InvokeModel. Standard IAM model-access permissions required. Model ARN in the AWS Bedrock model card.

OpenAI, Aug 18
ChatGPT for Teens

ChatGPT for Teens rolled out globally to eligible accounts (ages 13 to 17) on Free and paid personal plans. New system-level guardrails: no romantic language, no claims of feelings or consciousness, content restrictions on self-harm, eating disorders, and violence. Study Mode intercepts apparent homework shortcuts and redirects to guided, step-by-step problem-solving. Ninety-minute usage reminders included. Parental controls and age verification follow.

The irony of the timing is not subtle. OpenAI shipped teenager guardrails the same day it disclosed a model potentially too dangerous to ship to adults. The Model Spec update that accompanies the launch rewrites how ChatGPT should behave with minors. Developers building education products on the API should read it; the guidance will affect model behavior whether or not you are using the teen-specific product. Primary source.

Anthropic, Aug 18 to 19
Claude Code, two-day update

Two days of Claude Code updates landed this window. August 18 shipped GitLab MR badge integration in the footer and statusline: repos with a GitLab remote and an authenticated glab CLI now show the current MR number with draft, pending, or green state at a glance. Also August 18: the optional CLAUDE_CODE_PROJECT_DIR_NAME env var for per-project transcript directory naming, automatic session continuation at usage limits, faster session startup, improved memory and CPU usage during background cloud sessions, and expanded Remote Control syncing.

August 19 added ANTHROPIC_DEFAULT_MODEL, which sets the starting model for new sessions without touching per-project config (a /model pick still overrides it and persists across restarts). Also August 19: notify_when_idle for cross-session SendMessage, letting you ask another Claude Code session on the same machine to send one notification when it next goes idle. Opt-in, one-shot, macOS and Linux. Builder's move: Run glab auth login in a GitLab-backed repo and the badge appears. Set ANTHROPIC_DEFAULT_MODEL=claude-opus-5 in your shell profile to pin a default model.

Signal
Quiet on
the Wire

Google DeepMind was quiet this window. Koray Kavukcuoglu is three weeks into the SVP role after Demis Hassabis stepped to an advisory position; no product release has followed the transition yet. The new leadership's first major move will set tone for the lab's second half of 2026.

Mistral's Les Ulis inference datacenter (10MW, dedicated to inference operations) is on track for a Q3 opening per prior announcements. Leanstral 1.5, Mistral's formal proof engineering model, is in development. Neither item has a confirmed August date.

OpenAI's Ultrafast mode for GPT-5.6 Sol, powered by Cerebras at up to 14x standard speed, remains in limited preview from the August 13 announcement. Expansion is expected but undated. The only question that matters right now is the one OpenAI has not answered: when does Astra ship, and what does it look like when the safety work is done.

The Close
OpenAI's Preparedness Framework fired. First time in two years of thresholds.
Anthropic published the prompts. Also first time for that.
Same Tuesday. Not the same answer.
Aug 18 to Aug 19, 2026

Release
Log

Every confirmed release in the window, grouped by lab. Items that did not earn prose live here as one-liners.
Anthropic
3 releases
Research and tooling from the anchor lab. The protein paper and two days of Claude Code.
research
Autonomous de novo protein binder design with Claude
Claude Opus 4.8 and Mythos Preview ran 24 to 48-hour autonomous binder design campaigns against 16 protein targets using a single protocol prompt, with no human input into any design decision. Researchers only granted access approvals and monitored infrastructure. The paper and prompts are published.
Why it matters First Anthropic research result documenting autonomous biological design at campaign scale, with the methodology prompts in the public release. Dual-use capability explicitly flagged by the authors.
code
Claude Code, Aug 18 update
GitLab MR badge in footer and statusline for repos with a GitLab remote and an authenticated glab CLI. Optional CLAUDE_CODE_PROJECT_DIR_NAME env var for per-project transcript directory naming. Automatic session continuation at usage limits. Faster and safer sessions. Improved memory and CPU usage during background cloud sessions. Expanded Remote Control syncing.
How to use Run glab auth login in any GitLab-backed repo and the MR badge appears automatically in the statusline.
code
Claude Code, Aug 19 update
ANTHROPIC_DEFAULT_MODEL env var sets the default model for new sessions; a /model pick still overrides it and persists across restarts. notify_when_idle added to cross-session SendMessage: ask another Claude Code session on the same machine to send one notice when it next goes idle. Opt-in, one-shot, macOS and Linux only.
How to use Set ANTHROPIC_DEFAULT_MODEL=claude-opus-5 in your shell profile to pin a starting model without touching per-project config.
OpenAI
2 releases
The lab with the frozen training run had a full news day. A safety disclosure and a consumer product in the same 24 hours.
news
Pacing model development in an era of cyber-critical capabilities
OpenAI disclosed that its unreleased Astra model may have reached the Critical cybersecurity threshold under its Preparedness Framework, triggering a two-week pause in deployment-focused RL training (effective August 7). The largest planned frontier RL run remains on hold. Smaller-scale evaluations are running to assess model behavior. The Preparedness Framework is being rewritten with stronger monitoring and earlier safeguard checkpoints. A companion post details the broader safety infrastructure changes.
Why it matters First time OpenAI's safety framework has halted an active training run in two years of operation. The Critical threshold is defined as the ability to identify and develop functional zero-day exploits against hardened systems at scale without human intervention.
apps
ChatGPT for Teens
Global rollout to eligible accounts (ages 13 to 17) on Free and paid personal plans. System-level guardrails: no romantic language, no claims of feelings or consciousness, restricted content categories covering self-harm, eating disorders, and violence. Study Mode redirects apparent homework shortcuts to guided step-by-step problem solving. Ninety-minute usage reminders. Parental controls available. Companion Model Spec update covers teen-appropriate response guidelines.
How to use Available automatically to qualifying teen accounts. API developers building education products: read the Model Spec update, as guidelines apply to API behavior with minors regardless of product.
xAI
1 release
Grok 4.6 expanding into AWS infrastructure one week after its standalone launch.
news
Grok 4.6 on Amazon Bedrock
Grok 4.6 available in all AWS Regions via Amazon Bedrock. 500,000-token context window. Configurable reasoning levels: low, medium, high, xhigh. Pricing: $2 per million input tokens, $6 per million output tokens. Enterprise security, monitoring, logging, and cross-region inference included.
How to use Invoke via Bedrock Converse API or InvokeModel. Standard IAM model-access permissions required. Model ARN and full specs in the Bedrock model card.
Stay on the frontier

Get Shipped. in your inbox.

Daily digest at 9 PM ET. Weekly magazine every Friday morning. Six labs, one feed. No spam, one-click unsubscribe.