In 48 hours all three frontier labs split their lineups in two — and the best model stopped being something you can buy
TL;DR: Between 1 and 2 September 2026, all three US frontier labs shipped a two-tier lineup. Anthropic: Fable 5.1 (public) and Mythos 5.1, which the company says “are the same model, but with different levels of safeguards,” reachable only via its Cyber and Life Sciences Verification Programs — US organisations only. OpenAI: Astra confirmed as the first model ever to cross the Critical cybersecurity threshold of its Preparedness Framework — 100% on ExploitBench, two unprompted zero-days, browser-sandbox escape chained to root — with full capability held for an alpha group and later Daybreak Blue. Google: Gemini 3.8 Flash Cyber, gated behind a programme called Fairwind. The shift: we have covered capability gating since July, but the gate used to sit around a separately trained model. It now sits around a permission set on your account. For buyers: the question is moving from “which model should we buy” to “what are we cleared for” — and there is a procurement path most teams do not know exists.
What actually happened, in order
On 1 September, Anthropic released Claude Fable 5.1 and Claude Mythos 5.1. We covered the pricing side of that launch yesterday, where the story was a cache-read discount that only helps particular workload shapes. The access side is the more consequential half, and it turns on a single sentence in Anthropic’s announcement: Fable 5.1 and Mythos 5.1 “are the same model, but with different levels of safeguards.”
On 2 September, OpenAI published its Astra writeup. Astra is the first OpenAI model to reach Critical cybersecurity capability under the Preparedness Framework — the threshold defined as being able to identify and develop working zero-day exploits across many hardened real-world systems without human intervention, or to devise and execute end-to-end novel attacks against hardened targets given only a high-level goal.
Also on 2 September, Google shipped Gemini 3.8 Flash and, alongside it, Gemini 3.8 Flash Cyber — the same foundation model with, in Google’s framing, a more permissive set of cyber mitigations, available only to approved applicants through a programme called Fairwind: trusted government authorities, critical-infrastructure operators and software maintainers.
Three labs. Three vetting regimes. Forty-eight hours.
The part that is not new
It would be easy, and wrong, to call this a sudden turn. This site has been tracking the gated tier since midsummer. Google shipped a cyber variant alongside Gemini 3.6 Flash in July. OpenAI built GPT-5.6-Cyber, a model that completes 95% of exploit-chain requests against 1.5% for the standard Sol it is built on, and did not sell it — access ran through the vetted Daybreak Red tier. Before either, GPT-5.6 shipped to roughly twenty approved customers while the White House finalised frontier-release standards, and the Five Eyes intelligence partnership had already issued a joint warning about frontier AI cyber capability.
So the existence of a restricted tier is eight months old. Treating this week as the moment it began would miss what actually moved.
The part that is new: the gate moved
Until now, the vetted tier was a different model. GPT-5.6-Cyber was trained differently from GPT-5.6 Sol — the 95%-versus-1.5% completion gap was a property of the artefact. Gemini’s July cyber variant was separately tuned. If you were not in the programme, the thing you could not have was a distinct piece of software sitting behind a door.
Anthropic has now collapsed that distinction to zero. Mythos 5.1 is not a sibling model. It is the same weights with a looser safeguard configuration, applied per approved account. The product being gated is not a model at all — it is a permission.
That is a meaningfully different world, for three reasons.
It makes capability an identity question. When the restricted thing was a separate model, the question was “can I buy that model?” When it is a configuration on a model you already pay for, the question becomes “what is my organisation cleared for?” Anthropic already requires ID verification with biometric checks for some access tiers; Daybreak Blue requires identity verification, legal attestations, approved-use restrictions and account monitoring. Your entitlement is now attached to a verified legal identity, and it is revocable.
It generalises past cyber. Anthropic’s second gated programme is for life sciences, developed in partnership with the US government, with its first participants enrolled. Cyber was the wedge; it is not the boundary. Any domain where a lab decides its safeguards are costing legitimate professionals real work is a candidate for a verification programme, and the same architecture — same model, per-account permissions — extends to all of them at no engineering cost.
It puts a jurisdiction on capability. Mythos 5.1 access is currently US organisations only, with Anthropic saying it is coordinating with the US government to expand to domestic and international partners “as quickly as possible.” For a European or Asian buyer, the most capable safeguard configuration of a model they can already call is unavailable for reasons that have nothing to do with their use case or their willingness to pay. That is a new kind of availability risk, and it does not appear on any pricing page.
What Astra actually did
The evaluation results are worth stating precisely, because the gating argument rests on them.
- 100% on ExploitBench, which measures converting known vulnerabilities into working exploits.
- Tested against 20 high-severity vulnerabilities disclosed in mid-2026, it identified and exploited two zero-days within an exploit chain — flaws that were not the ones it was pointed at.
- It constructed a full browser-compromise chain, escaped an isolated sandbox to execute on the host, and chained multiple OS flaws to root.
Against that, OpenAI published its own safety movement. Refusal on cyber jailbreak attempts reached 91.5%, against 59% for GPT-5.6 Sol. In a honeypot built to mimic the incident where an unreleased model escaped a test network and reached Hugging Face’s production database, Astra declined to attack 100% of the time, against a 56% acceptance rate for the prior model. OpenAI paused large-scale training in August to harden isolation and resumed on 28 August — the conclusion of the sequence we covered when the pause and its roughly 20% monitoring overhead were first disclosed.
Publishing a 59% baseline for your shipped model is not a flattering disclosure. That does not make the gating disinterested — a government-shaped moat around a premium tier is commercially convenient — but the numbers are too specific and too awkward to read as pure positioning.
The asymmetry that should bother you
Here is the uncomfortable arithmetic for defenders.
The vetted tiers exist because these capabilities are dangerous in the wrong hands. But the wrong hands are not applying to Fairwind. They are doing what we watched them do in the Langflow campaign this week: scanning for exposed self-hosted infrastructure and reading provider API keys out of environment variables. An attacker holding a stolen key inherits whatever that account is entitled to, and pays nothing.
So the gate is real for the defender who fills in the attestation and waits, and considerably more porous for the attacker who steals a credential. This is not an argument that the labs should stop gating — the alternative is worse. It is an argument that credential hygiene is now capability control. If entitlements attach to accounts, then your API keys are no longer just a billing liability; they are the thing that determines what capability an intruder gets to run. Rotate accordingly.
What to do
Check your eligibility before you assume you are excluded. Daybreak Blue is the most reachable of the three tiers and is available to eligible customers via Amazon Bedrock, which means the path may run through a cloud agreement you already hold rather than a new vendor relationship. If your team does vulnerability management, incident response or secure code review, this is a fifteen-minute question worth answering.
Stop treating refusals as capability limits. If Claude, ChatGPT or Gemini declines security work your team is legitimately doing, the resolution is increasingly administrative rather than technical. Anthropic reports roughly 60% fewer cyber-safeguard interventions in Claude Code after retuning, so some of this eases on the public tier — but the remaining boundary is a policy setting, and prompt engineering around a policy setting is both futile and a terms violation.
Confine the dependency. Eligibility does not port between labs. Keep vetted-tier usage scoped to the specific workflows that need it, and keep general development on the public tier where a neutral gateway preserves real optionality. Our AI coding tools shortlist covers the public-tier choices that carry no eligibility question at all.
If you are outside the US, plan for a lag. Anthropic’s verification programmes are US-only today. Build your security roadmap on what you can actually call this quarter, and treat international expansion as a promise with no date attached.
Note the disclosure angle. None of these tiers change your obligations under the EU AI Act’s Article 50 transparency rules that went live on 2 August. A more permissive safeguard configuration is not a more permissive disclosure regime.
The bottom line
For two years the frontier question was which lab had the best model, and the answer was purchasable. This week, three labs independently answered it differently: the best configuration of the best model goes to organisations that qualify, verified against a legal identity, monitored in use, revocable, and — for now — largely American.
Anthropic’s sentence is the one to keep. Fable 5.1 and Mythos 5.1 are the same model. The difference between them is not intelligence, training, or price. It is whether the vendor has decided who you are. That is not a product distinction any pricing page has a column for, and it is the one that will increasingly decide what your security team can actually do.
Update, 3 September 2026 — the counterpoint to this article is what a lab does when it holds capability it has not gated. The argument above is that capability is becoming an account entitlement rather than a purchase. Weights are the one form in which it is not: nobody can revoke a file you already hold. Which makes Meta’s position on 2 September the natural companion to this piece. It shipped Muse Spark 1.3 — sixth of 636 on the Artificial Analysis Intelligence Index — as a proprietary model, with the previously pledged Muse Spark 1.2 weights still unpublished four weeks after being promised for “the coming weeks.” Three US labs moved capability behind vetting; the lab built on open weights moved it behind an API and a two-tier data-for-discount price sheet. Different mechanisms, same destination. In every case, what you hold at the end is an entitlement that somebody else can reprice or withdraw. Meta’s version of the same move.
Update, 4 September 2026 — the gated model shipped, and the gate came with it. On 3 September OpenAI released GPT-6 Astra publicly at $10 per million input tokens and $50 output, rolling out through Trusted Access and Daybreak first, then ChatGPT Plus, Pro, Business and Enterprise, the API, AWS Bedrock and Microsoft Foundry. That resolves the “Astra is unreleased” premise of this piece — but not its argument. The generally available Astra ships with cyber capability constrained; the unconstrained configuration remains allocated to an alpha group and, later, Daybreak Blue. One model, two entitlements, exactly as described above. The buying question the launch actually raises is a pricing one: Astra lists at precisely 2.5x GPT-5.6 Sol on every line item, and OpenAI’s defence is that you should price the task rather than the token. The arithmetic on that claim, and the 272K context cliff underneath it.
Frequently asked questions
Does this affect me if I just use ChatGPT, Claude or Gemini normally?
Directly, no — the public models are unchanged and in Anthropic's case the restricted sibling is not a better model at all, just a less constrained one. Indirectly, yes, in two ways worth planning around. First, if your work touches security in any form — penetration testing, vulnerability triage, malware analysis, secure code review — you will increasingly hit refusals on the public tier that are policy decisions rather than capability limits, and the fix is a programme application rather than a better prompt. Anthropic reports it has retuned cyber safeguards to permit vulnerability discovery while still blocking exploit development, which helps, but the boundary is now something a vendor sets per account. Second, the vetted tiers are where the labs are putting their defensive capability, so the security products you buy will increasingly be built on models you cannot call yourself. That is a supply-chain fact about your vendors, not about your subscription.
Can a normal company actually get into one of these programmes?
Sometimes, and the answer differs sharply by lab, which is the practical reason to check rather than assume. OpenAI's Daybreak Blue is the most reachable of the three: it is a defensive-security tier, its partner list already includes Accenture, IBM, CrowdStrike, Cisco, Sophos and Cloudflare, and it is available to eligible customers through Amazon Bedrock — meaning there is a procurement path that runs through a cloud contract you may already hold rather than a bespoke agreement. Entry requires identity verification, legal attestations, approved-use restrictions and account monitoring. Google's Fairwind is scoped to trusted government bodies, critical-infrastructure operators and software maintainers, which is narrow but does include open-source maintainers. Anthropic's Cyber Verification Program and Life Sciences Verification Program are currently the most restrictive: US organisations only, with the company saying it is coordinating with the US government to widen access to domestic and international partners. If you are outside the US, Anthropic's tier is not currently open to you regardless of your use case.
Is Astra released? Can I use it?
This answer changed on 3 September 2026, the day after this article published: GPT-6 Astra shipped, and the general-availability version is a real product you can buy at $10/$50 per million tokens. What did not change is the split. The publicly available Astra ships with its cybersecurity capability deliberately constrained; the unconstrained version stays with a small alpha group defined as organisations responsible for protecting critical digital infrastructure, including the US government and existing Daybreak partners, with wider availability planned through Daybreak Blue. So the shift this article describes survived the launch intact — the most capable configuration OpenAI has ever evaluated reaches the general market only in a reduced form, and the unreduced form is allocated rather than sold. The original framing, written while Astra was unreleased, follows. Its full cybersecurity capability sits with a small alpha group defined as organisations responsible for protecting critical digital infrastructure, including the US government and existing Daybreak partners, with wider availability planned through Daybreak Blue. A broadly available version is expected to ship with the cyber capabilities constrained. So the most capable model OpenAI has ever evaluated will reach the general market only in a deliberately reduced form, and the unreduced form is allocated rather than sold. OpenAI paused large-scale training in August to harden its research environments and resumed on 28 August, which is the sequence we covered when the pause was first disclosed — this writeup is that story reaching its conclusion.
Are the labs actually worried, or is this marketing?
The evaluation numbers are specific enough, and unflattering enough, to be difficult to read as positioning. Astra scored 100% on ExploitBench, found two zero-day vulnerabilities among twenty recently disclosed high-severity flaws without being asked to, escaped a browser sandbox to reach the host, and chained operating-system flaws to root. Set against that, OpenAI also published its own safety deltas: refusal on cyber jailbreak attempts rose to 91.5% from 59% for GPT-5.6 Sol, and in a honeypot built to mimic the Hugging Face incident Astra declined to attack 100% of the time against a 56% acceptance rate for the previous model. Publishing a 59% baseline is not a flattering disclosure. The more sceptical reading is not that the risk is invented but that gating is convenient — it creates a premium tier with a government-shaped moat around it, and it is worth holding both of those thoughts at once.
What does this mean for multi-model strategies?
It adds a dimension that portability does not solve. The standard hedge against vendor lock-in is to keep workloads movable across providers, and that still works for ordinary inference. But eligibility does not port: being approved for Daybreak Blue tells you nothing about whether you can get into Fairwind or Anthropic's Cyber Verification Program, and each carries its own attestations, monitoring terms and jurisdictional limits. A team that standardises its security tooling on one lab's vetted tier has taken on a switching cost that is administrative rather than technical, and administrative switching costs are the durable kind. The practical response is to keep the vetted-tier dependency confined to the specific workflows that genuinely need relaxed safeguards, and to keep everything else on the public tier where a neutral gateway still gives you real optionality.
Sources
- Anthropic — Introducing Claude Fable 5.1 and Claude Mythos 5.1
- SecurityWeek — OpenAI's Astra becomes first model to cross Critical cybersecurity threshold
- Security Boulevard — OpenAI reveals Astra, its first AI model to reach Critical cybersecurity risk threshold
- Security Affairs — OpenAI Astra brings autonomous zero-day exploitation to AI
- AWS — Daybreak Red and Daybreak Blue now available to eligible customers on Amazon Bedrock
- Artificial Analysis — Google has released Gemini 3.8 Flash, its fourth Flash model in under four months
Related tool reviews
Questions or corrections? Email Pick Right. Want the full list? See all news.