AI-generated content. This article was researched and written by an automated AI editorial system and published without prior human review. Every factual claim is checked against cited primary sources before publication, but no journalist read this page before you did — treat it accordingly, and report anything that looks wrong. How this works ›

Some links on this page are affiliate links. We may earn a commission at no extra cost to you.
Updated: Aug 23, 2026
·
openaisecuritymodel-launchgovernancedual-use

OpenAI built a model that writes exploits — and the interesting part is who is allowed to use it

TL;DR: On 10 August 2026, OpenAI announced GPT-5.6-Cyber, a model built on GPT-5.6 Sol and deliberately configured with lower refusal rates for offensive-security work. Vendor testing reports 95% completion on exploit-chain development, privilege escalation and authentication bypass — against 1.5% for Sol itself, and 57.3% for last generation’s GPT-5.5-Cyber. It is not generally available: access runs through Daybreak Red, a vetted tier of named consultancies, security vendors and red teams. The model has already produced CVE-2026-15903 (CVSS 8.8), an out-of-bounds flaw in Chrome’s V8 engine that chains with a second zero-day to escape the heap sandbox; Google has patched it. Reported pricing is $12.50 per million input tokens and $75 per million output. The structural news is not the capability — labs have had it for a while, as Anthropic’s models breaching three real companies showed in July. It is that a frontier lab has stopped treating offensive capability as something to suppress and started treating it as something to distribute selectively. Refusal is now a distribution decision, not a training outcome.

Update (11 August 2026): OpenAI and AWS have made both Daybreak tiers available to eligible customers on Amazon Bedrock — Daybreak Red (which fronts GPT-5.6-Cyber) and Daybreak Blue (GPT-5.6 Sol with defensive guardrails) — running in AWS US East (N. Virginia). Access still requires enrolment in OpenAI’s Trusted Access for Cyber vetting programme, so this widens the distribution channel without widening the gate: the two-tier governance model described below now extends onto a hyperscaler’s marketplace with the vetting intact. AWS states that prompts and outputs are not accessible to others and are not used for training. It reinforces the core reading here — offensive capability is being managed as a distribution problem, and the distribution surface is growing.

What was announced

OpenAI’s 10 August announcement had two parts, and most coverage led with the wrong one.

The first part is a model. GPT-5.6-Cyber is built on GPT-5.6 Sol, the flagship of the GPT-5.6 family that reached general release in July, and it is trained and configured for what OpenAI calls “advanced, authorized cybersecurity work”: finding zero-days, developing exploit chains, and the adjacent tasks — privilege escalation, authentication bypass — that turn a discovered flaw into a working intrusion.

The second part is a distribution system, and that is the actual story. OpenAI expanded its Daybreak Cyber Partner programme into two tiers. Daybreak Blue provides access to Sol and other general-purpose frontier models with guardrails tuned for defensive work. Daybreak Red provides access to the dedicated cyber models, including GPT-5.6-Cyber. There is no self-serve route into Red. Publicly named partners span the large consultancies — Accenture, Capgemini, Cognizant, EY, IBM, KPMG, PwC — the specialist offensive-security firms NCC Group and SpecterOps, and the major security vendors: Palo Alto Networks, CrowdStrike, Cisco, Sophos, Akamai, Fortinet and Cloudflare.

The capability numbers are what make the tiering necessary. On prompts covering exploit-chain development, privilege escalation and authentication bypass, GPT-5.6-Cyber completes 95%. The same prompts against GPT-5.6 Sol complete at 1.5%. Last generation’s GPT-5.5-Cyber managed 57.3%.

Two things follow from that trio of numbers, and both matter more than the headline figure.

The 1.5% is the interesting number

A 63-fold gap between Cyber and Sol on identical prompts tells you the refusal behaviour in the mass-market model is holding. Whatever else is true here, OpenAI did not loosen the product that hundreds of millions of people use. If you use ChatGPT for ordinary work, nothing about its safety behaviour changed on 10 August.

That is worth stating plainly, because the natural reading of “OpenAI reduced safeguards” is that the thing in your browser got more dangerous. It did not. A separate model, behind a separate contractual gate, got more permissive.

The jump from 57.3% to 95% in a single generation is the number that should hold a planner’s attention. That is not a safety-tuning delta — GPT-5.5-Cyber was already a permissive offensive model, and it still failed roughly two of every five tasks. The improvement is capability. The gap between “an AI that helps a skilled operator” and “an AI that completes the task” closed substantially in about one product cycle.

Refusal has become a distribution decision

For roughly three years, the industry’s answer to dangerous capability was to train it out, or at least train refusal in. The safety artefact lived inside the model.

GPT-5.6-Cyber inverts that. The capability is deliberately present. The safety artefact lives in the contract — who is vetted, what tier they hold, what they agree to. OpenAI’s stated rationale, that “democratizing access to frontier intelligence for defenders is crucial to accelerating and automating cyber defense,” is an argument about the defense gap: attackers are already using AI at scale and without permission, so restricting defenders’ access does not remove the capability from the world, it removes it from one side.

The argument is coherent. It is also, notably, the same argument the offensive-security industry has made about Metasploit, Cobalt Strike and Burp Suite for two decades — tools that are indispensable to defenders and routinely abused by attackers. The industry’s settled answer there was not to stop building them; it was licensing, vetting and monitoring. OpenAI has arrived at the same place, faster, with a more capable artefact.

What makes this different from Cobalt Strike is that the artefact is a model, and models leak in ways binaries do not. The pattern this site has tracked all year is that frontier capability arrives at the open-weights tier on a lag — a lag measured in months, not years. Nothing in Daybreak Red’s structure addresses what happens when a comparably capable model ships with downloadable weights and no partner agreement.

Why this matters

It confirms the trajectory, rather than starting it. In July, Anthropic disclosed that three of its models compromised real companies during adversarial evaluations after a partner misconfiguration granted unintended internet access. The same month, the Five Eyes agencies issued a joint warning about frontier-model cyber capability, and Anthropic published a severity framework for cyber jailbreaks. Google shipped cyber-tuned Gemini variants in the same window. Every major lab is now building in this direction; OpenAI is the first to formalise a two-tier access regime around it.

The vendor stack absorbed it immediately. Palo Alto Networks, CrowdStrike, Cisco, Fortinet, Akamai, Sophos and Cloudflare are all in the partner list. If you buy security products from any of them, this capability is arriving in your supply chain whether or not you ever see the model. That is a procurement question — what your vendors do with it, how they log it, what happens to findings — and it is worth asking before the renewal conversation.

The next threshold is already named. SecurityWeek’s reporting notes OpenAI’s own assessment that its forthcoming Astra model — the system that recently produced ten machine-verified mathematical proofs — could reach a “critical” cybersecurity risk threshold, potentially enabling autonomous zero-day discovery and exploitation. GPT-5.6-Cyber requires an operator. The stated concern about Astra is that the next one may not.

Pricing signals the intended user. At a reported $12.50 per million input tokens and $75 per million output — with cached input at $1.25 — this is priced far above Sol or its competitors at Anthropic and Google. That is consultancy-and-enterprise pricing, not researcher pricing, and it reinforces that the intended buyer is an organisation running paid engagements.

Honest caveats

The 95% figure is OpenAI’s own. It comes from vendor testing, on a vendor-selected prompt set, with no published methodology and no independent replication. It should be read as a directional claim about capability, not a measured fact. This site’s standing position on vendor benchmarks applies with extra force when the vendor is describing its own product’s willingness to do something dangerous.

“Completion” is not “success.” A model completing an exploit-chain prompt is not the same as producing a working exploit against a hardened target. The V8 finding is real and CVE-numbered, but one strong public result does not establish a base rate.

The partner list is not a security control on its own. Being a named Fortune 500 consultancy says nothing about the internal controls applied to the model once inside. OpenAI has not published its vetting criteria, monitoring approach, or what revocation looks like — and the honest reporting acknowledges that reduced-safeguard models “carry risks beyond standard model usage, whether from misuse or misalignment.”

Timing is unclear on the V8 chain. The CVE was patched by Google in July, before the August announcement, which suggests the discovery came from earlier access or pre-release testing rather than from the model as launched.

The verdict

The capability itself is not the surprise. Anyone tracking this beat expected frontier models to reach competent exploit development, and the July disclosures established they were already there.

The shift worth registering is governance. OpenAI has stopped pretending offensive capability can be trained away and started managing it as a distribution problem — vetted tiers, named partners, contractual control. That is a more honest posture than the alternative, and it is probably the right one for a capability that exists whether or not it is sold.

It is also a posture with an obvious failure mode. Contractual control works exactly as long as the capability stays scarce, and nothing in the last eighteen months suggests frontier capability stays scarce for long. The open-weights tier is roughly six to twelve months behind the frontier and closing. When an equivalently capable model ships with downloadable weights, Daybreak Red’s partner list will govern precisely nothing.

For buyers, the practical takeaway is narrow and unglamorous. You cannot get this model, your adversaries cannot get this specific model either, and both of those facts are temporary. What is not temporary is the direction: the cost of turning a known flaw into a working intrusion is falling fast. Patch latency and detection coverage were always the things that mattered. They now matter on a shorter clock.

Update (August 20, 2026). The governance-as-distribution posture described above has since acquired a second half: gating development, not just access. On 18 August 2026 OpenAI disclosed that it had paused frontier RL training for two weeks and left its largest planned run on hold, after determining on 7 August that its unreleased Astra model could not be ruled out as reaching Critical cyber capability under the Preparedness Framework — the first such designation. All Astra inference with tools now requires monitoring, and OpenAI put that monitoring’s cost at roughly 20% of the inference compute being watched. The scarcity argument below still holds: GLM-5.3 shipped open weights aimed at coding and defensive cyber work on 14 August, and open models do not pause.

Update (23 August 2026). The defensive mirror of this posture has now shipped. On 21 August Anthropic made Mythos 5 available inside Claude Security, where customers get vulnerability findings and patches but no direct model access, and said it is extending its Cyber Verification Program — reduced safeguards for vetted defenders — from Opus and Sonnet to Mythos-class models in the coming weeks. Both labs have now concluded that frontier cyber capability cannot simply be listed on a pricing page. They picked different gates: OpenAI vets the customer for an offensive product, Anthropic restricts the interface for a defensive one. The common finding is that “sign up and get the best model” is over for this capability class.

Frequently asked questions

Can I get access to GPT-5.6-Cyber?

Almost certainly not, unless you work at a named partner. Access runs through Daybreak Red, a vetted tier of OpenAI's Daybreak Cyber Partner programme. The publicly named participants are large consultancies and security vendors — Accenture, Capgemini, Cognizant, EY, IBM, KPMG, PwC, NCC Group, SpecterOps, Palo Alto Networks, CrowdStrike, Cisco, Sophos, Akamai, Fortinet and Cloudflare. There is no self-serve signup, and OpenAI has not published its vetting criteria.

What is the difference between Daybreak Blue and Daybreak Red?

Blue is the defensive tier: access to GPT-5.6 Sol and other general-purpose frontier models with guardrails tuned for defensive security work. Red is the offensive tier: access to dedicated cyber models including GPT-5.6-Cyber, which is configured to refuse far less on dual-use tasks. The split is the governance mechanism — the capability is separated from the general product surface rather than shipped inside it.

Does this mean ChatGPT will now help with hacking?

No. GPT-5.6-Cyber is a separate model behind a separate access tier. The consumer ChatGPT product and the standard GPT-5.6 Sol model retain their refusal behaviour — Sol completes 1.5% of the same offensive prompts, against 95% for Cyber. That 63-fold gap is the point: the safety behaviour of the mass-market product was not changed.

Has the model actually found real vulnerabilities?

Yes. It identified CVE-2026-15903, an out-of-bounds read/write flaw in Chrome's V8 JavaScript engine rated CVSS 8.8, allowing remote code execution via crafted HTML — and chainable with a second flaw to escape the V8 heap sandbox. Google patched it. Reporting also references findings in an unnamed mobile operating system, a database and an OS kernel.

What should a security team actually do about this?

For most teams the immediate answer is nothing, because access is closed. The planning question is different: if exploit development is becoming a capability your vendors can rent, assume your adversaries will reach equivalent capability on a lag, and prioritise patch latency and detection over hoping the capability stays scarce.

Sources

Related tool reviews

Questions or corrections? Email Pick Right. Want the full list? See all news.