OpenAI put a model on the market this week whose selling point is that it says yes. GPT-5.6-Cyber, announced on 10 August 2026, is trained on zero-day discovery and exploit-chain construction, and OpenAI's own internal evaluation puts its completion rate on advanced offensive-security prompts at 95.0%. The same evaluation scores stock GPT-5.6 Sol at 1.5%. Nothing about the underlying reasoning changed that much in a week. What changed is who is allowed to ask.

The timing is the part worth sitting with. Four days earlier OpenAI told the world it had slowed development of Astra, an unreleased frontier model, after internal evaluations put it at the Critical cybersecurity threshold of its Preparedness Framework. That was the first time a major lab publicly throttled one of its own models over offensive cyber capability. Now the same company has shipped a model built specifically for that work, with the refusal layer turned down, behind a partner allowlist. Both decisions are defensible. Read together, they describe where the industry has actually landed: the safety control is no longer the model's behaviour. It is the gate in front of it.

RelatedOpenAI Mandates Hardware Passkeys for Cyber Access

  • Two tiers, not one. Daybreak Blue gives vetted defenders GPT-5.6 Sol with system-level cyber safeguards lifted. Daybreak Red gives them purpose-trained models, GPT-5.5-Cyber and now GPT-5.6-Cyber.
  • The safeguard gap is enormous. 1.5% completion on the public model, 2.0% under Daybreak Blue, 57.3% for GPT-5.5-Cyber, 95.0% for GPT-5.6-Cyber on OpenAI's internal benchmark.
  • It has already produced findings. OpenAI says the model surfaced previously unknown bugs in V8, Chrome's JavaScript engine, one of which is now tracked as CVE-2026-15903 and rated high severity at CVSS 8.8.
  • Every major lab now has one. Google locked up Gemini 3.5 Flash Cyber in July. Microsoft shipped MAI-Cyber-1-Flash a week later. OpenAI is the third, and the most openly commercial about it.
Advanced cybersecurity completion rate by model and access tier Bar chart. Public GPT-5.6 Sol completes 1.5 percent of advanced offensive security requests, Sol under Daybreak Blue 2.0 percent, GPT-5.5-Cyber 57.3 percent, GPT-5.6-Cyber 95.0 percent. OPENAI INTERNAL EVAL Advanced cybersecurity completion rate GPT-5.6 Sol (public) 1.5% Sol + Daybreak Blue 2.0% GPT-5.5-Cyber 57.3% GPT-5.6-Cyber 95% Share of advanced offensive-security requests the model completes. OpenAI internal benchmark, Aug 2026. genztech.blog
Fig 1 · benchmark The jump from 57.3% to 95% is a model change. The jump from 1.5% to 57.3% is a policy change.

What is Daybreak Red, and who gets in?

Daybreak was already OpenAI's programme for defenders. This week it split in two. Daybreak Blue approves an organisation to use GPT-5.6 Sol with the system-level cybersecurity safeguards lifted, aimed at vulnerability discovery, secure code review, malware analysis, incident response and patch validation. Daybreak Red is the deeper tier: access to models trained specifically for the work, covering vulnerability research, exploit validation and security testing. GPT-5.6-Cyber lives only there.

Entry is by application and review, not by credit card. The Hacker News reports early partners including Cisco, Cloudflare, CrowdStrike, Fortinet, IBM, Palo Alto Networks, Accenture, Akamai, PwC and Sophos, which reads less like a research programme and more like the incumbent security industry getting first refusal. OpenAI says system-level screening still sits in front of requests, and that the model continues to refuse prompts it judges highly dual-use even with the standard guardrails removed. That last claim is the one nobody outside the allowlist can check.

Why does a 1.5% baseline matter so much?

Because it quantifies something the industry usually argues about in the abstract. The gap between 1.5% and 95% is not a measure of what a frontier model can do. It is a measure of how much capability sits behind the refusal layer of a model that millions of people already use. Same lineage, same weights family, different permissions. Daybreak Blue moving the public model from 1.5% to only 2.0% is the quiet detail here: lifting the system-level safeguards on Sol barely moves the number, which tells you the refusals were never the main brake. The capability difference is in the training, and the training is what Daybreak Red gates.

The counterargument OpenAI is making, and it is not a weak one, is that defenders are already outgunned by attackers who feel no obligation to respect a refusal. If offensive capability is going to exist regardless, better it exist inside a vetted programme with disclosure obligations than only outside one. CVE-2026-15903 is the exhibit: OpenAI ran the model at V8, found bugs in the engine that renders most of the web, and sent them to Google through coordinated disclosure rather than publishing them.

How OpenAI's three access tiers differ Diagram showing the public API, Daybreak Blue and Daybreak Red tiers, with an application and vetting gate sitting in front of the two Daybreak tiers. ACCESS TIERS Application, review and approval Public API 1.5% GPT-5.6 Sol, stock Full safeguards on Anyone can sign up Daybreak Blue 2.0% GPT-5.6 Sol, safeguards lifted for defence work Vetted organisations Daybreak Red 95% GPT-5.6-Cyber and GPT-5.5-Cyber Purpose-trained Percentages are OpenAI's internal advanced cybersecurity completion rate, not a public benchmark. genztech.blog
Fig 2 The capability is in the training, not the toggle. Lifting safeguards on the public model moves it from 1.5% to 2.0%.

How does this compare with Google and Microsoft?

All three big labs built an offensive-capable security model inside four weeks of each other, and all three chose a different door.

 OpenAI GPT-5.6-CyberGoogle Gemini 3.5 Flash CyberMicrosoft MAI-Cyber-1-Flash
Announced10 Aug 202621 Jul 202627 Jul 2026
Who can use itVetted commercial partnersGovernments and select partnersMicrosoft security stack customers
PositioningOffensive research, refusals reducedFind, prove and patchDefensive, cost-optimised
Headline number95% completion rate55 confirmed V8 bugs96% on CyberGym
Commercially open?Yes, by applicationNoYes, bundled

Google built the most capable of the three by its own account and then declined to sell it. Microsoft framed its model around defence and cost. OpenAI is the first to put an explicitly offensive-leaning model on a commercial application form. That is the genuine break in the pattern, and it is why this launch matters more than its benchmark table suggests.

  1. 21 Jul 2026Google DeepMind unveils Gemini 3.5 Flash Cyber Locked to governments and selected partners
  2. 27 Jul 2026Microsoft ships MAI-Cyber-1-Flash Defensive framing, 96% on CyberGym
  3. 7 Aug 2026OpenAI slows Astra First public Critical cyber rating under its Preparedness Framework
  4. 10 Aug 2026Daybreak splits into Blue and Red, GPT-5.6-Cyber ships 95% completion rate, application-gated
  5. NextFirst independent evaluation of a Daybreak-tier model No neutral harness has scored one yet

What should defenders actually do with this?

If your organisation is on the allowlist, the practical value is triage speed on your own code, not offence. If it is not, the change that reaches you is indirect but real: the pool of people who can generate a working exploit chain from a patch diff just got wider by however many companies OpenAI approves, and every one of those companies now has staff, contractors and turnover. Assume patch-to-exploit windows keep compressing and plan patch cycles against days, not weeks.

RelatedGoogle Built an Exploit-Writing Gemini, Then Locked It Up

Alex Goller of Illumio, quoted by Infosecurity Magazine, called the tiering a reasonable first step but pushed back on where the control belongs, arguing that the controls that matter follow zero trust principles and are enforced in your infrastructure, not in the model. That is the right instinct. A vendor allowlist is an access control on one supplier. Segmentation, least privilege and fast patching are controls you own.

What it means for the market

The security vendors named as partners are the obvious read. Palo Alto Networks (PANW), CrowdStrike (CRWD), Cloudflare (NET), Fortinet (FTNT), Akamai (AKAM) and Cisco (CSCO) all now have privileged access to capability their smaller competitors cannot buy at any price, which is a moat that did not exist a week ago. The signal for investors is consolidation pressure on mid-market security tooling: if vulnerability research quality becomes a function of which model tier you can get approved for, scale starts to compound in a way it did not when everyone bought the same scanners. Watch whether OpenAI prices Daybreak Red as a premium tier or as a partnership, because that choice decides whether this is a product line or a policy programme. This is analysis, not investment advice.

What to watch · 2026
  • Who gets approved next. The first independent researcher or university on the Daybreak Red list would tell us the gate is about competence. A list that stays large-vendor only tells us it is about liability.
  • An independent score. Every number here is OpenAI's own. No neutral harness has evaluated a Daybreak-tier model, and 95% on an internal benchmark is a vendor claim until someone else runs it.
  • Astra's status. If the paused model eventually ships through Daybreak Red, the Preparedness Framework's Critical threshold effectively becomes a routing rule rather than a stop.
  • Leak surface. The security question is no longer whether the model refuses. It is whether an approved partner's credentials stay approved.

Our take

The honest summary of this week is that a refusal rate turned out to be a business decision. OpenAI can hold both positions at once, pausing Astra and shipping GPT-5.6-Cyber, because they were never really about the same thing: Astra was about a capability ceiling nobody has agreed how to handle, and Daybreak is about who is on the list. That is a coherent policy. It is also a much smaller safety guarantee than the language around it implies, and it puts a lot of weight on a vetting process whose criteria have not been published.

The number to keep is 1.5% versus 2.0%. Lifting the guardrails on the public model changed almost nothing, which means the guardrails were never what stood between an ordinary user and this capability. The training did. Three labs have now done that training. Only one of them is selling it.

Primary sources

Original analysis by GenZTech. Primary announcement: OpenAI, 10 August 2026.