OpenAI's New Cyber Model Barely Refuses Anything. Paperwork Decides Who Gets It.
GPT-5.6-Cyber answers 95% of exploit-development requests. The standard model answers 1.5%. Same family, different account tier. If capability is gated by entitlement rather than by weights, your model risk assessment is aimed at the wrong object.

Sometime before 10 August, a small number of very large security and consulting firms signed a set of legal attestations with OpenAI and received something their competitors cannot buy at any published price: a model that will answer questions about exploit-chain development. On that date OpenAI published a post titled “Expanding Daybreak as the Cyber Defense Window Narrows”, unsigned and quoting no named OpenAI employee, describing a two-tier programme in which the same underlying model family behaves in fundamentally different ways depending on who holds the key.
The numbers in that post are the most consequential thing OpenAI has published about AI safety this year, and not for the reason the headlines suggested. “GPT-5.6-Cyber completes 95.0% of these requests,” the company writes, “compared with just 1.5% for GPT-5.6 Sol, and 2.0% when used with Daybreak Blue access.” One family, three behaviours, a spread of more than sixty-fold. The variable is not the weights. It is the entitlement attached to the account making the call.
What the 95% Actually Measures
Read OpenAI’s own definition, because it does the work: the Advanced Cybersecurity Completion Rate “measures how often models will respond to requests involving exploit-chain development, authentication bypass, privilege escalation, and other advanced cybersecurity scenarios.” Respond. Not succeed. This is a refusal-reduction score wearing the clothing of a capability benchmark, and OpenAI says so plainly. Much of the coverage did not; VentureBeat ran it as “95% completion on advanced cybersecurity tasks,” which a busy reader will hear as competence rather than compliance.
OpenAI publishes no prompt count for this evaluation, no grading methodology, and no external validation. The predecessor comparison sits in the same register. GPT-5.5-Cyber “completes only 57.3% of requests, addressing feedback from security researchers who encountered persistent refusals with the earlier model.” The product decision described there is a decision about friction, not about intelligence. Customers complained that the model said no too often. The company built a tier where it says yes.
The defensive case is real and, unusually, checkable. Per Infosecurity Magazine, the model “finds CVEs, helping identify vulnerabilities like CVE-2026-15903 in Chrome’s V8 engine, which Google subsequently patched.” Eduard Kovacs reported in SecurityWeek on 11 August that it also surfaced flaws in mobile operating system, database and kernel code. OpenAI’s stated rationale is a defensible reading of an arms race in which, by its own account, “threat actors will increasingly use AI to conduct cyberattacks at unprecedented speed and scale, including in fully autonomous ways.”
Two Levers, and Only One of Them Is the Model
OpenAI is operating two distinct controls, and the difference between them is the whole argument. Sam Sabin reported in Axios on 10 August that OpenAI delayed a separate model, Astra, after it demonstrated critical hacking capability in safety testing. That is capability control: the weights never ship. GPT-5.6-Cyber reached only the “High” threshold under the Preparedness Framework, a lower rung. It ships, with refusals relaxed and access gated by identity verification, monitoring, approved-use restrictions and legal attestations. Individual applicants must adopt hardware security keys from 1 September 2026.
Below the capability line, then, the account is the control. OpenAI concedes the residual: “Models running with reduced safeguards carry risks beyond standard model usage, whether from misuse or misalignment.” Notice what the hardware-key mandate tells you about the threat model OpenAI is defending against. It is not jailbreaking. It is credential theft. The company has concluded that the way someone gets an unrestricted cyber model is by stealing a session from someone who is entitled to one.
Alex Goller, principal solution architect for EMEA at Illumio, put the structural point to Kevin Poireault of Infosecurity Magazine on 11 August. He called the approach “a good first step,” then drew the line: “AI model guardrails were never the control plane for defense... controls that matter follow zero trust principles.” And: “enforcement lives in your infrastructure, not in the model.” He recommended sandboxing and scoped authorisation. He is describing OpenAI’s own design back to it.
The Question to Put to Your Suppliers
Which brings this to your desk. If capability varies by entitlement rather than by weights, a model risk assessment that assesses the model is assessing the wrong object. The live question is which of your third parties holds Daybreak Red, what they attested to in order to get it, and whether your contracts say anything at all about a supplier operating a reduced-safeguard model in an engagement scoped against your estate.
The rosters differ by outlet, so treat them as overlapping reports rather than one register. Axios named Accenture, IBM, CrowdStrike, Cisco and Palo Alto Networks. SecurityWeek added Capgemini, EY and Cloudflare. Ravie Lakshmanan’s list in The Hacker News on 11 August extends to Akamai, Fortinet, PwC and Sophos. The union takes in two of the Big Four and the largest global consultancies. If you buy audit, red teaming or MDR, you are almost certainly already downstream of an entitlement you did not procure and cannot see. eesel AI reported on 11 August that under the partner programme model access stays with the approved partner and is not transferred to the customer; OpenAI’s page does not confirm this, and it should be treated as unconfirmed until it is contractual.
Pricing is undisclosed. The public card shows a Cyber row with blank rate columns, against $5 in and $30 out per million tokens for GPT-5.6 Sol. Forkast called it a premium product category, which understates it. A price you cannot see is not a price; it is an admission ticket. And admission tickets, unlike model weights, can be revoked, transferred, subpoenaed and stolen. Ask your vendors which of those four they have planned for.
This analysis rests on OpenAI’s published post of 10 August 2026, which is unsigned and quotes no named employee. The 95.0%, 1.5%, 2.0% and 57.3% figures are OpenAI’s own internal evaluation, self-reported with no prompt count, grading methodology or external validation. Partner rosters differ across Axios, SecurityWeek, The Hacker News and TechCrunch and are attributed accordingly; eesel AI’s statement that model access is not transferred to end customers is unconfirmed by OpenAI, as is the price of GPT-5.6-Cyber.
Sources
OpenAI, Expanding Daybreak as the Cyber Defense Window Narrows, 10 August 2026 · Kevin Poireault, Infosecurity Magazine · Sam Sabin, Axios · Eduard Kovacs, SecurityWeek · Ravie Lakshmanan, The Hacker News · Lucas Ropek, TechCrunch · VentureBeat · eesel AI
Source: OpenAI