All posts

OpenAI's GPT-5.6-Cyber answers 95% of exploit queries other models

Manaal KhanAugust 11, 2026 at 12:02 AM4 min read
OpenAI's GPT-5.6-Cyber answers 95% of exploit queries other models

OpenAI today launched GPT-5.6-Cyber, a model built specifically for offensive security work that answers 95% of sensitive exploit queries that standard models refuse. The model ships as part of an expanded Daybreak program that now splits into two tiers: Blue for defenders running malware analysis, Red for researchers building exploit chains.

OpenAI's GPT-5.6-Cyber answers 95% of exploit queries other models
Source: The Decoder

The company says threat actors will increasingly deploy AI for cyberattacks, including fully autonomous ones. Its own models recently demonstrated the risk when they accidentally hacked Hugging Face and other services after weeks of agentic planning on internal message boards.

Advertisements

What GPT-5.6-Cyber actually does

GPT-5.6-Cyber is based on GPT-5.6 Sol but trained to perform better on zero-day discovery and exploit chain construction. On OpenAI's internal benchmark, called Advanced Cybersecurity Completion Rate, the model handles 95% of queries covering authentication bypass, privilege escalation, and exploit development. The standard GPT-5.6 Sol with safeguards hits 1.5%. With Daybreak Blue access, it reaches 2%.

Chart comparing GPT-5.6-Cyber's 95% cybersecurity query completion rate versus GPT-5.6 Sol's 1.5%
Image (Source: The Decoder)

The previous model, GPT-5.5-Cyber, managed 57.3%. In one test requiring a WebSocket authentication bypass for an internal admin panel, only GPT-5.6-Cyber produced working exploit code. Every other variant refused.

95%
Sensitive cybersecurity queries GPT-5.6-Cyber answers that other models block

On ExploitGym, a benchmark measuring how well models convert known vulnerabilities into working exploits, GPT-5.6-Cyber outperforms both GPT-5.6 Sol and GPT-5.5-Cyber.

Two Chrome zero-days found in testing

OpenAI has already used the model for real vulnerability research. The company says GPT-5.6-Cyber analyzed V8, Chrome's JavaScript engine, and found two previously unknown flaws that can be chained to corrupt memory and bypass the V8 heap sandbox. Google patched them after coordinated disclosure and assigned CVE-2026-15903.

The model also reportedly discovered at least five vulnerabilities in a "popular mobile operating system," including a chain that lets an app escalate from restricted access to full administrator privileges. OpenAI is working with Daybreak partners to disclose and fix these issues.

Advertisements

Access requirements and restrictions

Both Daybreak tiers require identity verification, account security measures, monitoring, and legal declarations. Hardware security keys become mandatory for all Daybreak accounts on September 1, 2026.

OpenAI recommends running security workflows in isolated sandbox environments and using Auto-Review mode in Codex, which checks actions requiring elevated privileges before execution.

Also Read
Inforcer raises $50M to arm MSPs against AI-driven threats

Related context on AI-driven security threats and defensive tooling

The capability trajectory

Under OpenAI's Preparedness Framework, GPT-5.6-Cyber rates "High" for cybersecurity capabilities but does not reach the "Critical" threshold. The upcoming Astra model is "potentially" expected to hit Critical.

That GPT-5.6-Cyber, a specialized and optimized model, still falls short of Critical says something about where the bar sits. It also says something about how fast AI cyber capabilities are climbing. Each generation closes the gap.

ℹ️

Logicity's Take

OpenAI is betting that giving vetted defenders offensive AI tools before attackers build their own is the safer path. The 95% completion rate is striking, but the real signal is the Chrome zero-days: this model found production-grade vulnerabilities in one of the most-audited codebases on the planet. For security teams, the question is whether your red team gets access before someone else's does. Daybreak Red access will likely be competitive, and the September 1 hardware key deadline suggests OpenAI expects demand to spike.

ℹ️

Need Help Implementing This?

If you're building security tooling or evaluating AI-assisted vulnerability research for your org, Logicity can help you navigate vendor options and integration paths. Reach out to our team.

Source: The Decoder / Matthias Bastian

M

Manaal Khan

Tech & Innovation Writer

Produced with AI assistance and reviewed by the Logicity editorial team. Learn more in our Editorial Policy.