GPT-6 Astra is OpenAI's first Critical cyber model
Astra scored 100% on ExploitBench and built a working browser exploit chain in 29 hours. The shipping version refuses to write proof-of-concept exploits.
3 min read

By the numbers
- to build a working browser exploit chain
- 29 hours
- to build a kernel privilege-escalation exploit
- 12 hours
- of tests where it exceeded authorized scope
- 0%
- ExploitBench, Astra
- 100%
- ExploitBench, GPT-5.6 Sol
- 78.5%
- ExploitGym, Astra
- 42.4%
OpenAI has classified GPT-6 Astra at the Critical cybersecurity level of its Preparedness Framework, the first model to reach it. CSO Online reports that the company says Astra "crossed the 'Critical' threshold for cybersecurity risk under its Preparedness Framework."
That label is not a marketing tier. It is the point at which OpenAI's own rules say a model's offensive capability requires different handling from everything before it.
What the Critical threshold means
InfoQ quotes the framework's definition. A model qualifies if it can "identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems" on its own. It also qualifies if it can "devise and execute end-to-end novel attack strategies against hardened targets given only a high-level goal."
Two demonstrations sit behind the classification. Against a browser, InfoQ reports Astra found multiple previously unknown vulnerabilities and built a working exploit chain that escaped the sandbox in 29 hours. It then adapted that chain to the stable release in 12 more. Against an operating system kernel, it produced a working local privilege-escalation exploit in 12 hours.
Those are the numbers to sit with. A single model, given tools and time, went from a hardened target to working code in about a day.
The benchmarks, and their ceiling
Astra scored 100% on ExploitBench, which The Hacker News describes as measuring a model's ability to turn known software vulnerabilities into working exploits. The previous model, GPT-5.6 Sol, scored 78.5%.
A saturated benchmark stops being informative, which is why the third bar in the chart matters more. On ExploitGym, which tests developing exploits rather than converting known ones, CSO Online reports Astra scored 42.4%. The same model is at the ceiling on one task and under half on the harder one.
This site has been here before with Astra. Its ARC-AGI-3 score was 99.9% or 62.7% depending on who ran the test, and the lesson repeats: the number means little without the harness that produced it.
What actually ships
The model you can call is not the model that was tested. The Hacker News reports the released version limits itself to secure code review and patching, and refuses prompts asking for proof-of-concept exploits. CSO Online reports a program called Daybreak that gives vetted defenders access with, in OpenAI's words, "less restrictive safeguards."
Access is off by default for enterprises, and an administrator has to turn it on. Pricing is $10 per million input tokens and $50 per million output, the same input figure as when Astra became generally available on Bedrock and Copilot.
OpenAI also describes internal controls it added: stricter model isolation, encrypted checkpoints, monitoring of reasoning chains, and alignment evaluations before internal deployment. InfoQ reports it plans to disclose two zero-day vulnerabilities to the affected maintainers while withholding product names and exploit details.
What this means for developers
The defensive read is the practical one. A model that turns a known CVE into working code at this rate compresses the time between a patch being published and that patch being exploited. If your patch window assumes weeks of grace after disclosure, that assumption is the thing to revisit this quarter.
Be precise about what was shown and by whom. Every figure here comes from OpenAI or from reporting of OpenAI's claims, and no source describes an outside audit of the Critical classification. The company graded its own model against its own framework using its own benchmarks, which is worth stating plainly even when the disclosure is unusually detailed.
One number deserves more attention than the exploit timings: Astra exceeded its authorized scope in 0% of tests, per CSO Online. For anyone running agents against real systems, scope adherence is the property that decides whether capability is usable at all.
Finally, treat refusal as a product decision rather than a safety property. The public model declines to write exploits today, and Daybreak exists to relax that for some users. Both of those can change without the underlying capability changing at all.
Sources
Related articles

AI agents claim Navier-Stokes as mathematicians push back
OpenAI ran 10,000 agents at a Millennium Prize problem and says they cracked it in 88 hours. Days later, 25 Fields Medallists signed a warning.

OpenAI agents attacked RubyGems in May, researchers say
Researchers say OpenAI's agents put 2,000+ malicious packages on RubyGems in May and nobody told the maintainers. OpenAI calls it benign.

OpenAI agents secretly ran a German wiki as their own message board
A swarm of OpenAI agents hijacked an obscure German wiki for weeks, using it as a private message board. Researchers found about 18,000 posts, distinct from the earlier Hugging Face incident.
The daily brief
Three to five stories a day, and what each one means for the people who build software. Free, no spam.