Skip to content
Tech AI Wire

GPT-6 Astra is OpenAI's first Critical cyber model

Astra scored 100% on ExploitBench and built a working browser exploit chain in 29 hours. The shipping version refuses to write proof-of-concept exploits.

By Tech AI Wire Team

3 min read

XLinkedIn
An empty office reception area at night, with the OpenAI knot mark in brushed steel on the wall behind the desk.

By the numbers

to build a working browser exploit chain
29 hours
to build a kernel privilege-escalation exploit
12 hours
of tests where it exceeded authorized scope
0%
Exploit benchmark scores
ExploitBench, Astra
100%
ExploitBench, GPT-5.6 Sol
78.5%
ExploitGym, Astra
42.4%

OpenAI has classified GPT-6 Astra at the Critical cybersecurity level of its Preparedness Framework, the first model to reach it. CSO Online reports that the company says Astra "crossed the 'Critical' threshold for cybersecurity risk under its Preparedness Framework."

That label is not a marketing tier. It is the point at which OpenAI's own rules say a model's offensive capability requires different handling from everything before it.

What the Critical threshold means

InfoQ quotes the framework's definition. A model qualifies if it can "identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems" on its own. It also qualifies if it can "devise and execute end-to-end novel attack strategies against hardened targets given only a high-level goal."

Two demonstrations sit behind the classification. Against a browser, InfoQ reports Astra found multiple previously unknown vulnerabilities and built a working exploit chain that escaped the sandbox in 29 hours. It then adapted that chain to the stable release in 12 more. Against an operating system kernel, it produced a working local privilege-escalation exploit in 12 hours.

Those are the numbers to sit with. A single model, given tools and time, went from a hardened target to working code in about a day.

The benchmarks, and their ceiling

Astra scored 100% on ExploitBench, which The Hacker News describes as measuring a model's ability to turn known software vulnerabilities into working exploits. The previous model, GPT-5.6 Sol, scored 78.5%.

A saturated benchmark stops being informative, which is why the third bar in the chart matters more. On ExploitGym, which tests developing exploits rather than converting known ones, CSO Online reports Astra scored 42.4%. The same model is at the ceiling on one task and under half on the harder one.

This site has been here before with Astra. Its ARC-AGI-3 score was 99.9% or 62.7% depending on who ran the test, and the lesson repeats: the number means little without the harness that produced it.

What actually ships

The model you can call is not the model that was tested. The Hacker News reports the released version limits itself to secure code review and patching, and refuses prompts asking for proof-of-concept exploits. CSO Online reports a program called Daybreak that gives vetted defenders access with, in OpenAI's words, "less restrictive safeguards."

Access is off by default for enterprises, and an administrator has to turn it on. Pricing is $10 per million input tokens and $50 per million output, the same input figure as when Astra became generally available on Bedrock and Copilot.

OpenAI also describes internal controls it added: stricter model isolation, encrypted checkpoints, monitoring of reasoning chains, and alignment evaluations before internal deployment. InfoQ reports it plans to disclose two zero-day vulnerabilities to the affected maintainers while withholding product names and exploit details.

What this means for developers

The defensive read is the practical one. A model that turns a known CVE into working code at this rate compresses the time between a patch being published and that patch being exploited. If your patch window assumes weeks of grace after disclosure, that assumption is the thing to revisit this quarter.

Be precise about what was shown and by whom. Every figure here comes from OpenAI or from reporting of OpenAI's claims, and no source describes an outside audit of the Critical classification. The company graded its own model against its own framework using its own benchmarks, which is worth stating plainly even when the disclosure is unusually detailed.

One number deserves more attention than the exploit timings: Astra exceeded its authorized scope in 0% of tests, per CSO Online. For anyone running agents against real systems, scope adherence is the property that decides whether capability is usable at all.

Finally, treat refusal as a product decision rather than a safety property. The public model declines to write exploits today, and Daybreak exists to relax that for some users. Both of those can change without the underlying capability changing at all.

Sources

  1. GPT-6 Astra Is the First Model OpenAI Classifies as Critical for Cybersecurity - InfoQ
  2. OpenAI launches GPT-6 Astra, its first model to cross a critical cybersecurity threshold - CSO Online
  3. GPT-6 Astra Scores 100% on ExploitBench as OpenAI Blocks PoC Exploit Requests - The Hacker News

Related articles

The daily brief

Three to five stories a day, and what each one means for the people who build software. Free, no spam.