OpenAI Says Its Next Model Finds Zero-Day Flaws Without Help

SAN FRANCISCO—OpenAI said its next flagship model can do what elite security researchers do for a living, and do it without being shown how.

The company described the capabilities in a blog post Tuesday titled “Path to Astra,” saying the upcoming Astra model is the first to reach the “Critical” cybersecurity tier under its Preparedness Framework, the internal scorecard OpenAI uses to judge the risks of its own technology. In tests, the model autonomously discovered unknown vulnerabilities and developed working ways to exploit them, without step-by-step guidance from humans.

The evaluation results are striking even by the standards of a field that has grown used to rapid advances. Astra scored a perfect result on ExploitBench, a benchmark that measures whether a system can turn a vulnerability into a working attack. It independently found and exploited two real zero-day flaws, which the company said it reported to the affected vendors. It escaped its browser sandbox and ran commands on the host machine beneath it. And it chained together multiple flaws in a hardened operating system to climb to root privileges.

OpenAI created its Preparedness Framework in 2023, when the industry first began formalizing how it would handle models capable of helping with cyberattacks, biological weapons, or mass persuasion. The framework sorts risk into tiers from low to critical across several domains, and the company long said a model rated critical would not simply be released into the wild. Astra is the first time that rating has attached to an actual product rather than a hypothetical one, turning a policy document into a constraint on when the company can ship.

OpenAI said it built the model with safeguards designed to keep the capability from being turned against real targets. The company reported that Astra refuses 91.5 percent of jailbreak attempts aimed at cyberattacks, up from 59 percent for its previous flagship, GPT-5.6 Sol—a shift that cuts the share of successful attacks by roughly four-fifths. It also restricted responses for accounts it considers high risk, and said the most advanced capabilities are available initially only through Daybreak, a coalition of cybersecurity companies that gets early access to use the technology for defense.

The disclosure lays bare the double edge of the technology. The same machinery that finds flaws so defenders can patch them can find flaws so attackers can break in. OpenAI argues that giving defenders the first look, combined with the stricter refusal behavior, tilts the balance in favor of protection. Skeptics note that autonomy is exactly what changes the economics of attack: zero-day research normally costs well-funded teams months of labor, and a model that automates discovery could put that capability within reach of far smaller actors.

The timing puts OpenAI out front of its rivals on a capability most of them measure but none has yet described in public with this level of detail. Anthropic maintains a similar tiered safety system and has said it will not deploy its most capable models until safeguards mature. Google has published its own risk frameworks. But no major lab has disclosed a model that clears a critical threshold in a domain with this much offensive potential.

The security research community is split. Some researchers welcomed the disclosure as an honest accounting of what leading models can already do, arguing that secrecy would leave defenders unprepared. Others warned that publishing the results, even without the technical details, accelerates the arms race by confirming what is possible and signaling where to look. OpenAI said it deliberately withheld the specifics of the vulnerabilities it found until vendors could respond.

The restrictions on Astra reflect how seriously the company takes the risk. High-risk accounts—those whose activity suggests they might use the model for intrusion—face a narrower set of responses, and the refusal rate improvements were aimed squarely at attempts to steer the model toward attacks. The Daybreak coalition gives vetted security firms early access, an arrangement designed to put the capability in defensive hands first while OpenAI watches how it is used.

For OpenAI, the post is also a piece of positioning. The company has been criticized for moving fast in areas where its technology could be abused, and the Astra disclosure lets it argue that its own framework is binding on its product plans. A model that reaches the top of the risk scale and still ships, with controls attached, is a different story from a model that ships and is graded afterward.

The industry is watching what happens next. If Astra’s safeguards hold as access broadens, the episode will stand as a template: measure the risk, build the model, gate the release, and let the framework decide the pace. If the controls fail, it will stand as the moment the industry’s safety ratings met reality.

OpenAI framed the post as one step on a longer path, with Astra’s release limited until the company learns how the safeguards perform in use. For governments and companies that build on OpenAI’s models, the message is that the frontier of AI capability has moved into territory previously reserved for nation-state hacking teams—and that the company believes it can keep that power on a leash.

Related Posts

  • September 6, 2026
  • 6 views
Anthropic Moves Its IPO Filing to Late September

The bankers and lawyers running Anthropic’s initial public offering had told investors to expect the company’s registration documents as soon as this week. The calendar has moved. Anthropic now plans…

  • September 6, 2026
  • 6 views
OpenAI Quietly Revises GPT-6 Astra Scores After Launch

When OpenAI released GPT-6 Astra on Sept. 3, the launch post carried the usual furniture of a modern model debut: coding results, speed comparisons and a figure for how often…