OpenAI says its forthcoming Astra model has reached the "Critical cybersecurity" capability threshold under the company's Preparedness Framework, a level that changes how the model will be released and monitored.

In a new safety post, OpenAI said Astra can identify and develop functional zero-day exploits in hardened real-world systems when given the right tools and access. The company also said the model scored 100% on ExploitBench, an evaluation for developing exploits from known vulnerabilities, and found two zero-day vulnerabilities during an internal benchmark involving recently disclosed V8 issues.

OpenAI says Astra was not involved in the earlier Hugging Face incident, but that the company folded lessons from that episode into its release plan. The company said it delayed parts of Astra's development and release while strengthening protections against cyber misuse and unauthorized model actions.

The initial rollout will not expose the model's strongest cybersecurity capabilities broadly. OpenAI says those workflows will first go to a small group of testers, with expanded defensive access planned through Daybreak Blue. The company also says it will use stricter behavior boundaries for accounts assessed as higher risk, additional monitoring of model reasoning and actions, and ongoing red-teaming.

TechCrunch noted that the release still leaves important questions unanswered, including who will test advanced access and how outside parties should evaluate OpenAI's claims. The conservative read is that Astra is both a capability milestone and a test of whether frontier labs can deploy powerful cyber models without making offensive use easier.