OpenAI published preliminary internal cybersecurity evaluations for Astra, an unreleased model, and concluded that it can no longer rule out the Critical capability level for cyber under its Preparedness Framework. This is the first time any OpenAI model has crossed out of the High band on that axis. Every prior system evaluated for frontier cyber capability, including GPT-5.6-Sol, was assessed at High rather than Critical.
The threshold itself is specific. Under the Preparedness Framework, a model reaches Critical for cybersecurity if it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or if it can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high-level desired goal. OpenAI says its evaluations of Astra over the past few days, combined with outside expert assessment, showed performance strong enough that the company concluded the night before publication that it could not rule the Critical level out. Benchmarking is still in progress, so this is a precautionary determination rather than a confirmed classification.
The response is a set of internal security controls rather than a deployment decision, because Astra is not deployed. OpenAI is implementing isolated testing environments, restricted network and tool access, enhanced model weight protections and encryption, additional monitoring and detection, and sandboxed execution for higher-capability models and the activities around them. Internal work involving Astra that does not yet meet those strengthened requirements has been paused. The company has also turned on universal monitoring for risky actions and misalignment across every agentic application of Astra, including training and evaluation runs; the monitors read the model's chain of thought and can trigger a security response that reviews and interrupts high-risk activity. External testing will run through relevant government agencies and selected AI safety organizations, and OpenAI says it will hand third-party testing partners a recommended set of security controls for running higher-risk evaluations safely.
OpenAI frames this as the same pattern it followed in June 2025, when its models approached the High biology threshold and it expanded safeguards, testing, and external expert involvement ahead of deployment. The framework was published in December 2023, well before models were anywhere near these levels, and this is the first time it has been used to declare a possible Critical crossing.
One clarification matters for reading the rest of the week's news: OpenAI states explicitly that Astra was not the model involved in the Hugging Face incident. Those were separate experimental models on a separate training track. TechCrunch's coverage frames the announcement as OpenAI slowing Astra's development over security concerns, which is accurate as to the internal pause but understates the framework claim, which is about capability measurement rather than schedule.
- OpenAI's own post leads with the Preparedness Framework threshold definition and stresses that benchmarking is unfinished.
- TechCrunch frames it as a development slowdown, emphasizing that the model can independently identify and carry out attacks against well-protected real-world systems.