OpenAI said it cannot rule out that its unreleased Astra model has reached the “Critical” cybersecurity capability level under its Preparedness Framework, the company’s system for tracking dangerous AI capabilities before public release. It is the first time any OpenAI model has triggered that threshold, and the company has paused parts of Astra’s internal development while it tightens safeguards.
What “Critical” means
Under OpenAI’s framework, a model reaches Critical cyber capability if it can independently identify and build functional zero-day exploits against many hardened, real-world systems, or devise and carry out an end-to-end cyberattack strategy against a hardened target from only a high-level goal. OpenAI said recent internal evaluations showed Astra’s coding and cybersecurity performance was strong enough that it could not confidently rule out the model had crossed that line. Every previous OpenAI model assessed for frontier cyber ability, including the recently released GPT-5.6-Sol, topped out one tier lower, at “High.”
Astra had already drawn attention after it reportedly solved ten unpublished math problems during testing; OpenAI has not said when, or whether, the model will be released publicly.
Tighter guardrails
OpenAI says it is pausing internal Astra work that doesn’t meet a stricter security bar, and has added isolated testing environments, restricted network and tool access, and stronger protections and encryption for the model’s weights. It is also rolling out universal monitoring for risky actions across all of Astra’s agentic uses, including training and evaluation, and says it is coordinating with government agencies and outside AI safety organizations on further testing, while giving third-party evaluators guidance on the security controls needed to test the model safely.
The company has taken comparable precautions before, citing similar restrictions it imposed in June 2025 when an earlier model approached a high-risk biological-capability threshold. For a broader look at how AI labs are building cyber-specific safeguards, see our explainer on cybersecurity foundation models.