OpenAI released GPT-6 Astra on September 3, describing it as “the most intelligent and aligned model” it has built, while confirming it is also the first OpenAI system to reach the “Critical” cybersecurity capability level under the company’s Preparedness Framework.
A stronger, riskier model
Astra posts state-of-the-art results across coding, computer use, browsing and scientific reasoning, according to OpenAI, saturating benchmarks such as FrontierMath Tier 4 (98%) and ARC-AGI-3 (99.9%). But the headline number for security teams is ExploitBench, an internal red-teaming benchmark: Astra scored 100%, up from 78.5% for predecessor GPT-5.6 Sol, and reportedly surfaced two previously unknown zero-day vulnerabilities during testing.
Under OpenAI’s own framework, “Critical” means a model can identify and chain novel exploits against hardened, real-world systems without step-by-step human guidance. Reaching that tier triggers mandatory extra safeguards rather than a change in who can use the model.
Guardrails before wider release
To offset the risk, OpenAI says it added chain-of-thought monitoring across all agentic uses of Astra, flagging and interrupting high-risk activity in real time. The consumer-facing model refuses requests to generate proof-of-concept exploit code, which OpenAI says it did successfully in all test cases, compared with roughly half the time for GPT-5.6 Sol. Vetted defenders can apply for deeper access through OpenAI’s Daybreak program, aimed at security teams rather than the general public — a step that echoes the joint cyber-defense pact OpenAI signed with Anthropic and Google earlier this year, and puts Astra alongside Google’s own Gemini 3.8 Flash cybersecurity variant as a frontier model built with an explicit cyber-risk tier in mind.
Specs, pricing and rollout
Astra ships with a 1.05-million-token context window, a 128,000-token output limit and an April 30, 2026 knowledge cutoff. It costs $10 per million input tokens and $50 per million output tokens — pricier than its predecessor. OpenAI is rolling it out first to a limited set of organizations, with access for ChatGPT Plus, Pro, Business and Enterprise users, plus the API, Microsoft Azure and AWS Bedrock, arriving over the following days; enterprise customers reportedly have the higher-risk capabilities switched off by default.
Analyst Sanchit Vir Gogia told CSO Online the “Critical” label is mainly a transparency milestone rather than a sudden capability jump, but cautioned that the same update that makes Astra refuse harmful requests more reliably has also made its reasoning harder for outside researchers to monitor.