OpenAI’s Preparedness Framework is the company’s internal rulebook for deciding how much danger a new model’s capabilities pose before that model ships — and what OpenAI promises to do about it. It splits a handful of high-stakes capabilities into two grades, High and Critical, and each grade is supposed to trigger specific safeguards, from restricting who can access a feature to pausing further development altogether. The framework became a live test case rather than a theoretical policy when OpenAI said its GPT-6 Astra model was the first of its systems to cross the Critical threshold for cybersecurity — a model capable of finding and exploiting unknown software flaws largely on its own.
What the framework tracks
OpenAI first published the Preparedness Framework in December 2023 and released a rewritten second version in April 2025. It deliberately covers a narrow slice of risk: not everyday problems like bias, misinformation, or copyright, which the company handles through separate policies, but capabilities that could open a genuinely new path to catastrophic harm. The current version tracks three such categories:
- Biological and chemical capabilities — could the model meaningfully help someone create a biological or chemical weapon?
- Cybersecurity capabilities — could the model discover and chain together unknown software vulnerabilities, known as zero-day exploits, against real, defended systems?
- AI self-improvement — could the model substantially speed up AI research itself, including work on its own successors?
An earlier version also graded persuasion and long-range autonomy as headline categories. OpenAI folded persuasion into other safety work and demoted autonomy to a research-only category it monitors but doesn’t yet score, in the 2025 rewrite.
How High and Critical differ
Each tracked category gets one of two ratings: High or Critical. (The original framework also had Low and Medium tiers, but OpenAI dropped them in 2025, saying they had never actually changed a real decision about a model.)
A High rating means a model meaningfully amplifies a path to harm that already exists — for instance, giving a complete novice enough uplift to attempt something that used to require specialist training. A Critical rating is reserved for a genuinely new path to catastrophic harm that has no real precedent. In cybersecurity terms, OpenAI defines Critical as a model that can identify and develop working exploits against many hardened, real-world systems without a human guiding each step, or one that can plan and execute an entire cyberattack from just a high-level goal.
What a Critical rating actually triggers
OpenAI’s written commitment is that a model isn’t deployed with a High or Critical capability exposed until safeguards can contain it, and that further advances in a Critical capability pause until adequate protections are defined. In Astra’s case, OpenAI says that meant delaying the model’s launch, retraining it to more reliably refuse harmful cyber requests, and shipping it with limits on access to its most advanced offensive capabilities rather than withholding the model entirely. The model’s system card is the main public document where OpenAI is supposed to show its work on this.
Why a self-graded framework matters
OpenAI isn’t required by any binding law to publish this framework or follow it — it does so as part of a voluntary pledge, alongside Anthropic, Google DeepMind, Microsoft, and more than a dozen other companies, made at the 2024 Seoul AI Safety Summit to publish safety frameworks and disclose how they handle severe risk. AI safety researchers have long pointed out the obvious tension: a lab is both the one setting the bar and the one grading its own homework against it. Anthropic runs a similarly purposed, differently structured system called its Responsible Scaling Policy; Google DeepMind and others have their own versions. None of these frameworks are enforced by regulators — they’re a bet that self-imposed rules, made public, are better than no rules while governments catch up.
In the news
OpenAI’s announcement that GPT-6 Astra crossed the Critical cybersecurity threshold is the first time a major lab has applied its own “Critical” label to a model it actually shipped, rather than to a hypothetical future system — making the Preparedness Framework’s promises concrete for the first time.