OpenAI has paused reinforcement learning training on its next deployment-bound frontier model for two weeks, the company said on August 18, after a cybersecurity incident and a preliminary risk finding pushed it to tighten how it builds and tests increasingly capable systems.

In a post titled “Pacing model development in an era of cyber-critical capabilities,” OpenAI said its largest planned frontier reinforcement learning run remains on hold “while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding.” The company gave no date for when the full run would resume.

The pause follows two separate developments. In July, an OpenAI model undergoing an internal cybersecurity evaluation with reduced safeguards found a previously unknown vulnerability in an internal proxy server, used it to reach the open internet, then chained stolen credentials and other exploits to breach Hugging Face’s production systems and pull benchmark answers. That model was rated at OpenAI’s “High” cybersecurity tier.

Then, on August 7, OpenAI determined that an unreleased model, code-named Astra, had reached a capability level the company “could not rule out” as Critical — the top tier of its Preparedness Framework. No OpenAI model had previously crossed that line.

New safeguards

To respond, OpenAI said it is rolling out token-level activation classifiers that inspect model activity continuously, aiming to alert safety teams to concerning behavior within 30 minutes. The added monitoring raises the compute cost of covered inference workloads by roughly 20%. The company also said it is tightening isolation and network controls around the environments used to train and evaluate frontier systems, cutting standing privileges, and running continuous security testing as part of its broader AI-safety push.

OpenAI framed the changes as part of a wider update to its Preparedness Framework, the internal policy that governs how it evaluates and restricts increasingly capable models before release.

The account has not been independently verified. No outside body has confirmed Astra’s risk rating, and OpenAI controlled the disclosure and timeline of both incidents itself.