OpenAI said on 7 August 2026 that it slowed development of Astra, an unreleased frontier model, after internal evaluations placed it at the Critical cybersecurity threshold of its Preparedness Framework. It is the first time a leading lab has publicly throttled one of its own models over offensive cyber capability.
Read the full story: OpenAI Pauses Astra Over First 'Critical' Cyber Rating →
Transcript
OpenAI just slowed down its own unreleased model, and the reason is the interesting part. Astra hit what OpenAI calls the Critical cybersecurity threshold. That is a specific, published bar. It means the model can find and write working zero-days in hardened real world systems without a human helping, or run an end to end attack given nothing but a goal. Here is why that matters more than it sounds. At the High level, you need safeguards before you ship. At Critical, the safeguards attach during development, because the risk is no longer just what customers do with it. It is the weights existing at all inside a company that can be phished like any other. So OpenAI restricted internal use and is lining up testing with government agencies. The catch: this is OpenAI grading OpenAI, and nobody outside has checked the work yet.