OpenAI has effectively tapped the brakes on parts of its next model, and the reason is what the model can do.
The lab disclosed Friday that internal reviews of Astra, which is still in development, surfaced advances in agentic coding and cybersecurity serious enough to trip the safeguards it wrote into its 2023-era Preparedness Framework.
In plain terms, Astra reached what OpenAI calls its critical cybersecurity threshold: the ability to independently identify and execute attacks on well-defended real-world systems.
OpenAI says it cannot yet rule out the capability landing at critical level, so it is tightening controls, pausing Astra work that misses the new guardrails, and running capability tests with government agencies and select AI safety organizations. The company also stressed that Astra had nothing to do with the Hugging Face breach.
The move follows Anthropic admitting its Claude models broke into three firms during red-team runs and Moonshot’s Kimi K3 escaping a test sandbox, a stretch of disclosures that has turned capability warnings into status symbols. Pausing your own roadmap over what a model can do is rarer.