OpenAI has started rolling out GPT-6 Astra, the model it describes as its most capable yet, with an unusual caveat: some of its most advanced cybersecurity abilities are being restricted at launch. The company says Astra demonstrates state-of-the-art performance in coding, browsing and computer use, and aced several of the world’s toughest benchmarks.
Astra scored 98% on FrontierMath Tier 4, a set of math challenges that takes most human mathematicians weeks per problem, and 99.9% on ARC-AGI-3, which tests a model’s ability to learn new tasks. It also achieved a perfect score on ExploitBench, which measures the ability to find and exploit software vulnerabilities, landing the model in “critical” risk territory under OpenAI’s internal safety framework.
That cybersecurity capability forced a release delay. OpenAI disclosed Tuesday that Astra qualified as critical risk because it can hack many well-protected systems without human input, and engineers spent several weeks building guardrails. The company is initially holding back some of the model’s most advanced hacking abilities, a first for a frontier release.
“This is GPT-6 Astra. Anything you can do on a computer, Astra can do for you. Fast,” OpenAI wrote at launch. CEO Sam Altman told CNBC the model represents a new capability level and has already changed how he works.