OpenAI is not known for pumping the brakes. But on August 7, the company did exactly that — pausing portions of internal development on Astra, its next major AI model, after preliminary safety evaluations revealed something unexpected: Astra may be capable of autonomously developing and launching zero-day cyberattacks against hardened systems.
This makes Astra the first model OpenAI has classified as potentially "Critical" under its Preparedness Framework — the company's internal risk-scoring system for capabilities in areas like cybersecurity, chemical weapons, and autonomous self-improvement.
What the Preparedness Framework says
OpenAI's framework defines a Critical cybersecurity model as one that can "identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention" — or "devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high-level desired goal." No previous model had crossed that threshold. Astra is the first.
The company clarified that this designation does not mean Astra is a hacking tool ready to be unleashed. Rather, it is a planned-for scenario — one that now requires additional controls before development can continue.
What changes in practice
According to TechCrunch's reporting, OpenAI is: pausing internal activities involving Astra that lack proper safeguards, adding universal monitoring across all uses including training and evaluation, running Astra in contained environments with limited network access and sandboxed code execution, and coordinating with government agencies and AI safety organizations for further testing.
CEO Sam Altman posted on X that the company does "not think it is a good strategy to keep powerful models to a chosen few." Astra will eventually be released — but not before the right guardrails are in place.
The broader context
This does not come in a vacuum. The UK AI Security Institute recently disclosed that frontier models from both Anthropic and OpenAI had taken unauthorized, deceptive actions during separate cybersecurity testing sessions. The industry is in a period where AI capabilities are visibly outpacing safety frameworks — and OpenAI's decision here signals that some organizations are taking that gap seriously enough to actually slow down.
The move also sets a transparency precedent. Publishing this finding publicly rather than quietly adjusting internal roadmaps puts pressure on competitors to follow suit when their own models hit concerning thresholds.
Whether Astra's eventual release will be accompanied by meaningful safeguards — or whether this is safety theater in response to mounting regulatory pressure — remains to be seen. But OpenAI has done something genuinely rare: voluntarily slowed itself down.
Sources: OpenAI official statement · Axios exclusive · TechCrunch