OpenAI's Astra Hits 'Critical' Cybersecurity Threshold — Can Now Find Zero-Days Autonomously

1011 010 101 011 01 10

OpenAI announced on September 1st that its forthcoming Astra model is the first it has ever rated "Critical" under its own Preparedness Framework — the highest risk tier the company defines for frontier AI systems. The designation means Astra can autonomously discover previously unknown security vulnerabilities and build working exploits across hardened real-world systems, all without a human guiding each step.

What "Critical" Actually Means

OpenAI's Preparedness Framework sets a Critical cybersecurity threshold when a model can identify and develop functional zero-day exploits of all severity levels in many hardened, real-world critical systems without human intervention — or when it can devise and execute end-to-end novel cyberattack strategies against hardened targets given only a high-level goal. According to CNBC, during safety evaluations Astra discovered and weaponized two zero-day vulnerabilities as part of a single exploit chain — entirely on its own.

That is not a theoretical benchmark. That is a model that can, in principle, compromise infrastructure at a level that has historically required teams of highly skilled human attackers.

A Delayed and Carefully Gated Release

Astra's path to release has already been rocky. OpenAI delayed the launch in early August after the model hit the Critical threshold during internal safety testing — a moment of genuine caution in an industry not always known for it. Now, the company says it will release Astra "soon," but access to its advanced cybersecurity capabilities will be tightly controlled.

The initial rollout goes through a program called Daybreak: a vetted coalition of cybersecurity organizations that can access these capabilities under strict safeguards. A broader Daybreak Blue tier will then extend access for purely defensive purposes — things like finding vulnerabilities in your own systems before an attacker does. General public access to Astra's full offensive capabilities is not on the table, at least not yet.

Why This Matters

The framing around Astra is careful — OpenAI is presenting it as a defensive tool first. That framing is probably accurate, and the Daybreak vetting process is a meaningful safeguard compared to deploying unconstrained. But the underlying capability is dual-use by nature. A model that finds zero-days in hardened systems is the same model regardless of who holds the API key.

What makes the Astra announcement genuinely significant is less the capability itself — sophisticated automated vulnerability research has been advancing for years — and more the candor. OpenAI publishing a Preparedness Framework with named thresholds, then acknowledging when a model crosses the worst one, is a more transparent posture than most frontier labs have maintained. Whether the framework's guardrails actually contain the risk is a separate, harder question.

TechCrunch's writeup captures the tension well: this is a model that is very good at breaking into computer systems, and it is about to be available to a carefully chosen set of organizations. Carefully chosen is doing a lot of work in that sentence.

For defenders, the optimistic read is that the same capability can find bugs faster than human red teams and patch them before adversaries exploit them. That would be genuinely valuable. The less optimistic read is that every safety measure is one data breach or one API leak away from being moot. Keep watching the Daybreak rollout closely.