GPT-6 Astra Is Out: OpenAI's First 'Critical' AI Hits 100% on ExploitBench and Takes Over Your Desktop

OpenAI's newest flagship model, GPT-6 Astra, began rolling out to the public on September 5, 2026 — the same day its benchmark scores went viral. The model earned its name partly from Greek for "stars," but what's turning heads is something more grounded: it can take over your desktop and actually get things done.

Computer Use That Actually Works

Astra's most talked-about feature is computer use. Unlike earlier attempts, the numbers here are hard to dismiss. On OSWorld 2.0, the standard benchmark for GUI-driving AI agents, Astra scores 72.6% — and completes tasks 47% faster than OpenAI's previous frontier model, Sol. On ScreenSpot-Pro, which tests whether a model can correctly identify and click specific UI elements in dense, complex interfaces, Astra hits 92.7% versus Sol's 76.9%.

In practice, this means Astra can fill out web forms, update CRM entries, run front-end QA on a site it just built, and troubleshoot whatever it sees on screen. The 1-million-token context window means it can hold an entire codebase or long research document in working memory while doing so.

Benchmark Saturation

The scorecard for reasoning and science is striking. Astra hits 97.6% on FrontierMath Tier 4 — a set of competition-level math problems that stumped every model released before 2025. On ARC-AGI-3, the successor to the benchmark that once seemed AI-proof, it scores 99.9%. And on ExploitBench, a security research benchmark measuring autonomous vulnerability discovery, it scores a flat 100%.

That last number is why access isn't fully open yet. GPT-6 Astra is the first OpenAI model rated "Critical" for cybersecurity under the company's Preparedness Framework. Its exploit-creation capabilities ship gated behind the Daybreak programme, available first to vetted enterprise customers. General API access and paid ChatGPT plan rollout follow over the coming weeks.

Where It Falls Short

Astra is not a clean sweep. On Humanity's Last Exam with tools — arguably the hardest multi-domain benchmark right now — it scores 57.2%, trailing Claude Fable 5.1's 65.0%. That gap matters for the kind of deep, multi-step professional reasoning where Anthropic's model has historically excelled.

Pricing

GPT-6 Astra is priced at $10 per million input tokens and $50 per million output tokens, with cached prompt prefixes at $1.00 per million. A Fast mode offering 2.5× standard throughput is available at 2× standard price. For context, that puts it at the same price point as Claude Fable 5.1 — making the benchmark comparison between the two models especially direct.

The short summary: if you need a model that can operate a computer reliably, Astra is the current leader by a wide margin. If you need the deepest multi-step reasoning, Claude Fable 5.1 still has the edge on the hardest evals. The frontier has two flagships now, and they're genuinely close.


Sources: OpenAI announcement · Artificial Analysis benchmarks · Fortune