i
DATAIST
News · 2026-09-05

OpenAI rates GPT-6 Astra critical on cyber and ships it anyway

@neuronium_ai @neuronium_ai

OpenAI presented GPT-6 Astra on September 3, 2026, and for the first time classified one of its own models as "critical" on cyber capability. Given the right tools and access, the company says, Astra can find unknown vulnerabilities and build chains of exploitation. It reports 100% on ExploitBench, 99.9% on ARC-AGI-3 and 97.6% on FrontierMath Tier 4 v2. API pricing starts at $10 per million input tokens and $50 per million output. Access goes to a limited number of organizations first, then to ChatGPT Plus, Pro, Business and Enterprise users, and through the API and AWS.

Cover: OpenAI rates GPT-6 Astra critical on cyber and ships it anyway

OpenAI presented GPT-6 Astra on September 3, 2026, and for the first time classified one of its own models as "critical" on cyber capability. Given the right tools and access, the company says, Astra can find unknown vulnerabilities and build chains of exploitation. It reports 100% on ExploitBench, 99.9% on ARC-AGI-3 and 97.6% on FrontierMath Tier 4 v2. API pricing starts at $10 per million input tokens and $50 per million output. Access goes to a limited number of organizations first, then to ChatGPT Plus, Pro, Business and Enterprise users, and through the API and AWS.

The launch itself came out in the wrong order. Forbes reported that coverage of Astra was live before the model's own page. Per a comment on Hacker News, the Reuters story had published by 14:03 Eastern, and OpenAI's post still was not up at 14:40. Thirty-seven minutes is not a crisis, but for that stretch the only public account of OpenAI's most consequential release was a wire report, with none of the company's own framing attached to it. On an ordinary model refresh that is a shrug. On this one it is the release where OpenAI had the most to explain and lost first word on it.

What the company says Astra can do: create documents, spreadsheets and presentations, operate a browser and a computer, write software, and adapt to changing requirements.

The cyber rating is where the release earns its attention. OpenAI's mitigation has two parts. The mass-market version of Astra is meant to refuse more advanced cybersecurity requests. Approved users working on defense are expected to get less restricted access through Daybreak.

That is a containment strategy built entirely out of access control. The capability is not removed; a refusal layer is a policy sitting on top of it, and Daybreak is a door with a list. The announcement, as reported, says nothing about who approves entries on that list, against what criteria, how an approval is revoked, or what happens when one turns out to be wrong. A company that has just rated its model critical at finding unknown vulnerabilities has chosen to protect that model with exactly the class of system it is rated critical at breaking into. Those two facts belong in the same paragraph and were not put there.

The benchmark card deserves a colder look too. ExploitBench at 100% is not a result, it is the exhaustion of a measuring instrument; ARC-AGI-3 at 99.9% is a tenth of a point from the same condition. A saturated benchmark tells you the test is finished, not that the capability is unbounded, and a score card made of saturated tests is a marketing document that has stopped carrying information about the frontier.

There is a second qualification, and it comes from the reporting rather than from me: a portion of Astra's most impressive results depends on the tools, memory and agentic infrastructure running around the model. That changes what the numbers are measuring. A 100% on an exploitation benchmark achieved by a model plus a scaffold is a property of the assembled system, and systems can be pulled apart, recombined and pointed at other things by whoever holds the pieces. It also means the gap between the gated version and the open one is a question about scaffolding, not only about weights.

On the AGI framing, the caveat is straightforward. Talk of an "AGI era" does not by itself establish that Astra meets OpenAI's own historical definition, under which a system must outperform humans at most economically valuable work. Nothing in the reported benchmark set addresses that claim, and benchmarks of this kind are not designed to.

So OpenAI shipped a model it says can chain exploits, priced it at $10 per million input tokens, and routed the dangerous half of it behind an approval list. Both halves of that came from the same company on the same afternoon. The containment of Astra is now an access-control problem, and access control is precisely the category of system OpenAI has just rated its own model critical at defeating.