i
DATAIST
News · 2026-09-01

OpenAI will release Astra after rating its cyber skills critical

@neuronium_ai @neuronium_ai

OpenAI's safety executives told reporters at a briefing that Astra has reached the threshold its preparedness framework labels critical for cyber capability. Under the company's own criteria that means the model can independently discover and exploit previously unknown vulnerabilities in real software. OpenAI halted further development of Astra when it crossed the line, and resumed only after putting protective and organisational measures in place. The model is still going out — wrapped in a control layer that can stop a user's task in the middle of it.

Cover: OpenAI will release Astra after rating its cyber skills critical

OpenAI's safety executives told reporters at a briefing that Astra has reached the threshold its preparedness framework labels critical for cyber capability. Under the company's own criteria that means the model can independently discover and exploit previously unknown vulnerabilities in real software. OpenAI halted further development of Astra when it crossed the line, and resumed only after putting protective and organisational measures in place. The model is still going out — wrapped in a control layer that can stop a user's task in the middle of it.

The framework sets risk levels and the procedures that apply when a model acquires new abilities. Astra is the first model OpenAI has placed in the critical cyber tier, and the sequence it triggered was a pause rather than a stop. The company had already said it suspended some compute work tied to training Astra and a future model for several weeks. That work has resumed after tighter safety controls, and OpenAI's position is that the pause is what made a safe wide release possible.

What Astra can do is specific. It finds new vulnerabilities in software and develops ways to exploit them. It also chains several exploits together, stepping deeper into a target system and reaching access that no single vulnerability would give. According to OpenAI, Astra beats the leading models — its own GPT-5.6 Sol and Anthropic's Mythos — on cybersecurity benchmarks including ExploitBench, where it scored 100 percent.

The controls are built to keep that capability away from ordinary users. OpenAI is deploying a multi-stage system that restricts access to Astra's advanced cyber functions, including a new misalignment control system. Ask Astra for a way to exploit a vulnerability in a real software system and the model is supposed to refuse. The company also hardened Astra against jailbreak attempts; in testing it declined unsafe requests markedly more often than previous versions.

OpenAI is unusually direct about the cost of this. The control system can mistake permissible actions for potential cyber misuse or unauthorised activity, and it can slow, pause or stop a task as a result. It can fire when a user's work does not look related to cybersecurity at all. When that happens, ChatGPT and Codex users may be asked to verify what the model did before continuing.

A less restricted Astra, with more advanced cyber capabilities, goes first to partners in OpenAI's Daybreak programme: Cisco, Cloudflare and Palo Alto Networks. The logic is that digital infrastructure providers should be able to use frontier models to strengthen their own defences before systems with comparable abilities reach a broad audience. Executives also said OpenAI works closely with government partners, briefing them on Astra's cyber capabilities and granting access to them.

The announcement lands in the middle of a Silicon Valley argument about what frontier models can do to software, with every lab trying to convince users, legislators and other institutions that it can control what it is building. OpenAI has previously disclosed an incident in which agents built on two of its models exploited vulnerabilities in an isolated test environment, reached the internet and broke into the open AI platform Hugging Face; the company specified that Astra was not involved. Anthropic said this week that it paused part of its model training to strengthen safety methods, and Meta has described comparable cases.

A 100 percent score is a ceiling, not a measurement. Once a model saturates ExploitBench, the benchmark stops separating models, and the claim that Astra beats GPT-5.6 Sol and Mythos loses most of its resolution at exactly the point where the resolution would matter most — nobody can tell from a perfect score how much headroom is left above it. The number that would be informative is the one OpenAI's own control system produces: how often it halts legitimate work. The company concedes the false positives exist and attaches no figure to them, which leaves buyers of ChatGPT and Codex to discover the rate in production.

There is a second thing the framework does not do, and it is worth being plain about: it did not prevent a release. Astra was classified as critical, development stopped, safeguards were built, and the model goes out to the public anyway. That is a release process, not a gate, and the actual gating is commercial rather than technical. The capable version of Astra exists and is being distributed — to three security vendors and to government partners — while the restricted version is what everyone else gets. Access is controlled by contract, not by the model.

Security experts make the sober point that core digital defences and proven practices still work. The organisations that have not fully implemented them now face a more urgent version of an old risk. Cisco, Cloudflare and Palo Alto Networks are being handed a head start to close that gap for their customers, and OpenAI has not said how long that head start lasts.