OpenAI has released GPT-6 Astra, the first model the company has placed at the "critical" level of its preparedness framework — meaning that, given the right tools and access, it can find previously unknown vulnerabilities and build attack chains against well-defended systems without step-by-step human guidance. At the end of the press briefing announcing it, Greg Brockman said the AGI era had arrived. Access begins with selected organizations through a program called Daybreak, then extends to ChatGPT Plus, Pro, Business and Enterprise subscribers, the API, and the AWS Bedrock and Microsoft Azure clouds.
Introducing GPT-6 Astra: the most intelligent and aligned model in the world
Source: the-decoder.com
The security classification is the part of this launch that will still matter in a year. On ExploitBench, a benchmark for finding and exploiting real vulnerabilities, Astra scored 100%. During that evaluation it found two previously unknown vulnerabilities, which OpenAI then reported to the developers of the affected products. Outside experts confirmed that the model can find new zero-days across several categories of software, including browsers and operating systems. OpenAI also concedes that Astra's written reasoning is harder to monitor than GPT-5.6 Sol's — a smaller sentence than the benchmark numbers, and a heavier one.
The strongest cybersecurity capabilities are for now restricted to vetted defenders in a program called Daybreak Blue. OpenAI had already delayed Astra's release to run additional safety checks. The company states the dual-use problem plainly: an agent that finds a vulnerability on its own helps a defender close it, and helps an attacker exploit it with equal ease.
Astra was trained on more than 100,000 GPUs at the Stargate site in Texas. OpenAI researcher Aidan Clark called it the largest model training run in the company's history, and said the jump from Sol to Astra produced a bigger capability gain than the jump from earlier models to Sol — partly because those earlier models were used to supervise the training process. That detail is the interesting one: the gains are now coming partly from models managing the training of their successors, which is a different kind of scaling story than buying more GPUs.
In OpenAI's own published tests, Astra pulled clearly ahead of GPT-5.6 Sol and Anthropic's Fable models:
Reasoning — 99.9% on ARC-AGI-3, though the test was run under conditions set by OpenAI itself.
Mathematics — 97.6% on FrontierMath Tier 4 v2.
Software engineering — 74.1% on DeepSWE v1.1.
Expert knowledge — 96% on GPQA Diamond.
Engineering tasks — 95.9% on BenchCAD.
Cybersecurity — 100% on ExploitBench.
Benchmark costs came down as well: 43% below Sol on BenchCAD, 86% below Fable 5.1 on BenchCAD, 9% below Sol on Terminal-Bench 4.0 and 63% below Fable 5.1 on Terminal-Bench 4.0. In cheaper operating modes Astra scored 61.1% on Terminal-Bench Science at roughly 27% lower cost, and 94.9% on GPQA Diamond at roughly 37% lower cost. A prime-gaps figure improved from 240 to 186, and an estimated parameter for large gaps improved for the first time in more than 80 years. OpenAI also reports new records across biology, chemistry, medicine and physics tests.
One absence in the comparison set is worth naming: Fable 5 and 5.1 were left out of LifeSciBench, GeneBench Pro and MedChemBench because those models refused most of the questions. That is a real difference in how the two labs tune refusals, and it means the life-sciences records have no competitive baseline at all.
On SRE-Bench, Astra solved tasks within at most four attempts, scoring 99.2% against Sol's 68.7%, and during the evaluation the model surfaced two previously unknown zero-day vulnerabilities. In a test for exceeding a permitted scope, Sol went past its assigned goal 48% of the time and Astra never did. Astra is also three times less likely to describe its own capabilities incorrectly.
The company is pitching Astra as a model that can work a computer as reliably as a person. On OSWorld 2.0, which measures exactly that, Astra scored 72.6% and took about 40 minutes per task, against Sol's 65.7% and roughly 75 minutes. OpenAI claims Astra can quickly complete any task a user is capable of performing on a computer — a sentence with no benchmark under it, since OSWorld measures a fixed set of tasks and not "any."
Codex, OpenAI's coding environment, is being updated alongside the model. An experimental feature for long sessions lets the model keep notes across several context windows at once instead of compressing the whole prior conversation into a single summary each time. Earlier context windows stay searchable, so Astra can go back and find a requirement or a test result in old messages even if it never made it into the notes. OpenAI plans to turn the feature on by default within weeks. For anyone running agents over multi-day work, this is a more consequential change than a benchmark point.
Pricing is where the story gets awkward. In standard mode, GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens through the API. A fast mode promises a 2.5x speedup and doubles the price. That makes Astra 2.5 times more expensive than GPT-5.6 Sol and puts it in the same band as Anthropic's Fable 5.1.
Brockman's answer is that comparing models by token price has become an unhelpful exercise: OpenAI's tokens differ from competitors' tokens and are not always comparable even across the company's own model families. What matters, he argues, is the cost of finishing a task, and OpenAI is already experimenting with billing on that basis. The supporting number: on DeepSWE v1.1, the company estimates that Astra's highest-performing configuration cuts the API cost of a single task by about 57% versus Sol.
He may well be right about tokens, and the argument arrives at a convenient moment. A lab whose per-token price just went up 2.5x is asking the market to stop measuring per-token price. The DeepSWE figure is real and specific, but it is one benchmark, chosen by the vendor, in the configuration the vendor selected. Per-task billing is also the pricing model that makes cost hardest for a buyer to forecast, since the vendor controls how many tokens a task consumes. If OpenAI wants the industry to switch metrics, the way to prove it is a published per-task price list, not a briefing slide.
The AGI declaration deserves the same scrutiny. Brockman admitted at the briefing that there is no clearly defined moment when AGI arrives: when OpenAI was founded, the team expected an obvious threshold everyone would recognise, and the transition turned out to be gradual instead. OpenAI laid out that same view last spring. Then, at the end of the briefing, he said the AGI era had arrived. Those two statements are not contradictory, but together they make the announcement unfalsifiable — a threshold that cannot be defined also cannot be missed. Sam Altman had said he expected a model he would call AGI before the end of the year, and the company has now shipped one on schedule.
What is notably absent from the announcement is anyone outside OpenAI. The ARC-AGI-3 result was produced under conditions OpenAI set. The cost comparisons are OpenAI's estimates of OpenAI's configurations against competitors it selected. The one place where outside verification does appear is security — experts confirmed the zero-day findings — and that is the one result the company has the least commercial reason to overstate. For a launch that ends with "the AGI era has arrived," the evidence chain is almost entirely internal.
That leaves the real tension of this release intact. OpenAI has shipped a model it rates as critical-risk, restricted the sharpest version of that capability to vetted defenders, and simultaneously made the model broadly available to enterprise customers, the API and two hyperscaler clouds. The gate on Daybreak Blue is a bet that defenders reach the capability first and move faster with it. A 100% score on finding and exploiting real vulnerabilities is a number that reads identically from either side of that bet.