OpenAI has replaced its image stack with ChatGPT Images 2.5, and it is two models, not one: GPT-Image-2.5 Flare, the default, which the company says produces higher-quality output than GPT-Image-2 at 50% lower latency, and GPT-Image-2.5 Sunburst, slower and built for harder image work and more precise edits. Both are in the API and rolling out worldwide in ChatGPT, ChatGPT Work and Codex, on desktop, mobile and web, free tier included. OpenAI says users already generate more than 3 billion images a week through ChatGPT Images and the GPT-Image models in the API. What the company does not say is which of the two models a given ChatGPT user is talking to.
The generation pipeline was reworked for more natural lighting and more accurate textures. The models hold objects from source photographs more faithfully and follow instructions more reliably across several consecutive edits. The faster of the two cuts generation latency by up to 50% against Images 2.0.
Pricing is where the launch gets interesting. Both models bill at the same rate: $8 per 1 million image input tokens and $30 per 1 million output tokens. Identical rates do not mean identical cost per image, because token consumption depends on the model and on the quality tier, and Images 2.5 adds two tiers, "xhigh" and "max", above the previous ceiling of "high".
For a 1024x1024 image, "low" costs roughly $0.006, the same as before. "High" lands near $0.053. "Max" runs about $0.21, burning roughly 7,024 output tokens. Put differently: "max" on Images 2.5 costs what "high" cost on GPT-Image-2. The quality ceiling moved up and the price ceiling moved with it, which is the honest reading of a tier structure that grows a new top end.
There is no cheaper batch rate for Images 2.5 yet. In early testing Sunburst generally cost more per image than Flare despite identical token pricing, most likely because it reasons longer. And unlike the previous launch, OpenAI does not publish an average cost per image this time. That omission is small and it is not accidental: an average is the one number a buyer can plan against, and it is the number that looks worse when a new top tier exists.
The routing question is the bigger gap. Neither the announcement nor the documentation says when ChatGPT reaches for the faster Flare and when it reaches for the more precise Sunburst. In the API you pick the model by hand. In the ChatGPT interface there is no such control. In testing, the dividing line ran mostly between Chat and Work modes: in Work, prompts changed only the elements named, regardless of reasoning settings. In Chat, follow-up generations kept altering extra details even at high reasoning effort. The stronger model appeared to engage only in "6 Pro" mode, and even then not consistently.
Source: the-decoder.com
Editing is the core of the release. The model is meant to change only what was asked and leave everything else intact, including complicated objects and backgrounds, and to preserve earlier edits across long conversations without quality decaying over successive rounds. OpenAI demonstrates this with a room redesign: the old model shifted additional details after every edit, the new one held the image stable across several iterations.
A quick test in ChatGPT Work with GPT-6 Astra (Max) shows the behaviour. The base prompt, then two edits in sequence, one small detail (the colour of a banana) and one large object (a big cat):
A hyper-realistic DSLR photo. A monkey holding a pink banana is sitting on a tiger in the foreground. In the background, a HORSE is RIDING AN ASTRONAUT. The astronaut is underneath, like a living "spacesuit horse saddle," and the HORSE is clearly on top, in control, as the rider. Make it 100% unambiguous: the HORSE is the rider and the ASTRONAUT is being ridden, NOT the other way around. High resolution, sharp focus, realistic lighting.
In these tests it is probably the best rendering of a horse riding an astronaut any OpenAI model has produced, with Image 2.0 in its Thinking variant shown for comparison. In Chat mode, each change to the banana's colour also touched other parts of the picture. In "6 Pro" the image held together on one run and not on another. Behaviour may shift as the rollout completes.
Images 2.5 is also meant to handle complex visual instructions better, render real-world facts more accurately, work with transparent backgrounds and build more complicated layouts.
An earlier test asked ChatGPT to lay text out as a 1980s magazine spread. That result looked like this.

Source: the-decoder.com

Source: the-decoder.com
GPT-Images-2.5 through Astra (Max) produced a detailed magazine, complete with a sample image from the monkey-and-astronaut prompt. It generated a weaker and a stronger variant, and held onto the original instruction even in the weaker one.
Source: the-decoder.com
Next the model was told to fix a deliberately planted error: the images on the first page were labelled GPT-Images-2.5 instead of 2.0. It corrected the label on command and left the rest of the layout alone. The same held for a translation issued as the short instruction Create a US English version.
Source: the-decoder.com
Regular ChatGPT Chat was checked for comparison. The weaker model changed other details while translating. In one case, though, it drew Microsoft's current chief executive more accurately.
Source: the-decoder.com
In ChatGPT itself, OpenAI is wrapping the new models in features. The main one is Sketch, invoked with @Sketch, which lets you draw directly in ChatGPT and use the drawing as a visual template for the finished image. OpenAI suggests it for diagrams, room layouts and posters. Alongside it come templates, ready-made prompts covering posters, logos, infographics, thumbnails, illustrations and ads, which replace the blank canvas with a structure and then sharpen the request through targeted questions. Users can leave comments directly on images and share the prompts they used so others can reproduce an idea with their own photos and details. OpenAI's example is the 1980s portrait prompt that went viral.
Source: the-decoder.com
On the Arena text-to-image leaderboard the two new models currently hold the top two places: Sunburst at 1421, Flare at 1399, GPT-Image-2 at 1381. Behind them sit Microsoft's mai-image-2.6 at 1331, SpaceXAI's grok-imagine-image-2.0 at 1315, and models from Reve, Meta, Google and Bytedance.
Those top two scores carry a "Preliminary" tag, and the vote counts explain why: about 3,100 for Sunburst and about 2,900 for Flare, against roughly 78,700 for GPT-Image-2. A 40-point lead built on 3,100 votes is not a result, it is an early reading, and the two new entries are separated from the model they replace by 40 and 18 points respectively. My read is that the leaderboard line is the weakest claim in this launch and the one most likely to be quoted anyway. The editing behaviour is the substantive change here, and it is exactly the thing an Arena head-to-head on single-shot text-to-image does not measure.
For provenance OpenAI continues to use the C2PA industry standard, embedding metadata for tracking, and is now adding Google DeepMind's invisible SynthID watermark across ChatGPT, Codex and the API. The company's stated reasoning is that no single provenance mechanism works, so it layers several. Adopting a rival lab's watermarking scheme is a small concession with a large signal inside it: on provenance, OpenAI has decided interoperability is worth more than owning the standard.
OpenAI promises shorter waits and higher usage limits with the rollout. Testing indicates Chat mode is currently served by the weaker model.
Source: the-decoder.com
Which leaves a company shipping its most controllable image model to date behind an interface that gives you no way to ask for it. Developers get a named model, a price sheet and a choice. Everyone else gets whatever the router decides, on a launch whose headline improvement, edits that change only what you asked, is precisely the thing that failed in the mode most people will use.