Google has shipped Gemini 3.8 Flash, its third budget model in six weeks, while Gemini 3.5 Pro and Gemini 4 remain unreleased. The new Flash scores 73.7% on DeepSWE v1.1, a benchmark for long-running software engineering work, against 74.0% for Claude Opus 5 — three tenths of a point separating a model priced at $0.75 per million input tokens from one priced at $5.00. Koray Kavukcuoglu, the new head of DeepMind, has indicated that Google is pushing on base capability and not only on price-performance. The shipping record points the other way.
Google's own figures put the rest of the field further back: Claude Sonnet 5 at 53.8%, GPT-5.6 Sol at 72.7%, and the outgoing Gemini 3.7 Flash at 65.3%. More than eight points of gain over the previous Flash in a single generation is the kind of number that normally carries a Pro launch.
Google claims Gemini 3.8 Flash beats Anthropic's Opus 5 and OpenAI's GPT-5.6 Sol on many benchmarks despite costing far less. These are still only benchmarks, and in real conditions performance can feel quite different
Source: the-decoder.com
The launch price is $0.75 per million input tokens and $3.75 per million output — exactly what Gemini 3.7 Flash costs. Google says the standard rate will later rise to $1.50 and $7.50. Claude Opus 5 charges $5.00 and $25.00, GPT-5.6 Sol $4.00 and $20.00, so even after the introductory window closes Flash stays far below both on token price.
Token price is not the same as cost, and Google says so itself. Part of the improvement over 3.7 Flash comes from additional reasoning steps on hard problems and repeated calls out to tools. The model, in Google's phrasing, works harder, spends more tokens, and partly cancels out its own price advantage. For workloads where compute efficiency matters, the company advises turning the reasoning level down or staying on Gemini 3.7 Flash, which remains supported.
The independent numbers make that trade-off concrete. Artificial Analysis puts Gemini 3.8 Flash at 59 on its Intelligence Index, three points above 3.7 Flash at 56, and level with GPT-5.6 Sol at xhigh reasoning and Grok 4.6 at medium. The rise came mostly from agentic benchmarks — tool use and coding. On cost per completed task the model sits on the Pareto frontier at $0.58, the cheapest anything reaches at that level of intelligence. But that figure is up roughly 40% from the $0.40 that Gemini 3.7 Flash spent on the same work, with token prices unchanged.
Run the announced price increase through that and the picture changes again. At double the token rate and the same consumption, a task measured at $0.58 lands near $1.16 — close to three times what the previous Flash cost to finish the same job. That is my arithmetic, not Google's, and the company has not published a projection of its own.
Google also claims progress in 3D generation. A clip shows a 3D game that Gemini 3.8 Flash supposedly built from a single prompt inside Antigravity, Google's coding tool, with textures generated by the Nano Banana image model.
Source: the-decoder.com
Developers reach the model through Google AI Studio, Google Antigravity and Android Studio, enterprises through Gemini Enterprise. Consumers get it in the Gemini app, in AI Mode in Google Search, and, with a paid subscription, in Google Sheets.
At high reasoning the model generates around 300 output tokens per second and averages 2.5 minutes per task. GPT-5.6 Luna takes 2.6 minutes, GPT-5.6 Terra 3.3, Claude Fable 5.1 2.1, and Gemini 3.7 Flash 2.2. At low reasoning the average drops to about 48 seconds.
Three Flash releases in six weeks with no Pro and no Gemini 4 admits two readings. The one Google would prefer is that the cheap tier is improving fast enough to be worth shipping every time it does. The other is that cadence is covering for a launch the company cannot yet make. I lean to the second, and the composition of the gains is why: Artificial Analysis traces the Index rise chiefly to agentic benchmarks, and Google attributes the DeepSWE improvement to more reasoning steps and more tool calls. That is a model that does more work per task, which is a real product improvement and a different thing from a model that knows more.
The security variant is the part of the release with the least public surface. Gemini 3.8 Flash Cyber is not sold openly. As with Gemini 3.5 Flash Cyber before it, Google distributes it through the Fairwind program to government agencies, critical infrastructure operators and software developers, with looser safety settings because the intended use is defensive. On CyberGym, which tests vulnerability discovery in C and C++, Google reports 86.2%, against 77.5% for Gemini 3.5 Flash Cyber, 83.6% for GPT-5.6 Sol and 85.6% for GPT-5.5-Cyber. On CWE-Bench, an external benchmark for automated vulnerability repair, it reaches 47.2% Pass@1 against the frontier leader's 47.8%, at substantially lower cost.
The base model holds up unusually well against prompt injection. On Gray Swan IPI, 5.5% of attacks on Gemini 3.8 Flash succeeded, against 60.1% for DeepSeek V4 Pro, 52.7% for Kimi K3 and 51.8% for Grok 4.6. Claude Opus 5 did slightly better at 4.8%, and lower still with extra safety settings inside the Claude ecosystem. A gap of that size between two frontier labs and everyone else says more about the state of injection defenses across the industry than about the half point between Google and Anthropic.
Google's own recommendation is the sharpest thing in the release: if compute efficiency is what you need, use the older model. The budget tier has become more expensive in the only currency that tracks real spending — tokens per finished task — and the models that would justify paying more are still not out.