i
News
News · 2026-09-30

Google’s Gemini 4 Argon leads in 13 of 18 benchmarks

@neuronium_ai @neuronium_ai

Google’s Gemini 4 Argon leads or ties in 13 of 18 benchmarks published by the company, putting it ahead of OpenAI’s GPT-6 Astra and Anthropic’s Claude Opus 5.5 by that measure. But the model is not broadly available yet: Google is starting with vetted cybersecurity specialists, while it completes security work and prepares a wider rollout. For enterprise buyers, the launch is a credible benchmark challenge—and a promise whose value still depends on access and performance in real workloads.

Cover: Google’s Gemini 4 Argon leads in 13 of 18 benchmarks

What the benchmarks show

Google’s comparison gives Argon the broadest spread of top results, not a clean sweep. It leads outright in 12 of 18 benchmarks and shares first place in one. GPT-6 Astra leads in three and shares one with Argon; Claude Opus 5.5 leads in two.

The strongest Argon results cluster around work businesses may pay to automate:

Harvey’s Legal Agent: 19.6% for Argon, versus 5.4% for GPT-6 Astra and 3.8% for Claude Opus 5.5.
AutomationBench, Zapier’s test of end-to-end business tasks: 51.3% for Argon, 42.5% for Claude Opus 5.5 and 41.4% for GPT-6 Astra.
GraphWalks, a test of graph traversal with large context: 84.2% for Argon, versus 71.8% for GPT-6 Astra and 66.8% for Claude Opus 5.5.
Vals Finance Agent v2: 65.4% for Argon, 58.6% for Claude Opus 5.5 and 53.5% for GPT-6 Astra.
DeepSWE v1.1, which evaluates long-running software development tasks drawn from real work: 77.9% for Argon, versus 74.2% for Claude Opus 5.5 and 74.1% for GPT-6 Astra.
LVBench, a test of long-video understanding: 91.7% for Argon, 87.5% for GPT-6 Astra and 83.7% for Claude Opus 5.5.

The contest is closer in other technical areas. Argon and GPT-6 Astra both score 68% on CWE-bench v1, which tests vulnerability remediation; Claude Opus 5.5 scores 67%. GPT-6 Astra leads Argon by 10.5 percentage points in both FrontierSWE v2 and Terminal-Bench Science 0.1. Claude Opus 5.5 leads Argon by 9 points on Terminal-bench 4.0 and also leads on PostTrainBench, with 49.3% against Argon’s 45.3%.

That makes “most first places in Google’s comparison” a better description than “best at everything.” The benchmark spread gives Google a stronger case for broad capability, while the gaps show that OpenAI and Anthropic retain advantages in particular coding, terminal and post-training tasks.

Argon also supports up to 1 million output tokens, compared with its previous limit of 64,000. That could matter for agents handling software development, audits, migrations and legal review, where a model’s usefulness can depend on how long it can continue a task before handing control to a person.

18benchmarks
13first or tied
1 millionoutput tokens

From benchmarks to Google’s own systems

Google says thousands of employees already use Argon for specialized coding, deep research and writing. The company also points to several internal applications:

Researchers in quantum computing used Argon to optimize space-time use for resource-intensive subroutines. Google says it beat the published baseline by 40% in a few minutes.
Agents analyzed profiling data across Google’s server fleet and identified memory optimizations. Once implemented, those changes are expected to free up more than 300 TiB in Google data centers, with total expected savings of 500 TiB to 1 PiB.
Agents are helping migrate codebases from C/C++ to Rust, including core libraries and the Zircon kernel of the Fuchsia operating system.

The most detailed software example is Google’s open-source video decoder, libgav1. According to the company, Argon agents took an existing Rust port and replaced 32,000 lines of SIMD code. They ran profiling-guided experiments, examined compiler output and produced memory-safe Rust that the compiler could automatically vectorize. Google says the resulting memory-safe decoder runs 2.7 times faster than the Rust version while producing identical video output.

Cybersecurity is the other major focus. Google says Argon can independently find and verify critical software vulnerabilities, then fix them. The company plans to give vetted defenders and its own teams access without cybersecurity restrictions, so they can use the model’s full capabilities for defensive work.

Google says Wiz is already using Argon through Scan for Good, an initiative to protect critical public infrastructure. The model found a critical vulnerability that could have exposed confidential personal data in medical software used by hospitals around the world.

Google is also highlighting safeguards before broader release. The company says it is strengthening protections against misuse in cybersecurity and chemical, biological, radiological and nuclear threats, as well as indirect prompt injection, model misalignment and unsafe AI-agent environments. In the Gray Swan benchmark for resistance to indirect prompt injection, Google reports a 0.7% attack success rate for Gemini 4 Argon. The reported rates are 1.0% for Claude Opus 5.5 and Claude Fable 5.1, 8.5% for GPT-6 Astra, 27.0% for GPT-6 Sol, 31.5% for GLM 5.3 and 51.8% for Grok 4.8.

I think the order of release is part of the product’s pitch: Google is leading with controlled cybersecurity access rather than immediate general availability. That may help establish a defensive use case, but it also leaves customers waiting to test the model on their own systems.

The price of a lead

During the introductory period, Argon’s API will cost $2 per million input tokens and $10 per million output tokens. Cached input tokens will cost 95% less, or $0.10 per million. After the introductory period, the prices will rise to $4 per million input tokens and $20 per million output tokens. If the cached-input discount remains, those tokens will cost $0.20 per million.

At the introductory price, Argon costs one-fifth of GPT-6 Astra’s listed API price of $10 per million input tokens and $50 per million output tokens. It is also half the price of Claude Opus 5.5, at $4 per million input tokens and $20 per million output tokens.

After the introductory period, Argon’s standard price will match Claude Opus 5.5 and remain below GPT-6 Astra. Anthropic lists Claude Opus 5.5 cached input at $0.20 per million tokens for reads and $5 for writes; Google’s published rate for Argon’s cached input is lower during the introductory period and comparable afterward.

Access is the immediate constraint. Google is first providing Argon to vetted cybersecurity specialists through its Fairwind program and is participating in a voluntary US government program for early access to models. The company plans to expand to developers, organizations and individual users, starting with paid API customers and Google AI Ultra subscribers.

My guess is the introductory price is meant to make trials easier, but it cannot answer the question enterprise buyers care about most: whether the benchmark advantage survives in production. The announcement says little about rate limits, data governance, deployment options or how Argon will integrate with Gemini API, Vertex AI, Google Cloud, Workspace and developer tools.

A response to the slowdown narrative

The timing gives the launch added weight. Last week, The Verge reported that new Google DeepMind head Koray Kavukcuoglu said Gemini 4 was in refinement and could arrive well before the end of the year. The publication noted that Google had not released a new flagship model since the Gemini 3 series in November 2025, while OpenAI and Anthropic had moved ahead with GPT-6 and newer Claude models.

Earlier reports had also described concerns about delays, costs and organizational changes. In July, Reuters reported that Alphabet investors were worried about delays to Gemini 3.5 Pro, rising AI infrastructure costs and departures from development teams. In August, Axios reported leadership changes: Demis Hassabis left his role as CEO of Google DeepMind to become Alphabet’s chairman and chief scientist; Jeff Dean left his role as chief scientist to start a company with other AI researchers; and Kavukcuoglu took over DeepMind, reporting to Sundar Pichai. Axios called it Google’s biggest AI leadership reshuffle since the turmoil at OpenAI in 2023, while noting that Google did not link the changes to model delays. MarketWatch reported tensions over whether Google’s AI unit should prioritize research, frontier capabilities or commercial products.

Argon gives Google a more substantial response to the idea that its research strength has not translated into the pace of frontier releases. But a benchmark table is still a company’s own comparison, and controlled early access is not the same as broad customer evidence. The next test is whether businesses can get the model, integrate it and reproduce its strengths on work that matters to them.

Daily AI news

Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.

Only what matters — every day

Follow on X