i
DATAIST
News · 2026-09-15

Gemini 3.8 Live undercuts GPT-Live-1 at $1.38 an hour

@neuronium_ai @neuronium_ai

Google DeepMind has launched Gemini 3.8 Live, a model it positions as the foundation for voice AI agents, priced at $0.005 per minute of audio input and $0.018 per minute of audio output. That works out to roughly $1.38 for an hour of voice conversation, against at least $3.00 an hour for OpenAI's GPT-Live-1. Example applications are published on GitHub.

Cover: Gemini 3.8 Live undercuts GPT-Live-1 at $1.38 an hour

Google DeepMind has launched Gemini 3.8 Live, a model it positions as the foundation for voice AI agents, priced at $0.005 per minute of audio input and $0.018 per minute of audio output. That works out to roughly $1.38 for an hour of voice conversation, against at least $3.00 an hour for OpenAI's GPT-Live-1. Example applications are published on GitHub.

The capability list is aimed squarely at people building production agents rather than at people watching a demo. Gemini 3.8 Live can call APIs in the background, process visual input, keep talking while it carries out other actions, and work in more than 97 languages.

The pricing structure is more revealing than the headline number. Output audio costs 3.6 times what input audio costs, which means the meter is driven almost entirely by how much the model talks, not by how long it listens. The $1.38 figure is also the ceiling rather than the expected bill: it assumes sixty minutes of input and sixty minutes of output, which is to say both parties speaking continuously for the full hour. A real support call, where the agent listens through long stretches of a customer explaining a problem, lands well below that. The comparison number runs the other way — $3.00 is described as a minimum for GPT-Live-1, so the true gap may be wider than the one the two figures show.

Source: the-decoder.com

Against that, OpenAI's model is probably the better conversational partner. GPT-Live-1 supports full duplex, meaning it can listen and speak at the same time, which is what makes an exchange feel like a conversation rather than a sequence of turns. Judging by the demos, its voice also sounds better.

So this is the trade Google has made again: cheaper and slightly worse, sold to developers rather than to reviewers. It is the same move the company has run through most of the Gemini line, and it works for the same reason it worked before — the person choosing a voice model for a contact centre is doing arithmetic across millions of minutes, and a 54 percent cut in the per-hour rate survives a certain amount of awkward turn-taking. It also happens that full duplex and per-minute audio pricing pull against each other: a model that holds the floor while you are still speaking is a model billing you for two streams at once.

The capability list is where the quiet part sits. Background API calls, vision, acting while talking, 97-plus languages — every item is about what the agent can do, and not one is about how the conversation feels. Interruption handling and end-to-end latency decide whether a voice agent is usable at all, and neither appears. "More than 97 languages" is its own small tell: 97 is a count, not a floor, and floors are not usually that precise.

Voice agents do not lose users on capability. They lose them in the first twenty seconds, when the model talks over someone or leaves a half-second of dead air after a question. Google has priced this one to be deployed at scale by companies that will measure it on cost per resolved call, and it may well win on that metric — while the axis it conceded is the one the end user actually experiences.