i
News
News · 2026-09-23

Google puts controllable voices into Gemini Flash TTS

@neuronium_ai @neuronium_ai

Google is introducing Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, two speech-generation models aimed at different workloads. Flash TTS can create a new voice from a text description, while both models support more than 100 languages and per-line instructions in dialogue. The pitch is less about simply turning text into audio than about giving developers control over character, delivery and long-form consistency.

Cover: Google puts controllable voices into Gemini Flash TTS

Two models, two jobs

Google positions the models differently:

Gemini 3.8 Flash TTS is for creative work such as game characters, audiobooks and podcasts.
Gemini 3.8 Flash-Lite TTS is for lower-cost, high-volume speech generation, including dubbing, audio content and voice agents.

Flash TTS can design a voice from scratch. A text prompt specifies its role, accent and vocal characteristics across multiple languages and dialects. Google also offers a library of more than 2,000 premade voices, including Mexican Spanish, Quebec French and Scottish English.

Voice cloning uses a 30-second audio sample to create a voice profile. The person being cloned must record spoken consent, and the voice in that consent recording must match the original sample. Google says every audio clip created by Gemini audio models includes an inaudible SynthID watermark to help identify AI-generated speech.

Google has also announced Voice Remixing, which will modify the timbre, pitch, speed and accent of library voices. The feature is not yet available.

Source: the-decoder.com

The useful part is control, not novelty

Both models accept instructions for individual lines or can interpret stage directions in a script on their own. Google says they can generate audio lasting several hours with minimal voice drift, meaning the voice changes little over time.

A two-voice mode creates dialogue from one script while preserving the distinction between the speakers. Scripts can also specify laughter, sighs and sounds such as “mhm” to control reactions and pauses.

Two tests used a premade voice with a style prompt designed to imitate an irritated Berlin resident speaking English with a strong German accent. The settings conveyed the accent and intonation convincingly, but both tests had a high-pitched background whistle in places. In one test, the voice also changed near the end of the clip.

Voice style prompt
A native German man from Berlin speaking English as a foreign language, with a thick, unmistakable German accent. He is clearly not a native English speaker: he pronounces English words the German way, applying German rhythm and intonation to English sentences. "Th" becomes "z" or "d" ("ze," "sink," "dat"), "w" becomes "v" ("vat," "vell"), and final consonants become harder ("goot," "bat" for "bad"). The "r" is guttural and throaty, never the English "r." Vowels are flat and short, lacking the softness found in American or British English.

The voice is nasal and slightly strained, in the mid-range, with a raspy quiver and a Kermit-like wobble, but grittier and less puppet-like. It sounds like a tired Berliner at 2 a.m.

Delivery: quick, clipped, choppy. He hesitates briefly when searching for an English word and sometimes inserts the German word flatly without translating it. Occasional exasperated upward pitch breaks. Dry, sardonic, unimpressed, and slightly annoyed that he has to explain anything at all—and even more annoyed that he has to do it in English.

The prompt specifies the speaker’s background, pronunciation of individual sounds, vocal texture, rhythm and attitude. It also asks the voice to hesitate, occasionally insert German words and speak in a dry, irritated manner.

The tests point to the appeal and the limitation at once. The model can follow unusually detailed direction, but a convincing character voice still depends on artifacts that become obvious over a long recording. I think that makes consistency more important than the size of the voice library: a voice that holds its identity for hours is more valuable than one that sounds impressive for a short sample.

Source: the-decoder.com

Access, pricing and the unanswered limits

Google is gradually making both models available through the Gemini API and Google AI Studio. Flash TTS is also available in Gemini Notebook, while Flash-Lite TTS is available in Google Vids. Access through the Gemini Enterprise API will arrive later, according to the company.

Agora, LiveKit, Pipecat and Vercel already support integration through the Gemini API. Google has not specified regional endpoints for the new models, although its previous speech-synthesis models supported data processing in the EU.

Data from the free tier is used to improve products, Google says. Data from the paid tier is not. Paid access is priced in US dollars per million tokens: text tokens are charged as input and audio tokens as output.

Google says one second of generated audio corresponds to 25 audio tokens, while one hour corresponds to 90,000 tokens.

25audio tokens per second
90,000audio tokens per hour
$0.81Flash TTS hour through 2026
$0.54Flash-Lite TTS hour through 2026

Through the end of 2026, an hour costs $0.81 with Flash TTS and $0.54 with Flash-Lite TTS. From January 1, 2027, those prices will rise to $1.62 and $1.08 respectively. Text tokens are charged separately.

The announcement is quiet about two practical questions: where developers can process this data, and how much quality control is needed when generated speech runs for hours rather than seconds. Google has addressed consent and watermarking, but the audible whistle and voice change in the tests suggest that production reliability remains part of the cost.

Source: the-decoder.com

Source: the-decoder.com

Daily AI news

Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.

Only what matters — every day

Follow on X