Two models, two jobs
Google positions the models differently:
Flash TTS can design a voice from scratch. A text prompt specifies its role, accent and vocal characteristics across multiple languages and dialects. Google also offers a library of more than 2,000 premade voices, including Mexican Spanish, Quebec French and Scottish English.
Voice cloning uses a 30-second audio sample to create a voice profile. The person being cloned must record spoken consent, and the voice in that consent recording must match the original sample. Google says every audio clip created by Gemini audio models includes an inaudible SynthID watermark to help identify AI-generated speech.
Google has also announced Voice Remixing, which will modify the timbre, pitch, speed and accent of library voices. The feature is not yet available.
Source: the-decoder.com
The useful part is control, not novelty
Both models accept instructions for individual lines or can interpret stage directions in a script on their own. Google says they can generate audio lasting several hours with minimal voice drift, meaning the voice changes little over time.
A two-voice mode creates dialogue from one script while preserving the distinction between the speakers. Scripts can also specify laughter, sighs and sounds such as “mhm” to control reactions and pauses.
Two tests used a premade voice with a style prompt designed to imitate an irritated Berlin resident speaking English with a strong German accent. The settings conveyed the accent and intonation convincingly, but both tests had a high-pitched background whistle in places. In one test, the voice also changed near the end of the clip.
A native German man from Berlin speaking English as a foreign language, with a thick, unmistakable German accent. He is clearly not a native English speaker: he pronounces English words the German way, applying German rhythm and intonation to English sentences. "Th" becomes "z" or "d" ("ze," "sink," "dat"), "w" becomes "v" ("vat," "vell"), and final consonants become harder ("goot," "bat" for "bad"). The "r" is guttural and throaty, never the English "r." Vowels are flat and short, lacking the softness found in American or British English.
The voice is nasal and slightly strained, in the mid-range, with a raspy quiver and a Kermit-like wobble, but grittier and less puppet-like. It sounds like a tired Berliner at 2 a.m.
Delivery: quick, clipped, choppy. He hesitates briefly when searching for an English word and sometimes inserts the German word flatly without translating it. Occasional exasperated upward pitch breaks. Dry, sardonic, unimpressed, and slightly annoyed that he has to explain anything at all—and even more annoyed that he has to do it in English.The prompt specifies the speaker’s background, pronunciation of individual sounds, vocal texture, rhythm and attitude. It also asks the voice to hesitate, occasionally insert German words and speak in a dry, irritated manner.
The tests point to the appeal and the limitation at once. The model can follow unusually detailed direction, but a convincing character voice still depends on artifacts that become obvious over a long recording. I think that makes consistency more important than the size of the voice library: a voice that holds its identity for hours is more valuable than one that sounds impressive for a short sample.
Source: the-decoder.com
Access, pricing and the unanswered limits
Google is gradually making both models available through the Gemini API and Google AI Studio. Flash TTS is also available in Gemini Notebook, while Flash-Lite TTS is available in Google Vids. Access through the Gemini Enterprise API will arrive later, according to the company.
Agora, LiveKit, Pipecat and Vercel already support integration through the Gemini API. Google has not specified regional endpoints for the new models, although its previous speech-synthesis models supported data processing in the EU.
Data from the free tier is used to improve products, Google says. Data from the paid tier is not. Paid access is priced in US dollars per million tokens: text tokens are charged as input and audio tokens as output.
Google says one second of generated audio corresponds to 25 audio tokens, while one hour corresponds to 90,000 tokens.
Through the end of 2026, an hour costs $0.81 with Flash TTS and $0.54 with Flash-Lite TTS. From January 1, 2027, those prices will rise to $1.62 and $1.08 respectively. Text tokens are charged separately.
The announcement is quiet about two practical questions: where developers can process this data, and how much quality control is needed when generated speech runs for hours rather than seconds. Google has addressed consent and watermarking, but the audible whistle and voice change in the tests suggest that production reliability remains part of the cost.
Source: the-decoder.com
Source: the-decoder.com
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X