Consistency across long recordings
ElevenLabs says v4 analyses tone, pace and script context. Users can direct delivery with tags or ordinary phrases, and supply phonetic spellings to clarify names and technical terms. The company says pronunciation control is more reliable, and that a narrator or character should keep the same sound even when individual lines are regenerated.
Tags in the script let users set emotions, pauses, and sound effects for each line. | Image: Elevenlabs
Source: the-decoder.com
A request can contain up to 10,000 characters, or roughly ten minutes of audio. Longer works such as audiobooks still need to be generated in sections; ElevenLabs says the voice’s pace and delivery should carry across those boundaries. In dialogue, AI speakers use the context of the whole scene rather than treating each line in isolation.
The model supports more than 90 languages, up from about 70 in v3. Cloned voices are meant to speak other languages with a natural accent without drifting back toward their original one. Professional voice clones, absent from v3, return in v4; Instant Voice Clone requires just ten seconds of audio.
Turbo puts latency in the foreground
ElevenLabs is also releasing v4 Turbo for real-time use, including customer-support calls and conversations with game characters. The company says developers previously had to trade expressive speech for speed. Turbo is its attempt to remove that trade-off.
In ElevenLabs’ tests, Turbo begins producing audible speech in 150 milliseconds. Cartesia Sonic 3.6 takes 262 milliseconds, while OpenAI’s GPT-4o mini TTS takes 814 milliseconds. ElevenLabs says it optimized Turbo together with its ElevenAgents platform.
In Elevenlabs' benchmarks, Eleven v4 Turbo starts producing speech much faster than the competing models tested. | Image: Elevenlabs
Source: the-decoder.com
On Artificial Analysis’s Provider Voice Arena ranking, Eleven v4 leads Cartesia Sonic 3.6 and Google’s Gemini 3.8 Flash TTS. It scored 91.7% on a pronunciation benchmark, compared with 85.6% for v3. In blind tests, about three-quarters of listeners preferred v4 to models from Cartesia, Inworld and Google.
In the company's blind tests, listeners rated Eleven v4 as more expressive in 65 to 81 percent of comparisons, depending on the competitor. | Image: Elevenlabs
Source: the-decoder.com
The improved cloning also gives ElevenLabs a route into dubbing: the company says an actor’s voice can be used across all supported languages. It licenses voices from their owners, and voice actors can list trained clones in the ElevenLabs library and earn money when paying customers use them. A separate marketplace already offers celebrity voices, including Michael Caine.
The price is temporary; the test is not
The standard API price is $80 per million characters for v4 and $40 for Turbo. Through October 12, those rates are reduced to $22 and $11. Subscribers to the $22-a-month Creator plan or higher can use v4 in ElevenCreative for two weeks at no extra charge, up to twice their monthly credit allowance. Artificial Analysis lists Sonic 3.6 at $49 per million characters and Gemini 3.8 Flash TTS at $16.49.
Both models are available through ElevenAgents, ElevenCreative and the API. By default, customer data is stored in the US. Enterprise customers can choose isolated storage in the EU, India or Singapore, though some processing may happen outside the selected region. In the EU, customers can keep API processing within the region by enabling zero-retention mode.
In mid-September, ElevenLabs also released Music 2.5, a model for creating fuller-sounding songs.
I think the more consequential claim is not the benchmark lead but the promise that a cloned voice remains consistent across languages and regenerated lines. That is what could make a voice useful across an audiobook, a live agent and a dubbed performance. But the announcement says little about how that consistency holds up in ordinary use, or how listeners judge a voice when they know it is synthetic. A lower introductory price can encourage trials; it cannot answer those questions.
Source: the-decoder.com
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X