A voice built for more than one language
MAI-Voice-2.1 is designed to speak 23 languages in the same voice, with an accent that sounds natural in each language. Its faster variant, MAI-Voice-2.1-Flash, is priced at $15 per million characters, compared with $22 for the standard model.
Both models can clone a voice from a few seconds of audio and include safeguards against misuse. They are available through Microsoft Foundry and MAI Playground, as well as other platforms. Both are also available on OpenRouter.
In one test involving 4,000 participants, about half thought the generated voices belonged to a real person. That is a striking result, but the announcement does not say what the participants heard or how the test was run.
The missing evidence
I think the more consequential detail is the combination of a short cloning sample and speech that can pass for human to some listeners. Microsoft says safeguards are built in, but provides no detail here on what they do or how well they work.
The 150-millisecond latency gives the Flash model a clear product claim; the test offers a more ambiguous signal about quality. Without its methodology, it is hard to tell whether the result reflects a broad improvement in generated speech or a particular listening test. For agent builders, the gap between a voice that sounds convincing and one that is safe to deploy is the part this release leaves unresolved.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X