Microsoft has released MAI-Transcribe-2, a speech-to-text model covering 60 languages at $0.10 per hour of audio. Five months ago the first model in the same line cost $0.36 and handled 25 languages. It is the third model in the line since the first shipped on April 2, and Microsoft claims first place on the FLEURS multilingual benchmark with a 5.2% word error rate, second place on the independent Artificial Analysis leaderboard, and inference ten times faster than OpenAI's GPT-Transcribe. The model is live in Microsoft Foundry, the company's model marketplace for developers, and in MAI Playground, its testing environment. The price is the announcement. The benchmarks are the argument for why Microsoft can charge it.
Most of what specialist vendors sell as add-ons is now in the base tier. MAI-Transcribe-2 ships with speaker diarization, which attributes each line to a specific person; word-level timestamps, which make a transcript searchable, editable and syncable to video; keyword customization for drug names, product codes and employee names; and automatic language detection, so the customer no longer has to declare the language of a recording in advance.
Microsoft highlights two further features. Configurable output style produces either a verbatim transcript that preserves every "uh", slip and stumble for legal and internal-compliance use, or a clean transcript with the filler stripped out for notes and subtitles. Language switching handles conversations where speakers change language in the middle of a sentence; Microsoft names Hinglish and Spanglish directly. Specialist providers have frequently charged separately for exactly these capabilities. Microsoft has folded them into the ten cents.
The company also says it trained on real business recordings — background noise, poor audio, several people speaking at once — rather than only on clean studio material.
Microsoft makes three performance claims, and each one rests on a different measurement.
The first is FLEURS: first place across 60 languages with an average word error rate of 5.2%. FLEURS is a benchmark published by Google researchers in 2022, in which native speakers read roughly 2,000 sentences in each of 102 languages, about 12 hours of speech per language. It is the standard way to compare multilingual systems on identical material — Swahili against Swedish, measured the same way. Word error rate counts substitutions, insertions and deletions against a human transcript, so 5.2% means roughly every twentieth word is wrong.
Two things qualify that number. FLEURS tests read, prepared text, not ordinary conversation. And the average went in the wrong direction: MAI-Transcribe-1.5 scored 3.7% in June. The likely explanation is coverage — the average now spans 60 languages instead of 43, including low-resource languages where every model does badly. That is a reasonable defense, and it is also the reason a buyer should ask for the per-language table rather than the headline.
The second claim is second place on Artificial Analysis by word error rate, plus a position on the Pareto frontier of accuracy and latency. Artificial Analysis is an independent evaluator that tests models through their public APIs, so it measures what a paying customer actually gets. Its index includes simulated AI-agent dialogues, European Parliament speeches and company earnings calls, weighted heavily toward English-language business conversation. In June it ranked MAI-Transcribe-1.5 third at 2.4% word error rate, behind Alibaba's Fun-Realtime-ASR-preview and ElevenLabs' Scribe v2, while making Microsoft the fastest model in the top ten. Moving to second place most likely means it has passed ElevenLabs. The Pareto claim means no competitor beats it on accuracy without giving up speed, or beats it on speed without giving up accuracy.
The third claim is throughput. By Artificial Analysis' measurements, MAI-Transcribe-2 runs 10 times faster than OpenAI's GPT-Transcribe, 7 times faster than ElevenLabs' Scribe v2, and 5 times faster than Google's Gemini 3.5 Transcribe. For batch transcription, speed is mostly a proxy for capacity and cost: a model running at 300 times real time consumes a fraction of the GPU-hours of one running at 30 times. That efficiency is the mechanism by which ten cents an hour can still be profitable.
The cadence is the part worth holding onto. MAI-Transcribe-1 arrived April 2 with 25 languages at $0.36 an hour. MAI-Transcribe-1.5 followed on June 2 with 43 languages, keyword customization and third place at Artificial Analysis. MAI-Transcribe-2 now offers 60 languages, diarization, timestamps, language switching, second place and $0.10. Language coverage has grown roughly 40% per step, and each step added a feature that competitors tend to keep behind a higher tier. That rhythm belongs to a team that has settled its architecture and is now scaling data and compute on a schedule — the phase in speech recognition where gains come quickly and predictably.
Mustafa Suleyman, who runs Microsoft AI, told The Verge in April that the first model came from a team of 10 people, deliberately insulated from bureaucracy, with a larger group handling vendors and data collection. He also said the model used half the GPU resources of other frontier models, which he called a significant saving for Microsoft. The Verge noted that Meta, Amazon, Google and Anthropic have tried similarly stripped-down structures. The transcription line is the clearest test of whether that structure produces commercial output rather than papers.
The strategic question has hung over Microsoft for four years. It has put more than $13 billion into OpenAI and runs OpenAI models inside Azure, Office and Copilot, which made building its own models look redundant. The answer is independence, and the timeline supports it. When Microsoft hired Suleyman out of Inflection AI in March 2024 along with most of that company's staff, Salesforce CEO Marc Benioff read it as a statement of intent; in January 2025 he told CNBC that Microsoft was building its own AI and would probably stop using OpenAI, because it wanted frontier models of its own, and that this was why it hired Suleyman. Benioff had his own interests — Salesforce competes with Microsoft and invests in Anthropic — but events largely proved him right. In October 2025 the two companies rewrote their partnership, and Microsoft's own announcement said it could for the first time pursue artificial general intelligence alone or with other partners. Suleyman told The Verge that the rewrite opened Microsoft's path to superintelligence; weeks later the company announced its MAI Superintelligence team. In April 2026 the terms changed again: Microsoft lost exclusive access to OpenAI's models, and OpenAI's revenue-share payments were cancelled. Each loosening was followed by new MAI models.
The second reason is margin. Every request Microsoft routes to an OpenAI model costs it money; a request to its own model on its own GPUs costs less. In July, Bloomberg reported that Microsoft had begun serving some user requests in Word and Excel with MAI models, products it had previously marketed as running on OpenAI and Anthropic. TechCrunch tied the shift to a broader pullback in AI spending, alongside cuts at Amazon, Uber, Meta and Accenture.
Transcription went first because it is the easiest thing to take back. The task is bounded and the quality is objectively measurable, and Microsoft happens to own the demand: Teams generates an enormous volume of meeting recordings, Nuance's medical documentation business runs on speech recognition, and Azure's speech services already serve thousands of companies. Every hour of audio moved onto MAI-Transcribe-2 is an hour Microsoft stops paying someone else for. Suleyman put the goal plainly to The Verge in April: the superintelligence conversation, in practice, comes down to whether models can deliver value to the millions of businesses that depend on Microsoft for frontier language models.
Here is how I read it. This is not Microsoft trying to beat GPT or Gemini at being a general model. It is Microsoft disassembling the frontier model into pieces small enough to own — pick a bounded domain, optimize inference cost hard, price under the frontier labs, distribute through Foundry, then quietly migrate its own products onto it. Microsoft AI now ships models for images, voice, transcription, code, reasoning and cybersecurity; at Build in June it introduced seven new MAI models in a single keynote. Speech is simply the first domain where the playbook has fully matured, which makes the MAI-Transcribe-2 release notes less a product page than a template for whatever comes next. And for all the two years Suleyman has spent talking about "humanist superintelligence" and assistants that act in the user's interest, the concrete result is this: five months ago Microsoft charged 36 cents to turn an hour of speech into text, and now it charges 10.
The competitor list is worth reading for who is on it. Microsoft names four products — OpenAI's GPT-Transcribe and Whisper V3-Large, Google's Gemini 3.5 Transcribe and ElevenLabs' Scribe v2 — and does not mention Deepgram, AssemblyAI, Speechmatics or Rev, all of which have been selling enterprise transcription for a decade. Positioning against frontier labs rather than incumbents is partly marketing and partly fair: the labs treated speech as a feature of a larger platform and priced it accordingly, and a dedicated model running 5 to 10 times faster at comparable accuracy genuinely is a different product. But the pricing pressure lands hardest on the specialists. At $0.10 an hour, Microsoft is at or below many high-volume enterprise contracts while including diarization, timestamps and 60 languages in the base tier. What the specialists retain is depth in specific verticals — medical vocabulary, legal formatting, industry integrations — and Microsoft's keyword customization aims squarely at that. The one competitor Microsoft does not claim to have beaten on accuracy is Alibaba, whose models have led independent rankings for most of 2026. TechCrunch reported in July that some American companies had started evaluating Chinese models as a cheaper option despite security concerns, and to those buyers Microsoft's pitch is legible: comparable accuracy, faster inference, lower price, and a vendor their compliance department already approved.
What the announcement does not say is the more instructive part. Microsoft calls $0.10 an hour a launch offer but gives no expiry date and no standard post-launch price, which is the first thing a finance team should get in writing. There is nothing about streaming: the post is built around batch processing and long recordings, and says nothing about real-time recognition, which is what voice agents and live captioning require. Artificial Analysis maintains a separate leaderboard for streaming models, which makes the silence conspicuous. There is no per-language breakdown, and a 5.2% average across 60 languages can easily mean 3% on common languages and 12% on thin ones. There is no diarization error rate or comparable metric, even though word error rate says nothing about whether a line was assigned to the right speaker — a transcript can score near-perfectly and still hand every other sentence to the wrong person. And there is nothing on data handling: no data residency, no retention period, no statement on whether audio uploaded to Foundry is used for further training. Microsoft's April materials described its training data as a mix of human-curated recordings, noisy audio collected by contractors, and "vast quantities of data from the open internet," per The Verge. Enterprise recordings contain medical, privileged and financial information, and for regulated industries those gaps are the questions to put to the vendor. None of this is unusual for a launch, but it is the distance between winning a leaderboard and getting into production.
The uncomfortable arithmetic sits with OpenAI. Microsoft's largest investment is also the vendor it is now undercutting by a factor of ten on a workload it controls end to end, and speech recognition is the easiest of those workloads, not the last.