i
DATAIST
News · 2026-09-04

Microsoft's MAI models run on half the GPUs and undercut AWS on price

@neuronium_ai @neuronium_ai

Microsoft released three of its own AI models on Thursday: MAI-Transcribe-1 for speech-to-text, MAI-Voice-1 for speech generation and MAI-Image-2 for images. Each was built by a team of under ten people. The transcription model runs on half the GPUs that competing frontier systems need, and all three are priced below what any other major cloud provider charges, Amazon and Google included. Until October 2025 Microsoft's contract with OpenAI barred it from doing any of this. None of the three models is a language model.

Cover: Microsoft's MAI models run on half the GPUs and undercut AWS on price

Microsoft released three of its own AI models on Thursday: MAI-Transcribe-1 for speech-to-text, MAI-Voice-1 for speech generation and MAI-Image-2 for images. Each was built by a team of under ten people. The transcription model runs on half the GPUs that competing frontier systems need, and all three are priced below what any other major cloud provider charges, Amazon and Google included. Until October 2025 Microsoft's contract with OpenAI barred it from doing any of this. None of the three models is a language model.

MAI-Transcribe-1 is the centrepiece and, according to Mustafa Suleyman, who runs Microsoft's superintelligence group, the first model the team has shipped. On FLEURS, the standard multilingual speech recognition benchmark, it posted an average word error rate of 3.8% across the 25 languages most in demand across Microsoft's products. In Microsoft's own testing the model:

beat OpenAI's Whisper-large-v3 on all 25 languages;

beat Google's Gemini 3.1 Flash on 22 of 25;

beat ElevenLabs Scribe v2 and OpenAI's GPT-Transcribe on 15 of 25 each.

The architecture is a transformer-based text decoder paired with a bidirectional audio encoder. It takes MP3, WAV and FLAC files up to 200 MB, and Microsoft puts its batch transcription speed at 2.5 times that of the company's existing Azure Fast service. Diarization, context handling and streaming are all listed as features arriving later. The model is already being tested inside Copilot's voice mode and in Microsoft Teams for transcribing meetings, which tells you how quickly Microsoft intends to swap out both third-party and older in-house models for its own.

MAI-Voice-1 produces 60 seconds of natural-sounding audio in one second and holds a speaker's voice steady across long material. Through Microsoft Foundry, a custom voice can be built from a few seconds of source recording. It costs $22 per million characters. MAI-Image-2 entered the Arena.ai leaderboard's top three model families on debut, generates images at least twice as fast as its predecessor inside Foundry and Copilot, and is being pushed into Bing and PowerPoint. It costs $5 per million input text tokens and $33 per million output image tokens. The advertising group WPP is among the first enterprise customers using it at scale.

The reason this launch is possible at all is contractual. The original 2019 agreement gave Microsoft a licence to OpenAI's models in exchange for building the cloud infrastructure OpenAI needed. When OpenAI decided to expand its compute beyond Microsoft and signed deals with SoftBank and others, the terms were renegotiated. Suleyman told Bloomberg in December 2025 that only weeks earlier Microsoft had held no contractual right to pursue general AI or superintelligence on its own. The new terms permit it to build frontier models while keeping its right to license anything OpenAI develops through 2032. He told VentureBeat that Microsoft reworked the contract in September of last year and then began assembling compute, hiring a team and buying data.

The partnership continues — Suleyman says Microsoft will work with OpenAI at least until 2032 and possibly longer, and called it an outstanding partner. Microsoft Foundry's API also serves Anthropic's Claude, part of Microsoft's pitch as a platform of platforms.

The timing is not incidental. Microsoft's shares closed their worst quarter since the 2008 financial crisis, and per CNBC are down roughly 17% year to date amid a broader slide in software stocks. Investors want evidence that the hundreds of billions poured into AI infrastructure will generate returns. Models that need half the GPUs cut Microsoft's own serving costs across Teams, Copilot, Bing and PowerPoint, while the same models are priced to compete for outside developers. Suleyman wrote in a March internal memo, reported by Business Insider, that his models had to deliver the unit-cost reduction required to serve AI workloads at enormous scale. These three are the first deliverable against that promise.

The team structure is the genuinely surprising detail. Ten people built the audio model. The image team is also under ten. Suleyman attributes most of the gains in speed, efficiency and accuracy to model architecture and the data used, and keeps a flat structure in which a small number of people get more autonomy. He described the working environment as closer to a young company's trading floor than a Microsoft engineering department: staff at round tables on laptops rather than large monitors, writing code together from morning to evening in rooms of 50 to 60. The implicit contrast is Meta, which by Suleyman's description to Bloomberg bet on recruiting large numbers of individual specialists rather than building one team, reportedly offering top researchers packages of $100 million to $200 million.

Suleyman has also been building a philosophy to go with the models. He calls it humanist AI, a term that appeared in his launch post and in the VentureBeat interview: superintelligence should serve humanity, people should remain the controlling force, and the system should act in line with human interests. In December he told Bloomberg that containment and alignment are red lines, and that superintelligence tools should not ship until their developers are confident they can control them. He frames training data provenance the same way, describing a conversation with Satya Nadella about building a clean lineage of models on quality data, and arguing that many open models were trained on improperly obtained material, which creates safety problems of its own.

Suleyman now describes Microsoft as a top-three lab, right behind OpenAI and Gemini. That is the one claim the launch does not support. All three models are narrow: audio in, audio out, images out. None of them touches the general reasoning and text generation that sit underneath ChatGPT and Copilot's core intelligence, which is exactly the capability the ranking is usually about. A world-class transcription model is a real accomplishment and not evidence about the thing it is being used to imply.

The small-team result deserves the same care. Speech recognition and text-to-speech are among the most tractable problems at the frontier: the objective is crisply defined, the benchmark is standard and public, and training data is abundant. Ten engineers beating Whisper is a strong signal about Microsoft's architecture and data work, and a weak predictor for a frontier language model, which is a different problem in data volume, compute and sheer complexity. Suleyman has the organisational authority, Nadella's public backing and the contractual freedom. What he does not yet have is a track record of building the hardest class of AI system inside Microsoft.

Some things are quiet in this announcement. FLEURS is a standard benchmark, but the head-to-head comparisons against Whisper, Gemini, Scribe and GPT-Transcribe are Microsoft's own tests, not audited ones. Diarization — knowing who said what — is the feature enterprise meeting transcription depends on most, and it is explicitly not shipped, yet the model is already being trialled in Teams. And the data provenance pitch, which is aimed squarely at enterprise buyers weighing AI vendors against a backdrop of copyright litigation, comes without any account of what Microsoft bought, from whom, or how a customer would verify the clean lineage being sold to them.

Suleyman has made clear that transcription, voice and images are the opening move. Asked about a language model that would compete with GPT at the frontier, he said Microsoft intends to build world-class models across every modality, and to be able to supply frontier-grade models itself, at high efficiency and minimal cost, if it needs to. He described a multi-year plan that includes standing up GPU clusters at the required scale. The superintelligence team was only formally created in October 2025. He gave the interview from Miami, where the whole team had gathered for one of its regular in-person weeks and where Nadella had come to lay out what Microsoft must achieve on AI self-sufficiency over the next two, three and four years, including the compute expansion that requires.

Two years ago, in MIT Technology Review, Suleyman proposed a modern Turing test: judge a system not by whether it can fool a person in conversation but by whether it can carry out real economic work with minimal supervision. Three narrow models shipped by small teams are a step toward that. They are also something more immediately awkward. Microsoft holds the right to license everything OpenAI builds through 2032, and it is spending that window building the models that replace third-party systems inside Teams, Copilot, Bing and PowerPoint. The licence runs another six years. Microsoft has started using them to make it optional.