What Xiaomi shipped
MiMo-V2.6-Pro uses a mixture-of-experts architecture with 1.02 trillion parameters, although only 42 billion are active for each request. Xiaomi is also releasing the smaller, more economical MiMo-V2.6-Flash.
Artificial Analysis places Pro above Kimi K3 and Qwen among available open models. Its score and pricing put it on what the company calls the Pareto frontier: the zone where quality and cost are considered together.
The release includes a faster Pro-UltraSpeed version, with generation speeds of up to 20 times higher than the standard version.
The reinforcement-learning bet
Xiaomi attributes the improvement to scaled reinforcement learning. The company expanded the process in three directions:
The training process took less than six days, according to Xiaomi. It cost about $2.62 million for Pro and $0.85 million for Flash.
On the DeepSWE programming benchmark, Pro improved from 58.4 to 72.6 points. Flash rose from 48.8 to 65.7.
To keep training stable at this scale, Xiaomi fixed the model’s internal routing mechanism and added several layers of protection against reward hacking — techniques that let a model earn a high score without solving the underlying task.
Xiaomi is also releasing the tools it used for reinforcement learning:
The tasks come from multiple sources. Some code was taken from real GitHub pull requests created by employees, as well as user requests. Other task descriptions were generated by a language model.
The cybersecurity tasks are based on OSS-Fuzz, a collection containing tens of thousands of real software vulnerabilities. The office environments were recreated synthetically.
The question behind the openness
The release’s open tooling contrasts with allegations Anthropic made two weeks earlier. In a report from its threat-analysis team, Anthropic examined abuse of Claude detected from December 2025 through August 2026 and named seven Chinese laboratories associated with campaigns targeting the model:
Anthropic says the laboratories generated about 190 million message exchanges in total to extract Claude’s capabilities for training their own models. It describes the practice as illegal distillation.
Xiaomi received a separate mention. In case GTG-16008, Anthropic tracked more than 400,000 exchanges over 20 days in March and April 2026. According to Anthropic, Xiaomi sent user conversations and coding sessions from its MiMo models through OpenClaw and OpenCode into Claude to expand the dataset for training future models.
The report contains little information about the origins of the original training data or the teacher data previously used for internal distillation of teacher models.
That leaves the release with a clear tension. Xiaomi is making its reinforcement-learning framework and evaluation tasks unusually visible, but the public material does not answer where the earlier training inputs came from. I think that omission matters more than the impressive $0.13 benchmark cost: openness around the training machinery is not the same as openness about the data that made the model competitive.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X