From meme to model
Le Chaton Fat began as an internet joke, complete with fake benchmark charts and absurd claims about a French model supposedly poised to beat US and Chinese competitors. Some versions assigned it more than 30 trillion parameters and the ability to produce “1,000 meows per second.”
Mistral CEO Arthur Mensch joined in, writing on X that it was “actually a big kitten.” In July, TechCrunch reported that the company was preparing a large open-weight model, while Mensch and investor Marc Andreessen were also playing along with the meme.
Mistral co-founder Guillaume Lample told VentureBeat the company liked the joke. He described ML4 as a first version of the idea behind it, adding that even larger models could come later. The name Le Chonk, Mistral executives said, deliberately nods to the community that wanted the company to build a huge frontier model.
The joke was fictional; the trillion-parameter model is not. But the nickname is the lighter part of the story. The harder question is what that parameter count buys.
A trillion parameters, with 49 billion active
ML4 uses a sparse architecture: the model has one trillion parameters, but only 49 billion are active during inference. That lets Mistral claim a much larger overall capacity without activating the whole network for each response.
The company says training required 4,000 Blackwell accelerators, a relatively modest amount compared with the largest frontier-model projects. That is still a substantial cluster, and a direct efficiency comparison is difficult because competitors do not publish training-compute figures in a consistent format. Mistral’s previous sparse model, Large 3, had 675 billion parameters, with 41 billion active, and was trained on 3,000 Nvidia H200 accelerators.
ML4 accepts multimodal inputs but produces text only, Lample told VentureBeat. Mistral is targeting software development, cybersecurity, financial analysis, satellite and aerial imagery, technical drawings, and chip design.
The company gives cybersecurity particular strategic weight. Its argument is that businesses should not have to rely entirely on closed providers whose security systems may reject legitimate defensive, dual-use requests. Open weights, Mistral says, give security teams more control over code scanning, defensive testing, and other large-scale workflows.
Benchmark claims need outside checks
Mistral’s preliminary score for ML4 on DeepSWE v1.1, a benchmark for long-horizon software development tasks, is 62%. The company’s chart also lists Beam from Reflection AI at 44%, Qwen 3.8 Max at 51%, DeepSeek V4 Pro 0813 at 57%, and GLM-5.3 at 61%.
Some competitor figures can be traced to public sources, but results depend on benchmark settings. Reflection AI’s announcement for Beam listed 44.4% for Beam, 51% for Qwen 3.8 Max, and 61% for GLM-5.3 under its comparison setup. Artificial Analysis reported 57% for DeepSeek V4 Pro 0813 on DeepSWE using the Codex AI-agent environment.
The current DeepSWE leaderboard looks different when it uses each model’s best published configuration: GLM-5.3 and Kimi K3 score about 69%, while GPT-6 Astra, Gemini 3.8 Flash, and Claude Opus 5 score around 74%. ML4’s 62% is competitive, particularly against Western open-weight models such as Beam, but it does not establish that ML4 leads software development across available model and agent setups.
Other results make a stronger case, though with their own caveats:
At publication, ML4 was not yet listed in Artificial Analysis’s public evaluations or the DeepSWE leaderboard. Mistral calls it the strongest open-weight model developed outside China and says it can compete with leading Chinese open-weight systems. Those claims remain provisional until independent evaluators test the final model and released weights.
The business is bigger than the model
Mistral was founded in 2023 by former DeepMind researcher Arthur Mensch and former Meta employees Guillaume Lample and Timothée Lacroix. Its first model, Mistral 7B, arrived in September 2023; sparse Mixtral 8x7B followed later that year.
Since then, the company has added developer and enterprise products, model adaptation services, inference infrastructure, and Mistral Compute. The broader strategy is to sell companies control over deployment and the engineering around a model, not just access to the model itself.
In September, Mistral announced a €3 billion Series D at a post-money valuation above €21 billion. It called the round the largest equity raise in European technology-company history. Reuters valued Mistral at about $24 billion. The company says it now works with more than 125 international organizations, including Airbus, ASML, and HSBC.
Lample told VentureBeat that customers increasingly need more than model access: deployment, infrastructure, adaptation, and engineering support for complex AI workflows are part of the business. The bet is that model weights will become more interchangeable, while value shifts to the systems built around them.
I think that makes the three-week wait more consequential than the nickname. Mistral says the weights will arrive after 27 October; then customers can test the performance, cost, control, and customization claims on their own hardware. If the model delivers, open weights become evidence for Mistral’s infrastructure strategy. If not, a trillion parameters will be a striking headline attached to an unproven business case.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X