A benchmark claim with a gap
Microsoft says Decision-1 outperformed other models tested across 36 benchmarks covering nearly 150,000 questions. It also says the model is 2.5 times faster than H2O-Lightning-4B, the runner-up.
But Cloudflare’s open Clef models, also built on Qwen, were not included in the comparison. That omission matters: the speed claim is specific to the models tested, not a measure against every relevant alternative.
Microsoft's benchmark comparison puts Decision-1 at the top with 83.5% accuracy and 85 ms latency, ahead of Jev 1.13.0. | Image: Microsoft
Source: the-decoder.com
Decision-1 is available through Microsoft Foundry and OpenRouter. Input costs $0.042 per million tokens; output tokens are free.
The market Jev helped create
Jev helped spark interest in decision models in mid-September. Since then, the underlying idea has spread to open small language models, which have reportedly surpassed Jev’s technology. Microsoft’s entry is another sign of how quickly the category has moved from startup novelty to competition among larger players.
I think the more revealing test will be whether Decision-1’s benchmark advantage holds up outside the comparison Microsoft describes. Its price makes trying it straightforward; the missing Clef comparison leaves an important part of the case unresolved.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X