i
DATAIST
News · 2026-09-10

Claude Fable 5.1 hedges 36% less and flatters a third less often

@neuronium_ai @neuronium_ai

Claude Fable 5.1 leans on fewer verbal props than the model it replaces, and the gap has been counted. Stock phrasing such as "pivotal" shows up 20% less often per 1,000 words. Hedges like "possibly" and "arguably" are down 36%. Praise and agreement with the user's stated position appear in 1.98% of responses, against 3.17% in the previous version. The share of long content words fell from 42.6% to 38.6%, and abstract nouns dropped by a quarter, from 4.39 per 100 words to 3.28.

Cover: Claude Fable 5.1 hedges 36% less and flatters a third less often

Claude Fable 5.1 leans on fewer verbal props than the model it replaces, and the gap has been counted. Stock phrasing such as "pivotal" shows up 20% less often per 1,000 words. Hedges like "possibly" and "arguably" are down 36%. Praise and agreement with the user's stated position appear in 1.98% of responses, against 3.17% in the previous version. The share of long content words fell from 42.6% to 38.6%, and abstract nouns dropped by a quarter, from 4.39 per 100 words to 3.28.

Every one of those figures describes how the model talks, not what it knows. That is worth saying plainly at the outset, because a release measured in adjectives per thousand words invites the reader to treat prose hygiene as progress.

The sycophancy number is the one that earns its place. Going from 3.17% of responses to 1.98% is a relative cut of close to 37%, which is a serious move on a behaviour that has been the most-complained-about failure mode of chat assistants for two years. It also means roughly one response in fifty still opens by telling the user they are onto something. A third fewer is not the same as gone, and the remaining share is large enough that any heavy user will keep meeting it.

The density numbers point the same direction from a different angle. Long content words and abstract nouns are the raw material of text that sounds authoritative while committing to nothing; cutting abstract nouns by 25% is the difference between a paragraph that names a thing and a paragraph that names a category. A four-point fall in long-word share is smaller than it looks as a percentage but large enough to change the texture of a page.

The hedging figure is where I would slow down. A 36% drop in "possibly" and "arguably" makes for cleaner sentences, and it also makes for a model that sounds more certain than its predecessor across the board. Nothing in these numbers says it is more often right. Qualifiers are cheap to strip out of prose and expensive to earn back in substance, and a tuning pass that removes them uniformly removes them from the answers that deserved them too. Whether the remaining hedges landed on the genuinely uncertain claims is exactly the thing this set of statistics cannot tell you.

What the figures do not come with, as presented, is a test set, a sample size, or a description of what was asked. Percentages to two decimal places carry an air of measurement that bare percentages do not, and 1.98 against 3.17 reads as precision until you ask precision over how many responses, drawn from where. These are house numbers until the method behind them is visible.

Style metrics are the cheapest improvements to ship and the easiest to report, which is why they get reported. The harder question sits underneath them: a model that hedges 36% less and flatters a third less often is a model whose prose has been made more decisive, and the reader now has one fewer signal for telling a confident answer from a correct one.