What Claude learned to write for
New AI models are improving quickly at mathematics, programming and reasoning. Their writing quality has followed a different path: it has stopped improving, and in some cases has deteriorated.
Kernion recently described Opus 4.6 as Anthropic’s last good “writing model.” His explanation is partly familiar: later versions were optimized more heavily for mathematics and programming.
The more unusual factor is that Claude was also trained to produce technical explanations for other AI models. Kernion says the model was “adapted to the psychology of large language models.” It learned to write for AI systems rather than people.
His analogy is deliberately uncomfortable. Kernion compares the result to people who communicate only with other autistic people: within the group, a style develops that works for its members but is difficult for outsiders to follow. Large language models have much greater working memory than humans and can track finer-grained details, so training can reward prose that suits them while leaving people with blocks of overloaded text.
The reward problem
Kernion traces the problem to how reinforcement learning rewards are structured. Some rewards improve whether an AI model understands an answer; others improve whether a human does. The more a model is trained in mathematics and programming, he says, the more aggressively developers must compensate by rewarding explanations that are simple and easy for people to follow.
Anthropic appears to have found a better balance with Opus 5.5. Kernion writes that he had not been this satisfied with a model’s written language since Opus 4.6.
That is not the same as saying Opus 5.5 surpassed the older model. Kernion says the problem remains difficult and that Anthropic will continue improving its models.
My read is that this is less a prose bug than an objective mismatch. A model can become more capable while becoming less pleasant to read if the training process treats dense, machine-friendly explanation as a success. “Smarter” and “clearer” are not the same target.
What I’d want to know is how Anthropic measures the human side of that tradeoff. Kernion explains why the drift happens and says Opus 5.5 finds a better balance, but the account offers no human-readability measure to show how much better that balance is.
That leaves Anthropic with a problem that cannot be solved by adding more intelligence alone: every gain aimed at machines may need a separate reward to keep the model speaking to people.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X