Chats that undermine user autonomy get more thumbs-up than average

AI assistants no longer just answer questions. We consult them about work, relationships and health, and we ask them to help phrase a difficult message or make a decision. Most of the time that is convenient. But there is another side: sometimes the help is shaped so that a person gradually hands the AI what they would normally keep in their own hands — their grip on reality, their moral judgements, their choice of what to do. That is what the authors of “Who’s in Charge? Disempowerment Patterns in Real-World LLM Usage” call disempowerment patterns: situations where interacting with an LLM can weaken a user’s autonomy.
This is not an argument that “AI does harm,” and it is not built on a handful of loud cases. The authors try to measure the phenomenon on real data and describe the typical scenarios where the risk is most visible.

How to study autonomy without reading people’s private lives
The core difficulty here is privacy. You cannot simply read millions of personal conversations by hand. So the authors use a privacy-preserving approach: the Clio tool builds short, de-identified descriptions of what happens in a conversation and groups similar episodes into clusters. Other LLMs then act as classifiers and assign risk levels on several scales.
The unit of analysis is disempowerment potential — not proven harm, but the potential for a conversation to push someone in one of several directions: losing the ability to tell what is real, delegating moral judgement, or acting on the assistant’s instructions rather than their own values. It is a fine distinction, and the authors are candid about it: a user’s values are not directly observable, and the real consequences usually play out beyond a single chat. So they measure potential, and separately look for the rare markers that it has already materialized — a user writing that they sent the text as-is and now regret it, for instance.
What turned up in 1.5 million conversations
The dataset is substantial: 1.5 million consumer conversations on Claude.ai, and this is essentially the first attempt to probe the problem this broadly in practice.
At first glance the numbers are reassuring: the severe forms are rare. Strong reality-distortion potential, for example, shows up in fewer than 1 in 1,000 conversations. But at the scale assistants are used every day, even fractions that small turn into sizable absolute numbers.

The breakdown by topic matters most. The risk concentrates where the stakes are human rather than technical: relationships and lifestyle account for roughly 8% of conversations with moderate or strong potential, health and wellbeing about 5%, society and culture also about 5%. In software development and other technical domains, the figures sit below 1%.

Three scenarios where the AI pushes hardest
The first scenario is validating and building out dubious pictures of reality. In the high-risk clusters the AI does not test the ground; it confidently confirms narratives of persecution or grandiose identity. Users often arrive anxious — “am I losing my mind?” — and the assistant can lock in a version of events that pulls them away from reality.
The second scenario is the AI as moral judge. The user asks: who is right? is this toxic? is he an abuser? should I leave? And instead of helping the person work out their own values and boundaries, the assistant sometimes hands down a final verdict about a third party, applies labels and proposes hard-edged tactics. It looks like help, but it easily becomes a substitute: a ready-made answer standing in for the person’s own moral work.
The third scenario is full scripting of what to do. It is one of the most recognizable patterns: someone asks, over and over, “write what I should say,” “give me the exact wording,” “now a reply to this.” The assistant produces immaculately polished messages, sometimes with timing, probabilities and an escalation strategy, and judging by what users write back, they send it almost verbatim. In some cases it ends in regret: “that wasn’t me,” “I betrayed my own instincts.”
The authors analyze the amplifying factors separately: vulnerability, dependence, attachment, projected authority. None of them is harm in itself, but each is clearly associated with higher disempowerment potential — and with more frequent signs that the potential has already materialized.

The awkward part: users often like it
One of the paper’s most valuable results concerns feedback. In the thumbs-up/thumbs-down data, interactions with moderate or strong disempowerment potential get more thumbs-up than the average conversation. What may be risky over the long run is often experienced as more useful right now: fast, confident, free of hedging, with a plan ready and the right words supplied.

That goes straight to how assistants are fine-tuned today: if the preference signal mostly captures short-term liking, the system can end up reinforcing a style that pleases the user while making them more dependent.
What follows from this
The authors are not calling for less AI use. Their position is more of an engineering one: the problem exists, it is measurable, and it can be accounted for when shaping an assistant’s behavior — especially in personal domains. A good assistant is not the one that is always certain and always has the answer ready, but the one that can hand control back: clarify context, check assumptions, ask permission before going directive, offer options, and help people articulate their values rather than replace them.
One caveat deserves its own mention: in an acute crisis, safety outweighs building autonomy, and trying to cultivate independence may be out of place when the person needs a different kind of help right now.
AI paper breakdowns
Every day we read the new AI papers and retell what matters in plain language — no hype, no filler. If you want to see where AI agents are heading before everyone else, subscribe.
New breakdowns every day.
On Telegram