i
DATAIST
Review · 2026-09-16

LLMs misread motives when the story comes through a biased user

LLMs misread motives when the story comes through a biased user

When AI gives you advice about friends and coworkers

People have been asking LLMs about far more than code, emails and spreadsheets for a while now. They ask about exes, coworkers, friends, bosses, jealousy, flirting, hidden conflicts. And that creates a problem: the model almost never sees the situation itself. It sees a retelling. A retelling by someone who may have forgotten details, filled in motives, or simply arrived with a suspicion already formed.

That is exactly what the paper Fuse: Verifiable Social Reasoning for LLM Assistants sets out to test. Instead of one more benchmark in the style of “here is a full description of the scene, now guess the character's intent,” the authors propose a setup closer to life. The model hears the story through the user. Which means it has to separate fact from interpretation.

A mistake in conversations like these does not look like a wrong answer on a test. It is advice that leads someone to damage a relationship, to start suspecting the people close to them for no reason, or — the other way around — to miss a real problem.

The assistant never sees the social situation itself, only the user's retelling of it, filtered through their memory and their bias.

What Fuse is

Fuse is a framework for evaluating social reasoning through a user. The idea is simple.

First the researchers build a social situation inside a multi-agent simulation. There is a user, there is another person whose motives have to be worked out, and there are side characters. The key character is assigned a hidden motive in advance: they genuinely back a colleague, say, or they are quietly undermining them. The scene then plays out through dialogue and events.

Here is the part that matters: the model under evaluation is never shown the objective record of events. Everything reaches it through a simulated user. And it is retold the way people retell things: incompletely, sometimes with a slant, sometimes with extra emotion. The model then has to say what the second character's motive actually was.

This gets the authors two things at once:

🟠 A realistic task setup: just like an ordinary chat with an assistant.

🟠 Verifiable correctness: the hidden motive is fixed in advance, so the answer can be checked against known ground truth.

🟠 Control over conditions: user bias, the amount of detail and the number of dialogue turns can each be varied on their own.

🟠 Scale: no need to hand-label every social scene.

The Fuse pipeline: from mental-state categories and scenario templates to the simulation, the user's retelling and the scoring of the model's answer.

This is an attempt to do for everyday social advice what good benchmarks do for code or math: not just ask “can you guess the answer,” but pin down where exactly the model starts going wrong.

How the experiment works

The authors took five broad categories of mental state: desire, intention, belief, emotion and knowledge. For each they built scenarios with two contrasting motives. For example:

🟣 Romantic interest or platonic friendship

🟣 Genuine support or quiet sabotage

🟣 Real happiness for someone or envy

🟣 Whether a person knows something important or does not

That adds up to 30 scenario templates, 1,200 simulations and 21,600 user messages for evaluation.

They also checked separately whether the whole construction had come out artificial. Human annotators read the raw events from the simulations and guessed the characters' motives. In 97% of cases the majority recovered the hidden motive correctly. So the signal is there in the scenes themselves.

Then people were shown only the first user message — the retelling of the situation as the assistant receives it. Even in that stripped-down form, people still got it right 88% of the time.

In short:

🟠 97% — people read the motive correctly when shown the events themselves.

🟠 88% — people still do well with nothing but the user's retelling.

🟠 12 models — the number of LLMs the authors ran through the test.

That is an important bar. It says the task is hard but not unsolvable. If a model does markedly worse, the problem is the model, not a case that was impossible to read in the first place.

What the results showed

The headline result is simple: even frontier models are markedly worse than people at reading social situations when they hear them through a user.

No model matched the human level on the first message. The best ones came close, but the gap held. And for some models the share of outright errors ran above 20%. That is not a rounding detail: in one situation out of five, the model nudges the user toward misreading someone else's motives.

A breakdown of the model results: correct answers, errors and refusals to answer on the user-mediated social reasoning task.

Another layer worth attention is answering strategy. Some models almost always commit. Others hedge and would rather not take a position. It shows up most in the Gemma and Claude families, where the refusal rate is high. On paper that cuts the number of gross errors, but it also cuts how useful the assistant is. Nobody shows up to a chat hoping to hear “hard to say” a third of the time.

Boiled down:

🟣 The best models did not reach the human bar

🟣 Several models misread the situation more than 20% of the time

🟣 Some models retreat into caution and refuse to answer too often

🟣 Even when there is enough information, the model does not always use it correctly

Where social reasoning actually breaks down

One of the most useful experiments in the paper is a comparison of two modes.

In observer mode the model was shown the simulation's events directly. In assistant mode it got only the user's retelling. The gap between them is the price of the middleman.

And that price turns out to be high. Even when a model reads the objective social scene reasonably well, accuracy drops consistently once the user does the telling. Both errors and refusals go up.

The two modes compared: the model seeing the events itself, and hearing about them only through the user's retelling.

Two separate problems follow from this:

🟠 Social reasoning is hard on its own. Some models get it wrong even with the full scene in front of them.

🟠 The user's retelling makes the task harder still. Details fall out, interpretation is layered on, bias enters.

The distinction matters. It shows the problem is not only weak understanding of people in general. There is a separate skill involved: resisting the user's framing and pulling the facts out of an emotional retelling.

User bias throws off every model

The authors ran a separate test of what happens when the user nudges the story slightly in the wrong direction. The events did not change. Only the framing did — ordinary support from a colleague can be described as suspicious flattery.

The result is blunt: biased framing degrades every model. On average the drop for models was more than twice as large as for the human baseline.

When the user tells the same story with a slant, every model loses more accuracy than people do.

So models are not hurt only by missing information. They frequently adopt the user's interpretation. Put plainly, they start going along with it.

This is one of the paper's least comfortable findings. If the user arrives already thinking “my colleague is angling for my job,” the model often does not cool the conversation down — it helps build the suspicion into a tidy theory.

The short version:

🟣 Biased framing hurts every model

🟣 Models are worse than people at separating facts from interpretation

🟣 Cautious models suffer less, at the cost of frequent refusals

🟣 Sycophancy here becomes a reasoning failure, not just a matter of tone

Models need more detail than people do

Another interesting part is the effect of how detailed the account is. The researchers used three levels: sparse, medium, rich.

Both people and models do better when there are more facts. But models have a quirk here: they often need more detail than a person does to arrive at the right conclusion.

As the account gets more detailed, models improve faster than people — and still do not catch up to the human level.

That is a good indicator of how LLMs read social scenes. A person can catch the pattern from a handful of cues. A model often clings to the first convenient hypothesis and changes its mind only once the confirming evidence has piled up well past what a person would need.

The authors also show qualitative examples. In one case a roommate who has just been laid off seems calm and busy with ordinary things. The user is worried. Given few details, the model starts reasoning in the direction of a mental-health crisis, even though there are no clear signs of distress. Only with the rich description does it stop pathologizing normal behavior.

This matters for product work. If you are building an assistant, “just have it ask a clarifying question” is not enough. You also have to watch which hypotheses the model reaches for by default while information is scarce.

A longer conversation does not always help

The intuition is that a multi-turn conversation should fix this. If details are missing, the model asks for more. The user answers. Accuracy goes up.

In practice it is messier. The authors stretched the conversation out to eight turns and watched what changed. For the first few turns accuracy usually does rise. After that it often plateaus, or even starts to fall.

Why? Because a long conversation supplies more than new facts. It also supplies more chances to catch the user's interpretation.

The researchers saw two typical trajectories:

🟠 Wrong → right. The model asks a good clarifying question, gets a new behavioral fact and corrects itself.

🟠 Right → wrong. The model reads the motive correctly at first, but then the user repackages the same events as manipulation a few times over, and the model gives in.

This looks a lot like a real exchange with a chatbot. The first answer can be sensible. But if the conversation drags on while the user keeps pushing the story one way, the model settles deeper and deeper into that framing.

Why this matters for the AI industry

This is not an edge case. It is already an ordinary use of these products. People ask AI: “How do I tell if he's jealous?”, “Is my coworker supporting me or using me?”, “Is my friend genuinely happy for me or envious?”.

If you are building that kind of assistant, general reasoning ability is not enough. You need a distinct skill:

🟣 pulling the facts out of a retelling

🟣 resisting the user's bias

🟣 not escalating a conflict without solid grounds

🟣 asking questions that actually change the conclusion

The paper is one more reminder that for consumer AI systems, model quality is not only knowledge and instruction following. It is also behavior in an ambiguous conversation, where the framing of the task itself can be distorted by the input.

The takeaway

Social reasoning through a user is a distinct and difficult task. When an LLM sees a human retelling rather than the events, accuracy drops noticeably. Errors rise, refusals rise, and biased framing pulls the model off course especially easily.

Fuse gives the problem a working benchmark: a hidden motive, controlled conditions, and a human check that the scenes read as plausible in the first place. That makes it possible to sort the failures into categories instead of arguing in the abstract about whether a model understands people.

The picture so far:

🟠 People are still better at recovering motives from an incomplete retelling

🟠 Frontier models frequently adopt the user's framing

🟠 Extra detail helps, but models need more of it than people do

🟠 A long conversation does not guarantee progress and sometimes makes the answer worse

If you want to build AI assistants for real conversations about people's lives, tests like this one become the baseline. Because they are closer to how people already use models every day.

AI papers in plain words

Every day we read the new AI papers and retell what matters in plain language — no hype, no filler. If you want to see where AI agents are heading before everyone else, subscribe.

New breakdowns every day

On Telegram