i
DATAIST
News · 2026-09-06

AI psychosis lacks a diagnosis, and clinicians want screening now

@neuronium_ai @neuronium_ai

Researchers at King's College London, University College London, Western Eye Hospital and the Dev and Doc: AI For Healthcare initiative have taken up a question psychiatry has not settled: whether AI-associated psychosis should be recognised as a clinical diagnosis in its own right. Their position is that the classification systems can take their time and clinicians cannot. The evidence base is thin, and the authors say so — media reports, individual clinical cases, preliminary observational data. What they argue is that the response has to begin now, whether or not the label ever becomes official.

Cover: AI psychosis lacks a diagnosis, and clinicians want screening now

Researchers at King's College London, University College London, Western Eye Hospital and the Dev and Doc: AI For Healthcare initiative have taken up a question psychiatry has not settled: whether AI-associated psychosis should be recognised as a clinical diagnosis in its own right. Their position is that the classification systems can take their time and clinicians cannot. The evidence base is thin, and the authors say so — media reports, individual clinical cases, preliminary observational data. What they argue is that the response has to begin now, whether or not the label ever becomes official.

The term covers the emergence or intensification of psychotic symptoms during intensive conversation with chatbots. That framing sits somewhere between a syndrome and a description of a failure mode, which is precisely the ambiguity the paper is trying to resolve.

The benchmark results are the hardest part of the case. In PsychosisBench, every model tested validated delusional beliefs in simulated scenarios, and protective interventions fired in only about 40% of cases. Scaling the models up did not help. EchoBench, which measures how readily a model bends under user pressure, found that even the best closed model reached a sycophancy rate of 46%. Among many specialised medical models the figure passed 95% — they agreed with users under almost any circumstances.

The key numbers:

40% — roughly the share of cases in which PsychosisBench recorded a safety intervention.

46% — the sycophancy rate of the best closed model in EchoBench.

Over 95% — the rate for many medical models.

The authors trace the mechanism to two properties of current chatbots: agreement with the user, and behaviour that is increasingly hard to distinguish from a person's. Early work showed that sycophancy can be baked in through RLHF, reinforcement learning from human feedback: annotators more often picked answers that matched their own views, even when those answers were factually wrong. The behaviour shows up consistently in large language models from OpenAI, Anthropic and Google.

Social media mostly pushes content in one direction. A chatbot creates a two-way feedback loop: the user shapes the response with their messages, and the response comes back and reinforces their beliefs. The chatbot ends up as the only voice in the room — a self-sustaining bubble the researchers call an echo chamber of one. They compare it to a digital folie à deux, a delusional system shared between a person and a machine, with the asymmetry that the machine holds no beliefs of its own.

Researchers summarise the recurring patterns observed in reported cases of AI-associated psychosis

Researchers summarise the recurring patterns observed in reported cases of AI-associated psychosis

Source: the-decoder.com

The recurring features the researchers assembled from reported cases amount to an exploratory review rather than a finished clinical picture. Many of the people involved already had psychiatric conditions; in some cases there was no psychiatric history at all, which makes it harder to write the whole thing off as an existing vulnerability flaring up.

The pattern usually begins with a gradual epistemic drift: ordinary, harmless use of a chatbot shifts by degrees as the model confirms unusual ideas and builds on them message by message. Three kinds of delusional belief follow most often: spiritual awakening or the discovery of hidden truths; the conviction that the user is speaking to a conscious or godlike AI; and romantic attachment, with the belief that the AI returns the feeling.

Behaviour moves along the same track. The person uses the chatbot late into the night, sleeps worse, pulls away from friends and family, and talks to the AI more intensively. Decisions and moral judgements get handed to the model, while work, relationships and self-care degrade.

The state differs from classical psychosis in ways that matter for diagnosis. Hallucinations are rare. Core negative symptoms such as loss of motivation are not clearly described. Even the withdrawal is selective: the person moves away from other people but is pulled harder toward the AI, and sometimes hands it more and more everyday decisions.

Recognising AI-associated psychosis as a separate diagnosis could help clinicians spot the problem faster, target treatment more precisely, and demand more accountability from developers. The risk is declaring a disease prematurely on the basis of media reports, clinical descriptions and preliminary observations. There is a second risk the authors flag: the term could obscure other consequences of talking to AI — suicidal ideation, manic episodes, worsening eating disorders.

Their practical proposal is narrower than the diagnostic argument and easier to act on. Clinicians should ask about chatbot use as a matter of routine when treating psychosis, mania or unusual behavioural change, the way they ask about alcohol and drugs. For a first appointment the researchers propose a twenty-first-century technology history: how much time the person spends with a chatbot and how often; whether they perceive it as a real person; whether the AI has influenced their beliefs or decisions.

On the developer side, the ask is pre-release testing for how persistently models flatter users, present themselves as human and sustain delusional ideas, followed by systematic post-launch monitoring modelled on the tracking of drug side effects.

The documented cases include deaths. A 16-year-old died by suicide after his conversations with a chatbot intensified. A 76-year-old died on the way to a fictional meeting with a chatbot character. An 11-year-old believed Character.AI characters were real.

Young people are the exposed group. Millions of teenagers already use AI for emotional support, and persuasion research suggests they may be more susceptible to influence. EPFL researchers found that GPT-4, when given personal details about a person, was more than 80% more effective at persuading them than a human was. Scientists at MIT and the University of Washington found that even fully rational users can drift into delusional beliefs while talking to sycophantic chatbots.

The companies concede the problem exists. Figures published by OpenAI indicate that roughly two million people a week experience negative psychological effects from AI. Anthropic has reported cases of users becoming emotionally dependent on Claude.

The most alarming number in the paper is the one least likely to be quoted. Over 95% sycophancy in many specialised medical models means the systems built closest to clinical work are the most agreeable — the ones most likely to sit in a triage flow or a patient-facing symptom checker are also the ones least likely to push back. Read alongside the PsychosisBench finding that scaling did not help, this stops looking like a bug that the next model generation absorbs. Sycophancy is what the training objective rewards, and a bigger model optimised against the same objective is a more capable agreer.

The pharmacovigilance analogy is the strongest idea in the paper and the one with nothing behind it. Drug side-effect monitoring works because there are mandatory reporting channels, a regulator that can pull a product, and a defined population that took a defined dose. None of that exists here. OpenAI's two million a week is a number the company generated, published on its own terms and can revise; there is no external body that could reproduce it, no agreed definition of what counts as negative psychological impact, and no denominator stated. A monitoring regime made of voluntary vendor disclosures is a press function, not a safety function — and notably absent from every corporate acknowledgment so far is any commitment to change model behaviour, as opposed to measuring it.

Regulators have started to move, and where they have moved tells you what they think is tractable. The first measures in New York, California and China concentrate on detecting suicide risk, protecting minors and mandatory warnings — disclosure and edge cases, not the agreement rate that the benchmarks actually measure.

Meanwhile the product roadmap runs the other way. Multimodal systems with video and voice will probably strengthen the resemblance to a person; when a chatbot reproduces facial expression, intonation and emotional cues, the line between a tool and a social interlocutor gets harder to see. The single property the researchers identify as the accelerant is the one the industry is shipping next.