i
DATAIST
News · 2026-09-16

Study finds AI disclosure itself triggers the trust penalty

@neuronium_ai @neuronium_ai

Telling a user they are talking to an AI is the thing that makes them trust the conversation less. That is the finding at the centre of "The AI Penalty and the Disclosure Paradox: Trust, Authenticity and Knowledge Assimilation in AI-Mediated Communication," by Siavosh Sahebi, Paul Formosa and Sarah Bankins, published in March 2026 in Computers In Human Behavior: Artificial Humans. People judge communication involving AI as less trustworthy and less genuine than communication between people — but the downgrade is triggered by the disclosure itself. Undisclosed, the same exchange usually passes without the extra bias. Honesty, which may be an ethical obligation and in several US states is already a legal one, carries a measurable cost.

Cover: Study finds AI disclosure itself triggers the trust penalty

Telling a user they are talking to an AI is the thing that makes them trust the conversation less. That is the finding at the centre of "The AI Penalty and the Disclosure Paradox: Trust, Authenticity and Knowledge Assimilation in AI-Mediated Communication," by Siavosh Sahebi, Paul Formosa and Sarah Bankins, published in March 2026 in Computers In Human Behavior: Artificial Humans. People judge communication involving AI as less trustworthy and less genuine than communication between people — but the downgrade is triggered by the disclosure itself. Undisclosed, the same exchange usually passes without the extra bias. Honesty, which may be an ethical obligation and in several US states is already a legal one, carries a measurable cost.

The effect has two halves, and they work at the same time. The "human premium" is the additional trust people extend to a conversation they believe comes from a person. The "AI penalty" is the trust they withdraw when they believe a machine is answering. Research has repeatedly found both.

The scenario the authors use is mundane on purpose. You have bought something, you want to return it, and support writes back: "Your case is of the highest importance to us, and I am ready to help you process your return." If you think a company employee wrote that, your trust is probably moderate — a little guarded, because returns are often harder than they should be, but with room for cautious optimism. Now assume the message came from an AI agent, and that no conversation between two people will take place at all. Trust falls. The machine is expected to produce templated answers, less flexibility, a less natural exchange. A routine return starts to look like a potential problem.

What makes the finding awkward for anyone building these systems is that the trust tracks the belief, not the source. To measure the depth of the bias, researchers told participants the wrong thing about who was answering. Tell someone an AI will handle their return when a human employee is in fact on the other end, and they apply the AI penalty anyway: they rate a human conversation worse purely because they think it is machine-generated. Run it in reverse — a person who believes they are speaking to an employee while an AI answers — and they apply the human premium and trust the exchange more.

That produces four cases: a matched human-to-human conversation, a matched human-to-AI conversation, and two mismatches, where the participant believes they are speaking to a person but is answered by an AI, or believes they are speaking to an AI but is answered by a person.

Most people assume they would catch the switch. They picture the conversation starting and immediately sounding wrong. That assumption is dated. Machine dialogue used to be stilted, and the system gave itself away quickly: push it with an unexpected question and it would stall and admit it could not answer. In ordinary conversation those tells are disappearing.

The paradox leaves developers with three options and no clean one. Tell the truth and say the user is talking to an AI. Lie for convenience and let them believe a person is answering. Or say nothing and let the user decide for themselves who is on the other end.

The case for the truth is straightforward: a user should always know when they are dealing with an AI, and anything else is dishonest. Staying silent can be read as an ethical violation and, increasingly, an illegal act — several US states have passed laws requiring AI developers to state plainly that the counterpart is a machine, in some cases repeatedly over the course of a conversation. On this view, whatever the user does with that information is their own responsibility. If they succumb to a bias, the developer is not obliged to protect them from their own beliefs.

The case against runs the other way. A user who reacts badly to AI purely because it is AI degrades the interaction themselves; they will rate the conversation as worse even when the system performs as well as a person or better. From that position the operator should keep the right to decide whether to disclose, on the grounds that the nature of the source is immaterial if the conversation goes fine. Concealment is reframed as protecting the exchange from the user's prejudice.

This is where the argument should be read carefully. The concealment case is the only one of the three with a commercial upside, and it arrives dressed as consumer protection. The study does not endorse it — it documents a cost and names a tension — but the finding is exactly the ammunition a deployment team needs to argue that a disclosure banner is a conversion problem rather than a duty. "Our users prefer not to know" is a claim that will be made on the strength of research like this, and it should be treated as what it is: an interest, not an ethic.

The third option, silence, is weaker than it looks. The developer poses as neutral while knowing full well what sits on the other end, and material information is withheld on purpose. Omission of that kind can be close to a straight lie. It is also the worst position legally. Where a law exists, non-disclosure invites civil and criminal exposure. Where one does not, users can still sue later on the claim that they were misled, and a court or jury is likely to find "I did not know I was talking to a machine" persuasive. "We simply said nothing" is a thin defence, and "we lied for the user's benefit" is likely to land worse.

A situational rule — disclose for health, finances and other personal matters, leave the rest to the operator — sounds like the compromise. It replaces one decision with a recurring one: every case now requires a judgement about whether the truth is owed, and each of those judgements is available to be used against the developer later. Asking the user whether they want to know is no better. The question itself plants the doubt. Having been asked, the user starts guessing, and may suspect deception whatever answer comes back — a problem that did not exist before it was raised.

One finding undercuts the whole framing, and it is the one getting the least attention. The bias sometimes runs the other way: an AI premium and a human penalty, a preference for the machine and distrust of the person. The assumption that people simply dislike talking to AI is not universal, and if the direction of the bias is unstable then the penalty being measured now may be an artefact of a moment rather than a fixed fact about human psychology.

That matters, because the laws now being written assume the bias points one way and will need disclosure to keep meaning something after it stops. Thomas Aquinas, whom the question of revelation long predates, held that human salvation requires the divine disclosure of truths that surpass reason. The AI version of the problem may end up in the same place: not a question about managing user bias at all, but about whether truthfulness is contingent on whether the truth helps.