i
DATAIST
Review · 2026-07-03

Readers can't spot AI literary translation but still prefer the human one

Readers can't spot AI literary translation but still prefer the human one

Good enough to read, still not the one readers pick

One question has hung over machine translation of literature for years: if a model can carry the meaning across more or less accurately, does that mean it can translate a novel as a novel — with voice, rhythm, atmosphere, and that feeling when a text carries you along?

A new paper that puts its conclusion right in the title answers honestly: AI translations are already fine, but readers still choose the human version more often. That is probably the most interesting result. Not because the machines lost. Because the gap turned out to be far narrower than many would like, and — more importantly — because of the specific reasons humans are still ahead.

The researchers did not reach for the usual automatic metrics. Instead they sat 15 avid readers down with real excerpts from contemporary novels — French, Polish and Japanese, translated into English. First the readers went through long stretches the way they would read any book; then they went back over the translations up close, passage by passage. It is a rare case of the question of translation quality being put not to models and not to annotators, but to readers themselves.

What the study actually tested

The authors assembled 15 book excerpts of roughly 8,000 words each. Each one came with a professional published English translation and a machine translation produced not by a single prompt but by a multi-step pipeline built on LLMs and coding agents. So this is not a weak baseline — it is closer to the best attempt at AI literary translation available today.

Participants read one version of an excerpt end to end, then the other. After each, they rated how fluent the text was, whether it was publishable, whether it helped them get absorbed, and whether they wanted to keep reading. Then they compared the two versions directly. A day later they returned to the same texts in slow-reading mode: aligned chunks of about 300 words, marking what worked and what didn't, and picking the better version.

The experimental design: ordinary reading of a long excerpt first, then a comparison of the two versions, and a day later a close side-by-side reading of short passages.

Machine translation is usually judged on short sentences or paragraphs. Literature doesn't work that way. A novel lives not only in the accuracy of its words but in how easily you enter a scene, hear a character's intonation, and get through three lines without stumbling. That is exactly the difference the authors were trying to capture: between "reads okay overall" and "this is genuinely good literary translation."

Why this matters

Because machine translation of fiction is no longer a lab experiment. Publishers are testing it on commercial titles. Platforms are offering authors a fast route into new markets. Which means that very soon — and in places already — a reader can buy a novel without knowing how deeply a human was involved in translating it.

The trouble is that the standard metrics are little help here. They pick up adequacy of meaning and general fluency reasonably well. But they barely see whether a text breathes, whether it holds its pace, whether the narrative voice falls apart between paragraphs. And this paper shows that if you look only at automatic scores, you will come away far too optimistic about the quality of AI translation.

How the machine translation was built

Notably, the researchers were not comparing humans against a naive take-a-model-and-ask-it-to-translate setup. They tested several pipeline configurations first and kept the best one.

The final design went like this: split the text into chunks, attach style guidance, then have one agent translate the chunks while others checked the output for both accuracy and literary quality. Anything flagged went back for revision. The assembled draft was then checked again at the level of the whole excerpt — for consistency, coherence and retention of voice.

The AI literary translation pipeline: chunking, translation, local revision and a final pass over the whole text.

So this is not a raw machine output but a fairly expensive, carefully designed process. That is what makes the results telling: even with all of that, the human translation stays ahead.

The headline result: AI is acceptable, humans are still better

Across whole long excerpts, the human translation wins, but not by a landslide: readers preferred it in 19 of 30 cases. In slow reading the gap widened considerably: of 772 short-passage comparisons, 522 went to the human translation and only 250 to the machine.

How reader preferences break down: after reading long excerpts and after slow reading of short passages, the human translation wins more often.

That contrast says a lot. Read a long stretch with nothing to compare it against and the machine translation often seems perfectly fine. It isn't necessarily irritating. It can be smooth, clear, at moments even gripping. Put the two versions side by side, though, and you can see where the human worked deeper and with a finer hand.

The authors also asked what specifically was better in each version. Human translations more often scored high on fluency, clarity and immersion. Readers said more often that such a text was easier to read, that it didn't make them reread a sentence, that it kept them inside the scene and sounded more natural in English.

None of which made the machine translation a failure. It won roughly a third of the passages. Sometimes on a well-chosen word. Sometimes because it sounded livelier or hit the register of a line more precisely. The AI can already make good individual choices. What it lacks is consistency.

Where the human actually wins

The most interesting part of the paper is not the raw percentages but why readers chose one version over the other.

The human translation usually won because the text flows more easily. Not in the sense of fewer errors, but in the sense of the reading experience: you don't stumble, you don't get thrown out of the scene, you aren't untangling odd word order. Readers kept pointing to clarity, naturalness, rhythm, and how easy it was to follow the action and the dialogue.

The machine translation, by contrast, broke down locally again and again: a phrase too literal, a word with the wrong shade of meaning, an overloaded sentence, a sudden roughness in dialogue. And the core problem was not that it was always bad but that it was uneven. One paragraph is excellent; two paragraphs later an awkward turn of phrase spoils the effect.

The machine translation has noticeably more passages with a high density of spots readers marked as weak.

The authors measured that unevenness through the readers' markup of good and bad stretches. Weak spots in the machine translation turned out to cluster far more tightly. The problem is not only average quality; it is that within a single excerpt the AI swings harder between a hit and a miss.

For fiction that is critical. In a news item you might forgive a couple of strange sentences. In a novel, one or two false notes at the wrong moment can wreck the atmosphere of a scene or blur a character.

AI is hard to spot, and that may be the most unsettling finding

Probably the most striking result: readers could not reliably tell machine translation from human translation. Even after comparing the two versions directly, they identified the machine version correctly in just 17 of 30 cases — barely above chance.

Accuracy at identifying the machine translation is close to chance, even after comparing the two versions of a text.

More than that, people almost always assumed the version they preferred was the human one. That is a powerful psychological effect. We don't simply judge a text — we build a story around it about authorship, intent and quality.

The study turns up another detail that is both funny and important: readers were regularly thrown off by false tells. One saw em dashes and decided it was the machine. Another figured an AI wouldn't swear like that. A third read an unfamiliar style as machine-made. Folk theories about how AI is "supposed to sound" work badly.

This goes beyond translation. There is a lot of overconfidence around AI text in general: people are certain they can spot machine writing, and in practice they often get it wrong.

Automatic metrics get it backwards

Here is where the paper gets uncomfortable for anyone who likes simple numbers. The researchers compared reader preferences against automatic quality metrics, including LLM-as-a-judge approaches. The metrics systematically lean toward the machine translation — that is, they run against what human reading says.

If you are building a product for a publisher and looking only at automatic scores, you might conclude the machine version is already close to better than the human one. Live readers say the opposite: yes, it is readable, but the human still delivers an easier, clearer, more absorbing experience.

The reason is clear enough. Metrics see closeness of meaning, surface fluency, sometimes consistency. They are poor at catching what matters most in literature: rhythm, voice, the sense of a sentence moving naturally, the density of weak spots, and the way one local rough patch breaks the whole effect.

One caveat: this is mostly a story about translating into English

The authors are upfront about the limits. The main experiment covers translations from French, Polish and Japanese into English. And English is the strongest language for LLMs today. Even in that best case, the human still wins.

In a smaller additional analysis of translations into other languages — French, Polish, Spanish and Japanese — the human advantage was larger still. There the AI was caught out more often and rejected more often.

The implication is straightforward: if the gap holds even in English, it may be wider in other language directions.

What it comes down to

This is a sober piece of work. It tips into neither techno-optimism nor alarmism.

The conclusion reads like this: machine translation of literature has reached the level of "you can read this". Sometimes it even wins. Sometimes it is hard to identify. Sometimes it looks surprisingly strong. But if the question is not whether the meaning comes through, but which text a reader wants to keep going with as a novel, the advantage stays with the human.

And the human wins not through magic but through the things that matter most in literature: flow, clarity, natural word choice, holding the voice, and steadiness across the length of a text.

For the industry, that is the practical takeaway. AI is already perfectly usable as a drafting tool, an assistant, an accelerator. But the idea of simply shipping novels in machine translation because readers won't notice is still too sure of itself. Readers may not always work out where the machine is. But when a good human translation sits next to it, they still pick it more often.

Which may be the best compliment the translator's profession could ask for.

AI papers in plain words

Every day we read the new AI papers and retell what matters in plain language — no hype, no filler. If you want to see where AI agents are heading before everyone else, subscribe.

New breakdowns every day

On Telegram