Walk into a cafe, look at the board of bagel sandwiches, and something is off before you can name it. Every item is flawlessly symmetric, too smooth, missing the small irregularities real food has. Sometimes the tell is obvious, like cheese on a burrito bubbling so hard the dish reads as avant-garde sculpture rather than lunch. More often the image passes until you actually look. Alex Lyle of Reality Defender, which sells tools for detecting AI-generated material and verifying authenticity, compares these pictures to an alien attempting a pizza without understanding the principles of one. A user on X put it more locally: in New York you can throw a rock and hit a mildly disturbing AI food ad.
The cause, according to Lyle, sits in how the models were trained. Ice cream comes out in perfect spheres. Shrimp look genetically engineered to eat their own tails. The result is a new category of what he calls Lovecraftian food horrors.
The mechanism is not mysterious. Large language models and the diffusion models behind ChatGPT and image generators like Midjourney learn patterns from enormous datasets and then predict what a user wants from a prompt like "make a menu for a burger restaurant." A great deal of the output looks like a Chili's menu circa 2015, which Lyle attributes to exactly that kind of material being in the training set when the models learned what a menu looks like.
New data matters enormously to the companies building these systems, which is why the sourcing has become aggressive. Amazon, it emerged, was acquiring rare books, scanning them for training, and destroying the originals once the material was uploaded.
Datasets that large inevitably take in AI-generated material of their own. Feed a model too much of its own output and you risk model collapse. Lyle likens it to mad cow disease: keep putting a model's results back into the model and the inbreeding becomes total until the system falls apart. But he does not think menus are a collapse case. They are a convergence case, which is milder and more durable. Convergence degrades output without rendering the system useless.
Ask for a fast-food menu and the model reaches for Wendy's, Burger King, McDonald's, or another large chain. Those menus already look like each other. The generated result copies the shared aesthetic, and then reinforces it if that output finds its way back into training data. Nothing breaks. The range just narrows.
Source: techcrunch.com
Advertising food has always looked better than food. In a McDonald's shoot, a prop stylist arranges each layer of a Big Mac to make the burger as appetising as possible. AI output amplifies the same instinct. Lee Rainie, director of Elon University's center for the study of the digital future, argues that optimising datasets for pleasantness and the absence of anything offensive gradually produces homogeneity. In images as in language, the model sands off the sharp and the unusual.
The sanding is easiest to see up close. A user on X called Labtec generated a menu in ChatGPT and then edited it 100 times. With each revision the food looked less like food, and he wrote that the final image made him uncomfortable. TechCrunch, which reported the phenomenon, reproduced the experiment and got a similar result.
That is the part restaurants should care about, because it maps directly onto how they actually work. Nobody generates a menu once. They change prices, rename dishes, adjust small design details, and every pass makes the buns a little rounder and the surfaces a little smoother. The damage is done in editing, not generation, which means it accumulates in exactly the businesses least likely to notice it happening.
There is a measured reason the results repel people. Researchers at the University of Duisburg-Essen found an uncanny valley effect in AI-generated food images: nearly realistic pictures produced more disgust and anxiety than pictures that looked plainly fake. The cultural mood around AI sharpens the reaction further. Rainie says people detect the difference between a generated image and a photograph almost inexplicably, without being able to articulate what is wrong. That is why the first stories about backlash against restaurants using AI menus travelled as far as they did.
Here is what I take from this, and it is not the usual complaint about slop. Convergence is not a defect that a better model fixes. It is the predictable output of optimising for inoffensiveness, and the companies doing the optimising consider it a feature. The menus are simply the cheapest place to observe it, because food is something everyone has looked at closely and nobody needs training to evaluate. What shows up in a bagel sandwich is happening in copy, in stock imagery, in product photography and in every other output where the model has been tuned to avoid the unusual. Most of those are harder to check by eye.
The thing nobody has put a number on is whether any of this costs money. The backlash stories resonated, the research shows the reaction is real, and the aversion is well documented at the level of feeling. Nobody has shown that a diner who notices the cheese bubbling wrong stops coming back. Until someone measures that, a restaurant owner comparing a free image generator against a photographer's invoice will keep choosing the generator, and being right about the economics.
Lyle's larger point is the one that outlasts the menus. For a long time people worked from the assumption that what they saw and heard counted as proof, and court systems were built substantially around that premise, with recorded confessions and video as the gold standard of evidence. That assumption has now moved, whether or not one wants to call the move good. The menu board is where ordinary people are learning, several times a week and without being taught, that an image is not evidence of anything. They are not going to unlearn it on the way into a courtroom.