One of the most unsettling images to circulate recently was an AI-made sandwich that appeared to contain noodles with no clear beginning or end. The strands seemed to merge with the filling, crawl over the bread, and continue into spaces where no noodle should exist. This kind of image is not an isolated mistake. It is a direct consequence of how diffusion models represent shape and texture.
AI-generated food has become its own genre of disgust. Restaurants, cafes, and brands are increasingly tempted to use machine-made pictures for menus, website banners, and social media posts. The results can look less like meals than like geological disasters. Shrimp resemble donuts, burgers look like broken concrete, burritos are covered with clustered holes, and ice cream takes on the texture of a pavement that has been through a freeze-and-thaw cycle. Even when an image is coherent enough to be recognized, it often feels deeply, viscerally wrong.
Diffusion builds pictures from noise
Most widely used image generators are diffusion models. During generation they start with a screen of pure visual static and then repeatedly remove noise, step by step, until an image appears. This order matters. Early operations establish broad color fields, silhouettes, and structure. Late operations add surface details such as skin, fabric, leaf veins, and shadows. Chris Russell, a professor of AI, government, and policy at the University of Oxford, explains that coarse structures are recovered first with fine texture details coming at the end.
If the model has already committed to wrong coarse geometry, texture details arrive too late to fix it. A burger could be given a boulder-like mass, then covered with realistic craters. A pastry could receive a structural shape that has no clear boundary, and the fine-grain texture layer will simply make the shape look even more alien. Russell compares this to the familiar failure of humans with six fingers instead of five. The overall skeleton was wrong, and later pixel-level refinement made it worse. In food, bizarre anatomy can lead to shrimp that look like donuts and pastries that belong nowhere near a plate.
Why noodles and holes cause trouble
Giovanbattista Califano, a behavioral scientist at the University of Naples Federico II, studies how people respond to AI-generated imagery. He points out that diffusion models are notoriously weak at generating thin, continuous, terminating structures. Noodles, strands, and tendrils are exactly the kind of geometry that trips them up. A model may start drawing something stringy and then struggle to decide where that string should end. Instead of stopping at the edge of a noodle, the artifact can bleed into a piece of bread, wrap around a piece of meat, or emerge from a sauce like a parasite.
Repeating textures cause a similar problem. Clusters of holes, bubbles, seeds, and small irregular shapes are difficult for the model to contain within sensible boundaries. The result is often a patch of holes that appears to have no natural end, spreading across an area where no such pattern should exist. That helps explain why so many AI food images are riddled with clustered cavities that trigger feelings of revulsion in people who are sensitive to trypophobia. The holes look like signs of infestation, decay, or disease.
AI has no understanding of what food is
Diffusion models also lack any meaningful understanding of the objects they are asked to depict. A model does not know that a sandwich is normally made of bread and filling, that a noodle is a kind of dough, or that a burrito is supposed to hold its contents inside a tortilla. It has learned a statistical approximation of how these things tend to appear in images, but it has not learned why they look that way or how they behave in the physical world.
Roland Meyer, a professor of digital cultures and arts at the University of Zurich, describes this as reproducing looks without proper knowledge about the world. The AI can copy the visual conventions of a food photograph, including glossy lighting and saturated colors, while having no grasp of the chemistry, structure, or cultural meaning of a meal. A model may have seen thousands of images of ice cream, for example, but it does not know that ice cream melts, that it is soft, or that it should not have the texture of cracked masonry.
Because there is no underlying physical reasoning, textures can flow across categories. Michael Cook, a senior lecturer in computer science at King’s College London, notes that a texture might look totally normal in an architectural context. Cracks in stone, irregular pits in concrete, and rough mineral surfaces are all perfectly acceptable in the right setting. But when those same textures show up on a burger or a scoop of dessert, they become deeply wrong because a human brain immediately categorizes them as inedible or contaminated.
A distorted training diet
The images used to train these systems can make the problem worse. Professional food photography is highly stylized. It often features sharp contrasts, intense colors, glossy reflections, and exaggerated shapes. Some of the food shown in those photographs is not actually edible. It may be styled with glue, shellac, artificial coloring, or other props. AI models are very good at absorbing these surface qualities, but they do not understand the professional strategies that produced them. They imitate the look of photography without understanding the intent behind the lighting, the styling, or the arrangement.
Training data scraped from the internet is also skewed by human behavior. Simon Colton, a professor of computational creativity, games, and artificial intelligence at Queen Mary University of London, points out that ordinary images of a simple red apple are probably rare compared with strange, funny, or shocking food images. People do not usually post boring pictures of an apple. They post bizarre creations, meme-worthy disasters, and highly processed visual oddities. As a result, an AI model can develop a peculiar set of associations about what food is supposed to look like.
There is also the growing problem of AI systems training on their own outputs. Many popular AI-generated videos involve bizarre food scenes, including people jumping into piles of food. When models are trained on material produced by other models, they can experience a form of model collapse. This can lead to visual degeneration, increasing sameness, and a gradual loss of the variety that makes images seem natural. The more synthetic slop appears in training sets, the harder it becomes for a new model to learn what authentic food actually looks like.
Prompts, upscaling, and hidden instructions
The way images are requested also contributes to the problem. Some prompts are extremely vague, asking only for a sandwich or a bowl of noodles. Others include language that makes sense for text, such as be precise or create a clear image, but whose meaning is very different for a diffusion model that operates on pixels rather than concepts. System-level instructions built into popular tools can also shape outputs in unpredictable ways, especially when they are written by people who do not fully understand the visual failure modes of the underlying technology.
Resolution and upscaling are another source of defects. Many image generators first create a low-resolution image and then enlarge it with a separate upscaling step. During that enlargement, the model must invent new details. If the original image contains a subtle crack, smudge, or strange edge, the upscaler can turn it into a much larger and more prominent feature. Small artifacts that might have escaped notice are therefore magnified into the kind of glaring deformities that ruin an otherwise plausible picture of food.
Evolution makes us experts at spotting bad food
Human beings are particularly good at noticing when food is wrong. Disgust likely evolved, at least in part, to protect people from parasites, pathogens, toxins, and other threats. The brain is finely tuned to treating unfamiliar or abnormal food with caution. Califano says this makes the uncanny valley for food even more visceral than the experience of seeing a robot that looks almost, but not quite, human.
AI-generated food can hit several of those evolved triggers at once. Strange noodle-like tendrils resemble worms or parasites. Clusters of holes suggest infestation. Unnatural colors, uneven textures, and glossy surfaces that look like mucus or slime signal contamination, spoilage, or chemical hazard. The brain does not require a conscious decision to reject the image. The feeling of disgust arrives automatically, before the viewer has time to think about whether the image is real or artificial.
This is why AI food looks the way it does. The failures begin inside the mechanics of diffusion, where coarse structure and fine texture are built in separate phases. They are strengthened by the absence of real world knowledge, by the strange contents of training data, and by the technical shortcuts used to upscale and refine images. At the receiving end, human beings are equipped with an ancient alarm system that reacts to anything resembling a contaminated meal. The result is a category of images that somehow look like food but feel like danger. The more AI is used to sell meals, the more clearly these limitations are exposed. A machine can imitate the surface appearance of a dish, but it does not know why food looks appetizing, what makes it safe to eat, or where a noodle should end. AI-generated food is a reflection of those blind spots, and the disgust it produces is a very human response to a very algorithmic mistake.
Source: The Verge News