HomeThe Science of ThoughtThe One Thing AI Image Generators Can't Do

The One Thing AI Image Generators Can't Do

They generate endlessly but can't judge what's interesting. New research shows why your taste still matters.

Image generator spouting an endless amount of images flying away in the distance.AI and computer scienceWhen AI image generators are tested in feedback loops, they converge on the equivalent of visual elevator music. (Science Reader)
When AI image generators are tested in feedback loops, they converge on the equivalent of visual elevator music. (Science Reader)
Share
The Science of Thought · Explore this series
January 9, 2026
Key Takeaways
  • AI image generators left to loop autonomously converge to just 12 generic visual motifs.
  • Convergence happened by iteration 100, regardless of starting prompt or temperature.
  • The language model describing images drove convergence more than the image generator did.

Arend Hintze expected his experiment to be straightforward. Let an AI image generator and an AI image describer play visual telephone, looping outputs back as inputs. The images should stay consistent with their starting prompts.

"I mean, how hard is it to consistently generate an image of a mountain with a village on it?" Hintze asked in a press interview.

Harder than Hintze, who studies intelligence through evolution, thought. After 700 trials, every trajectory collapsed into the same dozen destinations: lighthouses in storms, Gothic cathedrals, ornate interiors with red velvet. The researchers call it "visual elevator music."

The phrase captures something users of these tools likely recognize. A certain sameness. A particular aesthetic that's hard to articulate but easy to spot.

Now there's research explaining why.

Why can't ai image generators judge creative output?

AI image generators optimize toward the most probable outputs from their training data, not the most interesting ones. They receive reinforcement without critique - no mechanism asks whether what worked before should work again.

How AI Image Generators Lose Diversity

Hintze, a professor at Sweden's Dalarna University, designed the loop with collaborators Frida Proschinger Astrom and Jory Schossau of Michigan State.

They fed diverse prompts into Stable Diffusion XL. LLaVA described each output in 50 words. That description became the next prompt.

By iteration 100, the diversity was gone.

Using these specific models, the team identified exactly 12 visual motifs through clustering analysis. Sports imagery. Formal offices. Maritime lighthouses. Neon-lit streets. Cathedral interiors. Pompous rooms. Industrial scenes. Rustic spaces. Domestic settings. Palatial architecture. Pastoral villages. Dramatic landscapes.

"The kind of images you find in IKEA frames," Hintze wrote, "bought for the frame, while the picture gets thrown away."

Key figure

700 → 100 → 12

In 700 trials, 100 iterations lead to the same 12 visual motifs.

In their experiments, starting prompt didn't matter. Temperature settings didn't matter. The AI systems found their way to the same generic endpoints regardless.

The AI systems found their way to the same generic endpoints regardless.

AI Has Half of Creativity

Related reading

The Paper Mill Problem: Science's Fraud Industry Is Growing Faster Than Peer Review Can Handle

Paper mills, predatory journals, and AI-generated content are overwhelming peer review systems. Global retractions exceeded 10,000 in 2023, with Hindawi alone withdrawing over 8,000 articles.

The statistical analysis revealed something telling. Which language model described the images significantly affected how fast convergence happened. Which image generator created them barely mattered at all.

The bottleneck, it seems, sits in interpretation rather than generation.

"Creativity, I think, is two things," Hintze explained. "It's generating something novel, and then it's using a filter to decide: this is interesting, this is beautiful, this is stimulating, this is exciting. Right now, AI is really good at the first part, and they're really bad at the second part."

This might resonate with users who've noticed uniformity in AI-generated images - though the study tested research models, not commercial tools.

Why Human Curation Isn't Optional

The finding carries weight beyond academic interest. As AI-generated content floods the web and gets scraped into training datasets for newer models, convergence could compound.

Caterina Moruzzi, a philosopher at Edinburgh College of Art, observed in a Science.org story that the collapse happens because AI systems receive "reinforcement without critique." They optimize toward what worked before. No one asks whether it should work again.

Christian Guckelsberger researches AI creativity at Finland's Aalto University. In the same story, he hoped the limitation won't be dismissed as a mere engineering problem. Whether the constraint is structural or solvable remains an open question.

AI is really good at generating something novel. They're really bad at deciding if it's interesting.

Arend Hintze, Dalarna University

Human cultural transmission shows similar convergence patterns. But humans generate countercultures. We break rules deliberately. We value novelty for its own sake. These corrective mechanisms weren't present in the autonomous AI loops - their end result was just AI slop.

The Missing Filter

The implications land directly in creative workflows. If you use AI image generators, this research names a dynamic worth understanding: the tools produce things remarkably well, yet remain indifferent to whether those things are interesting.

That indifference is precisely what human judgment fills.

Hintze's team titled their findings "sobering for computational creativity." The word choice feels apt. Not damning, not dismissive, just clear-eyed about what autonomous systems can and cannot do.

The paper, published in December 2025, awaits independent replication. But the generation half of creativity scales beautifully.

The judgment half still needs us.


Sources

Fact Check: Claim-by-Claim Verification Verified

1 Supported
AI image-text loops converge to just 12 generic visual motifs
The peer-reviewed paper in Patterns (Cell Press) documents this finding across 700 trajectories using k-means clustering analysis.
2 Supported
Convergence happens by approximately iteration 100
Quantified in the original paper and confirmed in press coverage.
3 Supported
Language model effect is significant; image generator effect is not
ANOVA analysis in paper shows language model effect p < 10⁻¹⁹, image generator effect p=0.659.

Limits and uncertainties

This is a single peer-reviewed study published December 2025, not yet externally replicated. Results tested research models (SDXL, LLaVA), not commercial tools like Midjourney or DALL-E. The "visual elevator music" characterization is the authors' interpretation.

Bottom line

The core findings are well-supported by peer-reviewed research with robust methodology. Claims are appropriately hedged in the article to reflect single-study status.

Share
Related Articles
Why We Can Never Prove That Someone Else is Conscious

'Rival' scientists use category theory to show that while 'shapes' of experiences might be matched across minds, we can never observe the feeling itself.

AI Consciousness Is Unlikely, Says Neuroscientist Anil Seth

Neuroscientist Anil Seth argues AI consciousness is unlikely without biology. His TED talk lands amid a widening debate over conscious AI, not intuition.

AI In Science Connects the Dots, But Only In Fields That Are Fragmented

An analysis of 80 million papers shows AI boosts originality where knowledge is scattered and connections are weak, but contributes little novelty in structured science.

"Keep Humanity Safe From AI," Urges Pope Leo XIV

Pope Leo XIV's first encyclical reaches the same verdict on AI as the labs building it, then parts ways over the meaning of human limits.