HomeThe New IntelligenceWhen AI Passes the Test But Fails the Question

When AI Passes the Test But Fails the Question

Harvard's Keyon Vafa trained AI on Manhattan streets and found an invented city. New research reveals why AI world models break down under pressure.

Share
The New Intelligence · Explore this series
July 19, 2025
Key Takeaways
  • An AI trained on Manhattan routes was 99% accurate but held no coherent map of the city.
  • Blocking just 1% of roads collapsed the AI's routing performance entirely.
  • AI can explain rules correctly yet fail to apply them — researchers call this potemkin understanding.

Keyon Vafa had trained an AI on millions of Manhattan street directions. The model's AI world model looked impressive from the outside. Ask it how to get from the Lower East Side to Columbus Circle and it would tell you, accurately, 99% of the time.

Then Vafa asked a different question: what map of Manhattan does this model actually hold?

The answer, when his team at Harvard's Data Science Initiative reconstructed it, was unsettling. The AI had invented roads. It leaped across Central Park. It cut diagonally through a grid famous for its right angles.

Turn left, and you get one map of Manhattan. Turn right, and you get an entirely different one. The two maps share no coherent geography.

Key figure

99%

Accuracy giving Manhattan street directions, from an AI with no coherent map of the city

When AI World Models Break Down

Vafa, a postdoctoral fellow who completed his PhD at Columbia before joining Harvard, devotes much of his work to a question that sounds simple: does AI understand, or does it merely predict well? His criteria come down to whether a model carries a stable world model, a flexible internal framework that holds together under new conditions.

Sometimes, he notes, the answer looks promising. Ask a large language model to balance a marble on an inflatable beach ball on a stove pot on grass, and it gives the correct stacking order, even though that precise question likely never appeared in training. That suggests something like physical intuition.

But the Manhattan experiment told a different story.

Rather than building a model of the city, the AI had memorized a vast library of routing rules and stitched them together fresh for each query. Accurate, but without any underlying geography.

A related finding reinforced the point: when Vafa's team blocked just 1% of the virtual road network, the AI's routing performance collapsed. A person would navigate the detour without thinking.

What is a world model?

A world model is an internal representation of how reality is structured, not just a record of correct answers. A person who knows Manhattan has a world model: they can navigate detours, estimate distances, and reason about new routes. An AI without one is more like an oracle that memorizes outcomes without understanding the system that produces them.

Passing the Exam, Missing the Point

In 2025, Vafa and colleagues Marina Mancoridis, Bec Weeks, and Sendhil Mullainathan extended this inquiry into language itself, publishing a study at the International Conference on Machine Learning that introduced a term now circulating widely in AI research: potemkin understanding.

The term is deliberate. Rather than saying models "misunderstand," which implies something human, the researchers named the problem after a Russian stratagem: an elaborate facade built to impress without substance.

We as a field are making steps trying to understand, what would it even mean for something to understand? There’s definitely no consensus.

Keyon Vafa, postdoctoral fellow at the Harvard Data Science Initiative

Their test was revealing. Ask GPT-4o to explain the ABAB rhyming scheme, and it does so correctly and clearly: first and third lines rhyme, second and fourth rhyme. Ask it to write a four-line poem using that scheme, and the rhymes come out wrong. The model can describe the concept without being able to apply it.

Vafa's team found potemkin understanding to be widespread across every major model they tested, across literary techniques, game theory, and psychological biases. The pattern held: explain the rule, fail the application.

This matters beyond academic curiosity. AI benchmarks, including the AP exams and math competitions used to demonstrate AI progress, are tests designed for humans on the assumption that passing them indicates the same kind of AI understanding that passing them indicates in a person. Vafa's work suggests that assumption is wrong. A model can clear a benchmark without the underlying comprehension the benchmark was designed to measure.

What AI Understanding Actually Requires

The late Harvard philosopher Hilary Putnam posed a version of this puzzle decades ago. Imagine an ant crawling in sand, tracing a path that happens to resemble Winston Churchill's profile. Would you say the ant has drawn Churchill's portrait?

Putnam's answer was no: the ant would need to know about Churchill, about likenesses, about what a portrait is meant to do.

Vafa's research gives Putnam's thought experiment empirical weight. The AI routing directions around Manhattan is, in an important sense, that ant. Its outputs are accurate. Its comprehension is absent.

More On AI Understanding

AI Bias: How Language Models Amplify What They Copy

Researchers studying AI bias thought they were building digital twins. What they created instead were caricatures.

The stakes become concrete when AI is deployed in high-stakes settings: medical diagnosis, legal reasoning, infrastructure management. Cheryl Chen, a senior lecturer in philosophy at Harvard, points out that when a person says something, we assume they mean it in the full experiential sense. They have encountered rain, felt frustration, formed associations. With AI, that assumption is precisely what Vafa's work puts in doubt.

Vafa is careful not to suggest the situation is hopeless. Current models can still be useful even without coherent world models, he notes, and knowing how to measure the problem is itself progress.

The question now shifts from whether current AI understands to how understanding of this kind might be built, detected, and measured at scale.

Vafa argues that no serious path toward more capable AI can avoid that question.


Sources

Fact Check: Claim-by-Claim Verification Verified

All major claims are accurately supported by peer-reviewed research, author credentials, and institutional affiliations are correct, and the article appropriately represents the findings without misrepresentation.

1 Verified
Keyon Vafa is correctly identified as a postdoctoral fellow at Harvard's Data Science Initiative who completed his PhD at Columbia in 2023
2 Verified
The Manhattan navigation experiment is accurately described: AI achieved 99% accuracy on directions while lacking a coherent internal map, with failures when just 1% of the virtual road network was blocked
3 Verified
The "potemkin understanding" research is correctly attributed to Vafa, Marina Mancoridis, Bec Weeks, and Sendhil Mullainathan, published at ICML 2025
4 Verified
The GPT-4o rhyming scheme example correctly represents the core finding: models can describe concepts accurately while failing to apply them
5 Verified
Hilary Putnam's ant-drawing-Churchill thought experiment is accurately characterized as a philosophical precedent for understanding representation and intentionality
6 Verified
Cheryl Chen is correctly identified as a senior lecturer in philosophy at Harvard with relevant expertise in epistemology and philosophy of mind
7 Verified
The quote from Vafa about lack of consensus on what AI understanding means is verified in the Harvard Gazette source

Commentary

  • The article accurately represents research findings but uses accessible language ("invented roads," "leaped across Central Park") that are simplified descriptions of the formal technical findings. This is appropriate for popular science journalism and does not constitute error.
  • The article correctly notes that potemkin understanding was observed "across every major model" tested, reflecting the scope of Vafa's research as described in the ICML abstract
  • The article's framing that "passing the exam but failing the question" reflects the core insight of the potemkin understanding paper, which contrasts benchmark success with failure on conceptually equivalent but differently-framed tasks

Sources used for verification

Academic/Peer-reviewed:

Other reliable sources:

Share
Related Articles
AI Consciousness Is Unlikely, Says Neuroscientist Anil Seth

Neuroscientist Anil Seth argues AI consciousness is unlikely without biology. His TED talk lands amid a widening debate over conscious AI, not intuition.

AI In Science Connects the Dots, But Only In Fields That Are Fragmented

An analysis of 80 million papers shows AI boosts originality where knowledge is scattered and connections are weak, but contributes little novelty in structured science.

"Keep Humanity Safe From AI," Urges Pope Leo XIV

Pope Leo XIV's first encyclical reaches the same verdict on AI as the labs building it, then parts ways over the meaning of human limits.

AI Solves Erdős Math Problem: What's Next for AI in Mathematics?

An AI solved an 80-year-old Erdős math problem by walking a path mathematicians had collectively avoided.