- An AI trained on Manhattan routes was 99% accurate but held no coherent map of the city.
- Blocking just 1% of roads collapsed the AI's routing performance entirely.
- AI can explain rules correctly yet fail to apply them — researchers call this potemkin understanding.
Keyon Vafa had trained an AI on millions of Manhattan street directions. The model's AI world model looked impressive from the outside. Ask it how to get from the Lower East Side to Columbus Circle and it would tell you, accurately, 99% of the time.
Then Vafa asked a different question: what map of Manhattan does this model actually hold?
The answer, when his team at Harvard's Data Science Initiative reconstructed it, was unsettling. The AI had invented roads. It leaped across Central Park. It cut diagonally through a grid famous for its right angles.
Turn left, and you get one map of Manhattan. Turn right, and you get an entirely different one. The two maps share no coherent geography.
Key figure
99%
Accuracy giving Manhattan street directions, from an AI with no coherent map of the city
When AI World Models Break Down
Vafa, a postdoctoral fellow who completed his PhD at Columbia before joining Harvard, devotes much of his work to a question that sounds simple: does AI understand, or does it merely predict well? His criteria come down to whether a model carries a stable world model, a flexible internal framework that holds together under new conditions.
Sometimes, he notes, the answer looks promising. Ask a large language model to balance a marble on an inflatable beach ball on a stove pot on grass, and it gives the correct stacking order, even though that precise question likely never appeared in training. That suggests something like physical intuition.
But the Manhattan experiment told a different story.
Rather than building a model of the city, the AI had memorized a vast library of routing rules and stitched them together fresh for each query. Accurate, but without any underlying geography.
A related finding reinforced the point: when Vafa's team blocked just 1% of the virtual road network, the AI's routing performance collapsed. A person would navigate the detour without thinking.
What is a world model?
A world model is an internal representation of how reality is structured, not just a record of correct answers. A person who knows Manhattan has a world model: they can navigate detours, estimate distances, and reason about new routes. An AI without one is more like an oracle that memorizes outcomes without understanding the system that produces them.
Passing the Exam, Missing the Point
In 2025, Vafa and colleagues Marina Mancoridis, Bec Weeks, and Sendhil Mullainathan extended this inquiry into language itself, publishing a study at the International Conference on Machine Learning that introduced a term now circulating widely in AI research: potemkin understanding.
The term is deliberate. Rather than saying models "misunderstand," which implies something human, the researchers named the problem after a Russian stratagem: an elaborate facade built to impress without substance.
We as a field are making steps trying to understand, what would it even mean for something to understand? There’s definitely no consensus.
Keyon Vafa, postdoctoral fellow at the Harvard Data Science Initiative
Their test was revealing. Ask GPT-4o to explain the ABAB rhyming scheme, and it does so correctly and clearly: first and third lines rhyme, second and fourth rhyme. Ask it to write a four-line poem using that scheme, and the rhymes come out wrong. The model can describe the concept without being able to apply it.
Vafa's team found potemkin understanding to be widespread across every major model they tested, across literary techniques, game theory, and psychological biases. The pattern held: explain the rule, fail the application.
This matters beyond academic curiosity. AI benchmarks, including the AP exams and math competitions used to demonstrate AI progress, are tests designed for humans on the assumption that passing them indicates the same kind of AI understanding that passing them indicates in a person. Vafa's work suggests that assumption is wrong. A model can clear a benchmark without the underlying comprehension the benchmark was designed to measure.
What AI Understanding Actually Requires
The late Harvard philosopher Hilary Putnam posed a version of this puzzle decades ago. Imagine an ant crawling in sand, tracing a path that happens to resemble Winston Churchill's profile. Would you say the ant has drawn Churchill's portrait?
Putnam's answer was no: the ant would need to know about Churchill, about likenesses, about what a portrait is meant to do.
Vafa's research gives Putnam's thought experiment empirical weight. The AI routing directions around Manhattan is, in an important sense, that ant. Its outputs are accurate. Its comprehension is absent.
More On AI Understanding
AI Bias: How Language Models Amplify What They Copy
Researchers studying AI bias thought they were building digital twins. What they created instead were caricatures.
→The stakes become concrete when AI is deployed in high-stakes settings: medical diagnosis, legal reasoning, infrastructure management. Cheryl Chen, a senior lecturer in philosophy at Harvard, points out that when a person says something, we assume they mean it in the full experiential sense. They have encountered rain, felt frustration, formed associations. With AI, that assumption is precisely what Vafa's work puts in doubt.
Vafa is careful not to suggest the situation is hopeless. Current models can still be useful even without coherent world models, he notes, and knowing how to measure the problem is itself progress.
The question now shifts from whether current AI understands to how understanding of this kind might be built, detected, and measured at scale.
Vafa argues that no serious path toward more capable AI can avoid that question.
Sources
- Primary Research: Potemkin Understanding in Large Language Models (ICML 2025, Mancoridis, Weeks, Vafa, Mullainathan)
- Additional Context:
- Does AI understand? (Harvard Gazette, July 2025)
- Evaluating the World Model Implicit in a Generative Model (NeurIPS 2024, Vafa et al.)
Fact Check: Claim-by-Claim Verification Verified
All major claims are accurately supported by peer-reviewed research, author credentials, and institutional affiliations are correct, and the article appropriately represents the findings without misrepresentation.
Commentary
- The article accurately represents research findings but uses accessible language ("invented roads," "leaped across Central Park") that are simplified descriptions of the formal technical findings. This is appropriate for popular science journalism and does not constitute error.
- The article correctly notes that potemkin understanding was observed "across every major model" tested, reflecting the scope of Vafa's research as described in the ICML abstract
- The article's framing that "passing the exam but failing the question" reflects the core insight of the potemkin understanding paper, which contrasts benchmark success with failure on conceptually equivalent but differently-framed tasks
Sources used for verification
Academic/Peer-reviewed:
- Potemkin Understanding in Large Language Models - arXiv
- Evaluating the World Model Implicit in a Generative Model - NeurIPS
- ICML 2025 Poster: Potemkin Understanding in Large Language Models - ICML
- Keyon Vafa Faculty Profile - Harvard Data Science Initiative
- Cheryl Chen Faculty Profile - Harvard Department of Philosophy
Other reliable sources:
- Does AI understand? - Harvard Gazette
- Keyon Vafa Personal Website - keyonvafa.github.io
Fact-checked by Perplexity Sonar Pro on 2026-03-09
