- OpenAI's o1 model passed graduate-level linguistics tests that most AI models failed.
- The model correctly inferred phonological rules from 30 invented mini-languages it had never seen.
- It generated two distinct syntactic trees for an ambiguous sentence without being asked.
Noam Chomsky declared in 2023 that AI models can't truly reason about language, they just marinate in big data without understanding the complicated rules that make language work. A new study from UC Berkeley challenges that view head-on.
As reported in Quanta Magazine by Steve Nadis, Gašper Beguš and his colleagues at Berkeley put several large language models through rigorous linguistic tests - the kind given to graduate students in linguistics.
Most models failed. But OpenAI's o1 model did something remarkable: it analyzed language with genuine sophistication, diagramming sentences, resolving multiple ambiguous meanings, and handling recursion with ease.
Key figure
30
invented mini-languages used to test the AI's ability to infer phonological rules from scratch
The Recursion Challenge
The researchers created tests that models couldn't have memorized during training. One focused on recursion, which is the ability to embed phrases within phrases infinitely.
"The sky is blue" becomes "Jane said that the sky is blue" becomes "Maria wondered if Sam knew that Omar heard that Jane said that the sky is blue."
Recursion is what gives human language its infinite potential from finite rules.
What is recursion?
Recursion is the grammatical rule that lets you embed one phrase inside another, indefinitely. It's why "the cat sat on the mat" can become "She said that the cat sat on the mat" and then "He heard that she said that the cat sat on the mat" – with no theoretical limit. Linguists consider it one of the defining features of human language.
The o1 model parsed tricky sentences like "The astronomy the ancients we revere studied was not separate from astrology." It correctly identified the nested structure, then added another layer of recursion on its own.
Beguš didn't expect to find this metalinguistic capacity, the ability not just to use a language but to think about language.
Making Up Languages
The team invented 30 mini-languages with made-up words to test phonology, the patterns of sounds. Each language had 40 nonsense words following specific rules the models had never seen.
The o1 model correctly inferred the phonological rules from scratch, writing precise descriptions like "a vowel becomes a breathy vowel when it is immediately preceded by a consonant that is both voiced and an obstruent."
It also handled ambiguity with surprising skill. Given "Rowan fed his pet chicken," o1 produced two different syntactic trees, one for a chicken kept as a pet, another for chicken meat fed to a different pet.
"Ambiguity is famously difficult for computational models to capture," said Tom McCoy, a computational linguist at Yale, to Quanta Magazine.
What Makes Us Unique?
David Mortensen at Carnegie Mellon called the results "attention-getting." Some linguists have argued that language models just predict the next word without deep understanding. "This looks like an invalidation of those claims," he said.
No model has yet discovered something about language we didn't already know. But the steady progress is chipping away at abilities once thought exclusively human.
"It appears that we're less unique than we previously thought we were," Beguš said.
Fact Check: Claim-by-Claim Verification Verified
All claims verified against the original Quanta Magazine article, UC Berkeley materials, and independent coverage. Researcher affiliations, study methodology, quotes, and specific examples all confirmed.
Commentary
- The study tests metalinguistic reasoning (analyzing language structure) rather than general linguistic competence.
- Results are specific to OpenAI's o1 model; other LLMs including GPT-4 performed poorly on the same tests.
- The study is a preprint and has not yet undergone full peer review.
Sources used for verification
Academic/Peer-reviewed:
- In a First, AI Models Analyze Language as Well as a Human Expert - Quanta Magazine
- Berkeley linguistics coverage - UC Berkeley
Other reliable sources:
- Chomsky on AI - Fortune
- Secondary coverage - NeurotechNUS
Fact-checked by Perplexity Sonar Pro on 2026-03-14
