- Markov chains model dependent events using only the current state.
- A 1913 mathematical feud over free will sparked the theory.
- The Monte Carlo method and Google's PageRank both rely on Markov chains.
Andrey Markov sat down with a copy of Alexander Pushkin's Eugene Onegin and began counting letters.
It was 1913 in St. Petersburg, and the mathematics that would become known as Markov chains started with a grudge.
Markov's rival, the Tsarist-aligned mathematician Pavel Nekrasov, had claimed that the law of large numbers proved the existence of free will.
Markov, a committed atheist who had recently requested his own excommunication from the Russian Orthodox Church, thought this was absurd.
Key figure
20,000
Letters Markov counted by hand from Pushkin's Eugene Onegin
A Feud Over Probability
The dispute centered on a 200-year-old assumption. Since Jacob Bernoulli first proved the law of large numbers in 1713, mathematicians had assumed it required independent events.
Nekrasov took this further. If social statistics like marriage rates converged to stable averages, the decisions behind them must be independent acts of free will.
Markov saw a flaw. He stripped Pushkin's poem down to vowels and consonants, 20,000 characters pushed together without spaces. If letters were independent, vowel-vowel pairs should appear about 18% of the time.
They appeared only 6%. The letters were clearly dependent on what came before.
What is a Markov chain?
A Markov chain is a system where the next state depends only on the current state, not on the full history of previous states. This "memoryless" property makes complex dependent systems far simpler to model.
Markov built a prediction machine with two states, vowel and consonant, and transition probabilities calculated from the poem. When he ran the chain, the vowel-to-consonant ratio converged to the exact 43-to-57 split he had counted by hand.
Dependent events followed the law of large numbers after all.
Nekrasov's argument for free will collapsed.
From Poetry to Nuclear Physics
Markov himself showed little interest in practical applications. He wrote that he was concerned only with questions of pure analysis.
Three decades later, his mathematics found an unexpected purpose.
Stanislaw Ulam, a Polish-American mathematician at Los Alamos, was recovering from a near-fatal case of encephalitis in 1946. He spent his convalescence playing solitaire. A simple question nagged him: what fraction of randomly dealt games could actually be won? With 52 factorial possible arrangements, analytical solutions were hopeless.
Ulam's insight was to skip the impossible calculation and just play hundreds of games.
Back at the laboratory, he realized the same statistical sampling could model neutron behavior inside nuclear cores. His colleague John von Neumann recognized that the neutrons formed a chain of dependent events, each step influenced by the last.
They needed a Markov chain.
The pair ran their simulations on ENIAC, one of the world's first programmable electronic computers. They named the technique the Monte Carlo method, a nod to Ulam's gambling uncle and the casino in Monaco.
It is still an unending source of surprise for me to see how a few scribbles on a blackboard could change the course of human affairs.
Stanislaw Ulam, mathematician
The Chain That Built Google
The method's most visible descendant arrived in the 1990s.
Stanford PhD students Sergey Brin and Larry Page realized that web links functioned like endorsements. A page with many incoming links from well-connected pages was probably more valuable than one with few.
More On Mathematics
Can Math Finally Prove We Live in a Simulation? It's Complicated
A physicist proves that self-simulation is possible. Another team proves it is impossible. Both may be right.
→They modeled the entire web as a Markov chain. A hypothetical random surfer clicking links would, over time, spend more time on important pages. A 15% chance of jumping to a random page at each step prevented the surfer from getting trapped in loops.
They called the system PageRank and launched it in 1998. Today Alphabet, Google's parent company, is valued at roughly $2 trillion.
What began as one mathematician's determination to humiliate another produced a framework now woven through modern computing.
Nuclear physics simulations, search engine rankings, speech recognition, and the token-prediction systems inside large language models all trace a line back to Markov's analysis of Pushkin's vowels.
Sources
- Primary Source: The Russian Math Behind Google's Trillion Dollar Algorithm (Veritasium)
- Additional Context:
- First Links in the Markov Chain (American Scientist)
- Andrey Markov (Wikipedia)
Fact Check: Claim-by-Claim Verification Verified
The article accurately recounts the historical origins of Markov chains, their development from Markov's analysis of Pushkin's Eugene Onegin amid his feud with Nekrasov, and key applications in Monte Carlo methods and PageRank.
Commentary
- Article approximates vowel-vowel percentage (~6% observed vs. ~18-19% independent expectation) and simplifies Markov chain definition slightly but correctly captures "memoryless" property.
- Modern applications (Google, AI, nuclear) are standard and well-established; Alphabet valuation updated from article's ~$2T (pre-2025) to current ~$3.7T.
- Dramatic framing ("grudge match," "humiliate") acceptable for popular science; core facts match peer-reviewed histories.
Sources used for verification
Academic/Peer-reviewed:
- First Links in the Markov Chain - americanscientist.org
- [PDF] First Links in the Markov Chain - uconn.edu
- Why Markov Requested Excommunication - probabilityandfinance.com
- [PDF] The Life and Work of A. A. Markov - ufl.edu
Other reliable sources:
- Stanisław Ulam - wikipedia.org
- 8.5.1 Markov Chains (PageRank) - artint.info
Fact-checked by Perplexity Sonar Pro on 2026-03-11
