The Markov Chain Reaction
How an argument about free will travelled through poetry, nuclear physics, Google and AI
On 23 January 1913, Andrei Andreyevich Markov stood before the Imperial Academy of Sciences in St Petersburg carrying an unusual piece of literary criticism.
He had taken the first chapter and part of the second chapter of Alexander Pushkin’s Eugene Onegin and removed everything that usually matters in a novel: the characters, the plot, the irony, the rhythm.
Pushkin had written a novel in verse about love, boredom and Russian society.
Markov classified each one as either a vowel or a consonant. Then he counted what followed what.
A vowel after a vowel.
A consonant after a vowel.
A vowel after a consonant.
He had reduced Pushkin to a sequence of two states.
Markov was not trying to understand the poem. He was trying to settle an argument that had begun years earlier and had little to do with literature.
The mathematics of free will
At the centre of the dispute was Pavel Nekrasov, another Russian mathematician.
Nekrasov was interested in the regularity of social life. The actions of one person may be unpredictable, yet across large populations, rates of marriage, crime and suicide can remain surprisingly stable.
To Nekrasov, this stability carried a moral and religious implication. The law of large numbers seemed to show how independent human choices could produce an orderly society. Individuals remained free, while the collective still obeyed statistical laws.
The mathematics appeared to leave room for both free will and divine order.
Markov objected.
He opposed the Russian state’s entanglement with the Orthodox Church and disliked attempts to recruit probability theory into theology. More importantly, Nekrasov’s argument rested on a mathematical assumption: stable averages required independent events.
Markov set out to show that they did not.
In papers published from 1906 onward, he studied sequences in which one event could influence the next. Under the right conditions, dependent events could also settle into predictable long-run patterns.
A system could carry information from one step into the next without preserving its entire history.
Today, we call such a system a Markov chain.
Dependence does not destroy statistical order. When the sample became large enough, the proportions stabilized.
Why Pushkin mattered
The poem gave Markov something his equations could not: a visible sequence from the real world.
Letters in Russian are not independent. Writers do not reach into an alphabetic bag and pull out the next character at random. The structure of a word constrains the letter that follows.
Markov’s count showed this clearly. In his sample, a vowel was much more likely to follow a consonant than another vowel. Each letter altered the probability of the next.
Yet as the sample grew, the overall proportions became stable.
Dependence had not destroyed statistical order.
The result did not prove that people lacked free will. It did something narrower and more damaging to Nekrasov’s case: it removed the mathematical necessity beneath his claim. Social regularity could no longer be treated as evidence that individual actions were independent.
An argument about theology had produced a new mathematical object.
The feud faded. The object travelled.
A game of solitaire
In 1946, Stanisław Ulam was recovering from illness. To pass the time, he played Canfield solitaire.
He began wondering about the probability of winning. Calculating every possible game was prohibitively difficult. Playing a hundred games and recording the results was crude, but possible.
Then Ulam saw what electronic computers might change.
A machine could play the imaginary games.
Instead of solving an enormous problem directly, it could sample many possible paths through it. Enough simulated paths would reveal the approximate shape of the answer.
Ulam discussed the idea with John von Neumann. At Los Alamos, they and their colleagues applied statistical sampling to neutron transport. A neutron might travel a certain distance, collide, scatter, lose energy or trigger further reactions. Each event changed the probabilities governing the next.
The first large Monte Carlo calculations were prepared for ENIAC in 1947 and run soon afterwards. What began with a game of cards became a way to study processes that branched too quickly for conventional calculation.
Monte Carlo methods and Markov chains are not identical. But they fit together naturally. Markov chains describe movement between states. Monte Carlo methods simulate large numbers of possible movements when exact calculation becomes impractical.
The computer turned probability from something mathematicians proved into something machines could repeatedly perform.
The surfer who never stops
Half a century later, Larry Page was trying to understand the World Wide Web.
Search engines already existed. Most paid close attention to the words written on a page. This made them vulnerable to websites stuffed with popular search terms, whether or not the page deserved to be found.
Page became interested in the links between pages.
A link resembled a citation. But merely counting links was insufficient. A link from an obscure personal page should not carry the same weight as one from a widely trusted site. The value of a page therefore depended partly on the value of the pages pointing towards it.
This created a circular problem.
To know which pages mattered, you needed to know which important pages linked to them. To know which pages were important, you had to perform the same calculation again.
PageRank resolved the circle by imagining a person wandering through the web. At each page, the surfer chooses a link and moves elsewhere. Occasionally, the surfer becomes bored and jumps to a different part of the web.
Web pages become states. Hyperlinks become possible transitions. After enough movement, some pages receive the surfer more often than others.
Those pages receive a higher rank.
The next move depends on the surfer’s current page, not on the complete route used to arrive there.
Markov had counted transitions between vowels and consonants. Google counted transitions between webpages.
The objects had changed. The question had not.
Given where the system is now, where is it likely to go next?
The machines that continue sentences
Claude Shannon asked a related question about English in 1948.
He generated approximations of language in stages. First, he selected letters with equal probability. The result was noise. Then he used the actual frequency of English letters. Later approximations accounted for pairs and groups of letters.
As more of the preceding sequence was considered, the output began to acquire the shape of English. It was still nonsense, but increasingly convincing nonsense.
This became one strand in the development of statistical language modelling: estimate which symbol or word is likely to follow a given sequence.
Modern large language models still learn partly by predicting the next token. That does not make ChatGPT a giant Markov chain.
A basic Markov chain has deliberately limited memory. A transformer can use attention to draw on information from a much longer passage. Its “current state” is not merely the previous word or letter, but a dense numerical representation constructed from the available context.
The family resemblance lies in the question each system asks.
Markov showed that sequences could be studied through conditional probabilities. Shannon showed that those probabilities could reproduce some of the visible structure of language. Modern models use much larger datasets, parameter spaces and computing systems to carry next-token prediction far beyond either man’s experiments.
The distance between counting vowels and generating an essay is immense.
But both begin by asking:
What tends to come next?
The argument that outlived its participants
Nekrasov wanted probability to preserve a place for free will. Markov wanted mathematics kept clear of theology.
Neither man settled the larger argument.
Markov showed that dependence and statistical regularity could coexist. He did not show that human beings were machines. A stable pattern across a population says little by itself about the freedom of the individual inside it.
The mathematical system he helped describe would later be used to model particles, rank webpages, recognise speech, translate languages and generate text. Each application involved a different machine and a different meaning of “state”. None followed automatically from his work.
That is what makes the 20,000 letters worth remembering.
The technologies around us were not assembled from a single master plan. They accumulated through questions whose original stakes have disappeared: a quarrel over religion, a convalescent mathematician’s card game, a physicist’s communication problem, a graduate student trying to organise an unruly web.
Pushkin’s poem survived Markov’s counting.
The counting set off a chain reaction.









