The early twentieth century split combinatorial writing into two streams that would later converge in software: an avant-garde of chance procedures and a science of statistical language. Both reduced authorship to a system of parts plus a rule for combining them — exactly the model a grammar generator follows.
Tristan Tzara's "To Make a Dadaist Poem" (1920) is a literal generative algorithm: cut a newspaper article into words, shake them in a bag, and copy them out "conscientiously" in the order drawn.[1] The Surrealists' cadavre exquis formalized the same idea as a collaborative slot-filling game, where each contributor supplies one grammatical part.[2] These are random rule systems — the recreational ancestors of weighted production rules.
"The poem will resemble you." — Tristan Tzara, 1920
In 1913 Andrey Markov analyzed the sequence of vowels and consonants in Pushkin's Eugene Onegin, founding the theory of chains in which each symbol's probability depends on the previous one.[3] Markov chains remain the basis of the simplest text generators.
Claude Shannon's "A Mathematical Theory of Communication" (1948) showed that English could be approximated by sampling from n-gram statistics, producing eerily plausible nonsense and proving language is partly predictable.[4] This bridged Tzara's chance and Markov's mathematics into a single, computable model of text.