What was ruled out, and why
The ‘we’ of this page means computer programs. They took the measurements and wrote these pages. No person has checked these numbers by hand. The checking that exists is programs re-running each other's work from the manuscript's own text.
We tested seven explanations of the Voynich text. Five of them failed. The evidence against each of the five comes first. The two that did not fail come after. Each explanation says how the text was made. For some of them a computer program writes text by that explanation's rule.
Text written that way is an imitation. We scored every imitation on 23 measurements. A measurement is one thing measured on the text, such as the average length of its words. One measurement is how easy the next sign is to guess from the sign before it. A sign is one letter of the manuscript's own alphabet. On this page the word letter is kept for real languages.
We made the same measurements on 243 samples of real writing. The samples are passages cut from real texts, and real lists of words. Among the lists are glossaries of the early fifteenth century. A glossary is a list of words with their meanings. The seven explanations are these.
- A real language written in a made-up alphabet.
- A simple cipher of a real language. A simple cipher always writes the same letter or word the same way.
- A changing cipher. A changing cipher changes its spellings as it goes.
- A message hidden in some feature of the writing, not in its words. The first signs of the words, read in order, are one example.
- Random meaningless text, meaning words set down by chance. In such text a word used on a page is no likelier to be used again there than anywhere else.
- A list or a catalogue, such as a list of names.
- A procedure. A procedure is a fixed set of steps carried out by hand with dice and tables of words.
Each explanation gets one of five verdicts. An explanation fails when text written by its rule comes out unlike the manuscript on the measurements. One fact runs through the verdicts. Every sample of real writing repeats some of its phrases. The manuscript almost never does. A cipher disguises a message. Nobody can see the message behind a cipher. So nobody can see whether that message repeated its phrases. The verdicts assume it did, since every real text sampled does. From the most certain verdict to the least, they are:
- Ruled out: the tests exclude it. No assumption is needed.
- Very unlikely: no way of doing it, tested or untested, can pass if the disguised text repeated its phrases. Every real text does. This verdict rests on that assumption.
- Unlikely: every way we tested fails. A way we did not test could pass. The assumption does not settle it.
- Undecided: nothing found supports it. Nothing rules it out.
- The best guess, not confirmed: closer to the manuscript than any other explanation. It does not pass every test.
A real language in a made-up alphabet
A real language written in a made-up alphabet is very unlikely. Every sample of real writing repeats its four-word phrases. A four-word phrase is a run of four words. It repeats when the same run appears twice in the same order. The manuscript is about 35,000 words long. In Latin prose of that length, 67 to 92 four-word phrases appear at least twice. The manuscript repeats one.
A made-up alphabet only renames the letters. So it cannot remove a repeated phrase. The manuscript's next sign is also easier to guess from the signs before it than any sample's next letter. That is a second measurement where no sample comes close to the manuscript. A language not sampled would fail too, as long as it repeats phrases like every sampled text does. That assumption is needed. So the verdict is very unlikely and not ruled out.
A simple cipher of a real language
A simple cipher of a real language is very unlikely. A simple cipher always writes the same letter or word the same way. So it keeps every repeated phrase of the text it disguises. The manuscript has almost none. So every simple cipher fails, as long as the text it disguises repeated its phrases. Every sampled text does repeat them.
A simple cipher could still hide the breaks between words. So we ran the test again with the spaces between words ignored. This time we counted any stretch of thirty signs written twice. The manuscript has none. The 243 samples include passages cut from real texts. There are 140 such passages. They come from texts in twenty-six languages. Most of them have at least 28 such repeated stretches each. To be exact, 131 do. The other nine do not. Each of those has 26 or fewer.
Four of those are verse by Dante, the Italian poet. That verse has none, like the manuscript. But cut to the manuscript's length, Dante's verse repeats 113 to 155 four-word phrases. The manuscript repeats one. So the verse is unlike the manuscript on the test that matters most.
A changing cipher
A changing cipher is unlikely. A changing cipher changes its spellings as it goes. We tested twenty-two of them. All fail. Michael Greshko published one of them in 2025, with working code, under the name Naibbe. It cuts Latin into pieces of one or two letters. For each piece it draws a card from a shuffled deck. The card chooses which of six tables spells that piece. So the spellings change as the cipher goes.
We built the other 21 for these tests. Each cipher fails in at least one of these ways. It keeps the repeated phrases of the text it disguises. Or its signs are far harder to guess than the manuscript's. Or it writes far fewer words that appear only once. One kind was not tested. That kind keeps one spelling of each word through a page. It changes the spelling at the next page. A changing cipher can hide the repeated phrases of the text it disguises. So the assumption about repeated phrases does not settle it. The verdict is unlikely.
A hidden message
A hidden message of 3,000 words or more is ruled out. A message could hide in some feature of the writing, not in its words. The first sign of every word, read in order, is an example. We wrote a program to search for such messages. To test the program, we planted Latin messages in each of 29 features. Each message was written into a copy of the text.
The program found every planted message of 3,000 words or more. It found none in the manuscript. It missed every planted message of a thousand words or fewer. So a shorter message could be missed. We did not plant messages between a thousand and 3,000 words. So the test says nothing about them.
Random meaningless text
Random meaningless text is ruled out. Random meaningless text is words set down by chance. In the manuscript a word used on a page tends to be used again on that page. Words set down by chance do not do that.
Meaningless text can still reuse its words on a page. Suppose a text is written from a short list of words. The list changes only slowly. Then one list serves for many lines together. The same words come back on a page. The procedure, the best guess, is of that kind. So the verdict rules out random text only, not every kind of meaningless text.
Is the text about anything?
This question is not one of the seven explanations. It is asked here because of the two explanations left, the list and the procedure. A text about something leaves two traces. First, its words go with its pictures: pages with similar pictures use similar words. Second, its commonest words are the little grammar words, such as ‘the’, ‘and’ and ‘of’. Those words are used evenly across the pages.
In the manuscript, pages with similar pictures do not use similar words. And the commonest words crowd onto some pages and stay off others, as rarer words do. So we found no trace of a subject. That counts against any explanation where the text is about something. The language and cipher explanations are of that kind. A procedure writes about nothing. So this finding does not touch it. A list is about something but has no sentences. Whether this counts against it is taken up below.
Not ruled out: a list or a catalogue, and a procedure
The list is undecided and the procedure is the best guess. Here is why.
A list or a catalogue is about something. But it has no sentences. So it might repeat no phrases. Real lists do repeat them. Every list sampled from the early fifteenth century, cut to the manuscript's size, repeats at least 24 four-word phrases. What would repeat none is a catalogue of names, each used once. Nothing found supports that. Nothing rules it out. So the list stays undecided.
A procedure is the best guess, not confirmed. A procedure is a fixed set of steps carried out by hand with dice and tables of words. The tables hold words taken from the manuscript itself. We built three procedures. The first has large tables of 18,000 entries. The second has small tables of a few thousand. An entry is one word or part of a word. The third has rules worked out from the manuscript's own pages.
Those rules come from counting which words copy a nearby word and which are spelled sign by sign. A copied word repeats a word already standing nearby on the page, such as the word directly above. At most one sign is changed. A word spelled sign by sign is built afresh, one sign at a time. Each sign is chosen by the sign before it.
The procedure's imitations come closer to the manuscript than any other explanation's. The large tables match 21 of 23 measurements. That is the easier test. In it the tables are filled from the whole manuscript. The procedure is then scored on the whole manuscript.
The harder test fills the tables from half the manuscript only. So they hold only the words of those pages. It then asks the procedure to write the other half. On those pages the small tables write too few words that appear only once. So does the third procedure. The manuscript's distinct words are its words with repeats dropped. About two thirds of those appear only once. In the pages the procedure writes, far fewer do.
So the procedure is not confirmed. The large tables were never given the harder test. How they would do on it is not known. The best guess page describes the three procedures and that test.