Beinecke MS 408 · independent examination
Voynich Text Examination

What the text of the Voynich manuscript is like, measured on the page scans and on full transcriptions, and compared with real writing in thirty-one languages.

What the tests conclude

The answer

Nobody can read the Voynich manuscript. The tests on these pages found no translation either. The manuscript is a book of about 35,000 words written in an alphabet of its own. We call each of its letters a sign. We keep the word letter for real languages.

The ‘we’ of this page means computer programs. They took the measurements and wrote these pages. No person has checked these numbers by hand. The checking that exists is programs re-running each other's work from the manuscript's own text.

One finding about how the book was made has its own section near the end of this page. Others published it first.

We did not try to read the text. We measured it instead. We made 23 measurements. A measurement is one thing measured on the text, such as the average length of its words.

We made the same measurements on 243 samples of real writing. The samples are passages cut from real texts, and real lists of words. The passages come from forty texts. Those texts are in thirty-one languages. The lists come from the early fifteenth century.

We then tested seven explanations of how the text was made. For some of them a computer program writes text by that explanation's rule. Text written that way is an imitation. We scored every imitation against the manuscript on the 23 measurements. The seven explanations are these.

  1. A real language written in a made-up alphabet.
  2. A simple cipher of a real language. A simple cipher always writes the same letter or word the same way.
  3. A changing cipher. A changing cipher changes its spellings as it goes.
  4. A message hidden in some feature of the writing, not in its words. The first signs of the words, read in order, are one example.
  5. Random meaningless text, meaning words set down by chance. In such text a word used on a page is no likelier to be used again there than anywhere else.
  6. A list or a catalogue, such as a list of names.
  7. A procedure. A procedure is a fixed set of steps carried out by hand with dice and tables of words.

Each explanation gets one of five verdicts. An explanation fails when text written by its rule comes out unlike the manuscript on the measurements. One fact runs through the verdicts. Every sample of real writing repeats some of its phrases. The manuscript almost never does. A cipher disguises a message. Nobody can see the message behind a cipher. So nobody can see whether that message repeated its phrases. The verdicts assume it did, since every real text sampled does. From the most certain verdict to the least, they are:

ExplanationVerdictWhy
A real language written in a made-up alphabetVery unlikelyEvery sample of real writing repeats dozens to thousands of four-word phrases, and the manuscript repeats one.
A simple cipher of a real languageVery unlikelyA simple cipher keeps the repeated phrases of the text it disguises, and the manuscript has almost none.
A changing cipherUnlikelyWe tested 22 changing ciphers and all fail, but one kind was not tested, and it could hide the repeated phrases.
A message hidden in a feature of the writingRuled out for messages of 3,000 words or moreOur search found every planted Latin message of that size, and none in the manuscript.
Random meaningless textRuled outA word is used again on its page far more often than it would be if the manuscript's words were put in a random order.
A list or a catalogueUndecidedEvery list sampled repeats phrases and the manuscript does not, but a catalogue of names, each used once, would not.
A procedure carried out by hand with tables of words and diceThe best guess, not confirmedIts imitations come closer to the manuscript than any other explanation's. With tables filled from half the manuscript, it writes the other half with too few words used once.

Not a language in a made-up alphabet

A real language written in a made-up alphabet is very unlikely. Every sample of real writing repeats its four-word phrases. A four-word phrase is a run of four words. It repeats when the same run appears twice in the same order. Latin prose of the manuscript's length repeats 67 to 92 of them. The manuscript repeats one four-word phrase. A made-up alphabet only renames the letters. So it cannot remove a repeated phrase.

The manuscript's next sign is also easier to guess than the next letter in any of the thirty-one languages. The guess is made from the signs before it. A computer does the guessing, from how often each sign follows the one before. That is a second measurement where no sample comes close. A language not sampled would fail too, as long as it repeats phrases like every sampled text does. That assumption is needed. So the verdict is very unlikely and not ruled out.

Not a simple cipher

A simple cipher of a real language is very unlikely. A simple cipher always writes the same letter or word the same way. So it keeps every repeated phrase of the text it disguises. The manuscript has almost no repeated phrases. So every simple cipher fails, as long as the text it disguises repeated its phrases. Every sampled text does.

A changing cipher

A changing cipher is unlikely. We tested twenty-two changing ciphers. All fail. Michael Greshko published one of them in 2025, with working code, under the name Naibbe. It cuts Latin into pieces of one or two letters. For each piece it draws a card from a shuffled deck. The card chooses which of six tables spells that piece. We built the other 21 for these tests.

Each of the 22 fails. There are three ways to fail. Either it keeps the repeated phrases. Or its signs are far harder to guess than the manuscript's. Or it writes far fewer once-used words. One kind was not tested. That kind keeps one spelling of each word through a page and changes it at the next. A changing cipher can hide the repeated phrases of the text it disguises. So the assumption about repeated phrases does not settle it. The verdict is unlikely.

No hidden message found

A hidden message of 3,000 words or more is ruled out. A message could hide in some feature of the writing, not in its words. For example, take the first sign of every word, read in order. We wrote a program to search for such messages. To test the program, we planted Latin messages in each of 29 features. Each message was written into a copy of the text.

The program found every planted message of 3,000 words or more. It found none in the manuscript. It missed every planted message of a thousand words or fewer. So a shorter message could be missed. We did not plant messages between a thousand and 3,000 words. So the test says nothing about them.

Random meaningless text

Random meaningless text is ruled out. Random meaningless text is words set down by chance. In the manuscript a word used on a page tends to be used again on that page. The reuse is not only in the next few lines. Suppose the manuscript's words were shuffled, put in a random order. A word is used again 1.7 times as often in the real manuscript as in the shuffled text. Words set down by chance do not do that.

The verdict rules out random text only, not every kind of meaningless text. Meaningless text can still reuse its words on a page. Suppose a text is written from a short list of words. The list changes only slowly. Then one list serves for many lines together. The same words come back on a page. The procedure, the best guess, is of that kind.

Is the text about anything?

This question is not one of the seven explanations. But its answer matters for the list, one of the two explanations left standing. A text about something leaves two traces. First, its words go with its pictures: pages with similar pictures use similar words. Second, its commonest words are the little grammar words, such as ‘the’, ‘and’ and ‘of’. They are used evenly across the pages.

We found no trace of a subject. In the manuscript, pages with similar pictures do not use similar words. And the commonest words crowd onto some pages and stay off others, as rarer words do. That counts against any explanation where the text is about something. The language and cipher explanations are of that kind. It does not touch the procedure. The procedure writes about nothing. Whether it counts against a list is taken up next.

A list or a catalogue

A list or a catalogue is undecided. It is about something. But it has no sentences. So it might repeat no phrases. Real lists do repeat them. Every list sampled from the early fifteenth century, cut to the manuscript's size, repeats at least 24 four-word phrases. A catalogue of names, each used once, would repeat none. Nothing found supports that. Nothing rules it out.

The best guess: a procedure

A procedure is the best guess, not confirmed. A procedure is a fixed set of steps carried out by hand with dice and tables of words. The scribe is the person writing the book. The tables and dice would lie on the scribe's desk. The tables are numbered lists of words. In our imitations the tables hold the manuscript's own words. The scribe's real tables, if there ever were any, are lost.

We built three procedures. None is confirmed. The first has large tables of 18,000 entries. The second has small tables of a few thousand. An entry is one word or part of a word. The third has rules worked out from the manuscript's own pages. The large tables miss two of the 23 measurements on the whole manuscript. The small tables match all 23 on the whole manuscript. But they fail one harder test. So does the third procedure.

There is an easier test and a harder test. The easier test fills the tables from the whole manuscript. Then it scores the procedure on the whole manuscript. It is easier because the tables already hold every word of the pages being scored. The harder test fills the tables from half the manuscript only. So they hold only the words of those pages. It then asks the procedure to write the other half.

The harder test is done five times over. Each of the five is one run. In each run we divide the pages in two a different way. We fill the tables from one half and test on the other. Then we swap the halves and do it again. So each run tests both halves.

In the harder test the small tables write too few words that appear only once. The manuscript's distinct words are its words with repeats dropped. About two thirds of the manuscript's distinct words appear only once. In the pages the small tables write, far fewer do.

How far short the small tables fall is measured against the manuscript's own wobble. We cut the manuscript into two halves at random, many times over. The measurement here is the share of words used only once. We noted how far apart the two halves' figures usually fall on it. That distance is the wobble. A miss of more than three times the wobble is a clear miss.

The small tables have a clear miss on most of the five runs. The large tables were never given the harder test. So how they would do on it is not known. The best guess page describes the three procedures and that test.

Two scribes, shared words

One more observation is about the scribes, not the seven explanations. It leaves every verdict as it is. Five scribes wrote the book. Their handwriting can be told apart, as Lisa Fagin Davis showed in 2020. Two of the five use many of the same uncommon words. What produced that overlap is not settled.

One finding about how the book was made

Layfield and Davis found that the two halves of a folded sheet use words more alike than neighbouring pages do. They conclude the book was meant to be read sheet by sheet. Two scholars, Colin Layfield and Lisa Fagin Davis, published this finding first, in July 2026. We came to it on our own, in one bundle of the book, without having read their paper.

A book of this age is made of sheets of parchment folded in half. Fold one sheet and you get two leaves, four pages. A bundle is a set of folded sheets stacked inside each other and sewn. The book is made of about twenty of them.

What is added here is modest. In the last bundle of sheets, take the pages that were once the two halves of one folded sheet. They share more of the book's middling-common words than any other way of folding those same leaves would give. Middling-common words are the words that are neither the book's commonest nor its rare ones.

Those leaves can be paired off with each other in 120 ways. Here is how. Ten of the bundle's leaves are in the pairing. Five are from the front half of the bundle. Five are from the back. A folded sheet always joins one front leaf to one back leaf. The first front leaf can join any of five back leaves. Each front leaf after it has one fewer to choose from. Multiplied together, those choices give 120 ways to fold them into sheets.

One of the 120 is the real folding, the pairing the book's own sheets give. The real folding was not established here. It comes from the published description of the manuscript. The question was whether the words prefer it over the other 119 ways. Ranked against all 120, the real folding comes first. That is the strongest thing the test can say.

The finding lives on one choice of words and dies on another. We counted two different sets of words, chosen two different ways. The result appears in one set and not the other. The first set is the book's middling-common words. There the real folding comes first of 120. The second set is the bundle's own words. Those are the words common inside it and rare elsewhere in the book. There the same test puts the real folding eleventh of 120.

No text our own programs write comes near the manuscript on its middling-common words. We also took real books and wrote them out into the manuscript's own page and line lengths. It is as if a scribe had copied those books into this one. Measured the same way, they fall short of the manuscript on the same kind of words.

The finding is not found in the other bundle tested. The page ‘One bundle of sheets’ gives their result, the checks added here and the limits.

The answer in short

Nobody can read the manuscript. These tests did not read it. A real language in a made-up alphabet is very unlikely. So is a simple cipher. A changing cipher is unlikely. A hidden message of 3,000 words or more is ruled out. So is random text. A list is undecided. The best guess is a procedure carried out by hand with dice and tables of words. It is not confirmed.

One finding is about how the book was made. Layfield and Davis published it first. Take the two halves of a folded sheet. They share more words than any other way of folding the same leaves would give. That holds in one bundle and on one choice of words.