The 23 measurements
The ‘we’ of this page means computer programs. They took the measurements and wrote these pages. No person has checked these numbers by hand. The checking that exists is programs re-running each other's work from the manuscript's own text.
We did not try to read the Voynich text. We measured it instead. We made 23 measurements. A measurement is one thing measured on the text, such as the average length of its words. They are listed below. Several are about signs. A sign is one letter of the manuscript's own alphabet. The word letter is kept below for real languages.
The verdicts on the page ‘What the tests conclude’ rest on these 23 measurements. A verdict is the answer these pages give to one of the seven explanations. A few more measurements, made the same way, each bear on one verdict. One is how often a run of four words appears twice in the same order.
What the measurements are compared with
A measurement on its own says little. So we made the same measurements on 243 samples of real writing. The samples are passages cut from real texts, and real lists of words. Of the 243, 140 are passages from texts. Those texts are in twenty-six languages. Another ninety-two are passages from medieval texts. The other eleven are lists and glossaries of the early fifteenth century. A glossary is a list of words with their meanings. In all, the samples come from forty texts. The texts are in thirty-one languages.
Some measurements grow with the length of a text, as the number of repeated phrases does. For those we cut every sample to the manuscript's length, about 35,000 words, before measuring. Then every figure for the manuscript can be set beside another. The other comes from a text of the same length whose origin is known. On some measurements the manuscript stands far outside every sample.
How an imitation is scored
Each of the seven explanations says how the text was made. Four are these: a real language in a made-up alphabet, a simple cipher, a changing cipher, and a hidden message. A cipher is a disguise that writes each letter or word as something else. A hidden message is one hidden in some feature of the writing, not in its words.
The other three are random text, a list, and a procedure. A procedure is a fixed set of steps carried out by hand with dice and tables of words.
For some of them a computer program writes text by that explanation's rule. Text written that way is an imitation. We made the 23 measurements on every imitation and compared each figure with the manuscript's.
A miss is measured against the manuscript's own wobble. No imitation ever matches a measurement exactly. So a miss needs a measure of its own. To find the wobble, we cut the manuscript into two halves at random, many times over. On each measurement we noted how far apart the two halves' figures usually fall. That distance is the wobble on that measurement.
An imitation matches a measurement when its figure is within three times that wobble of the manuscript's figure. A miss of more than three times the wobble is a clear miss. Only clear misses count as failures.
The explanation we favour is a procedure. These pages call it the best guess. We built three procedures. The first has large tables of 18,000 entries. The second has small tables of a few thousand. An entry is one word or part of a word.
The third has rules worked out from the manuscript's own pages. They come from counting which words copy a nearby word and which are spelled sign by sign. A copied word repeats a word already standing nearby on the page, such as the word directly above. At most one sign is changed. A word spelled sign by sign is built afresh, one sign at a time. Each sign is chosen by the sign before it.
The first procedure, with the large tables, matches 21 of 23 measurements. It has two clear misses. They fall on measurements 4 and 20 in the list below.
The list
The measurements are about the signs, the words, the order of the words and the lines of the page. Some are about guessing. There a computer does the guessing, from how often each sign or word follows another. The manuscript's distinct words are its words with repeats dropped.
| No. | Measurement |
|---|---|
| 1 | How unevenly the signs are used |
| 2 | How easy the next sign is to guess from the sign before it |
| 3 | The average length of a word |
| 4 | The spread of word lengths, how much they vary around the average |
| 5 | The whole pattern of word lengths, the share of words of each length |
| 6 | How steeply the use of words falls away from the commonest word to the rarest |
| 7 | The share of the text taken by the 10 commonest words |
| 8 | The share taken by the 100 commonest words |
| 9 | The share of the manuscript's distinct words that appear only once |
| 10 | How hard the next word is to guess when the word before it is known |
| 11 | How much easier that guess is than a guess with no word known |
| 12 | The part of that gain that word order supplies, found by taking away the gain the same words give when put in a random order |
| 13 | How alike neighbouring words are in their spelling |
| 14 | How often the same word is written twice in succession |
| 15 | How often neighbouring words differ by a single sign |
| 16 | How far the spelling of the first word of a line differs from the spelling of other words |
| 17 | How far the spelling of the last word of a line differs from the spelling of other words |
| 18 | How unevenly the first signs of words are used |
| 19 | How unevenly the last signs of words are used |
| 20 | How much the start of a word tells about the signs that follow |
| 21 | How much the end of a word tells about the signs before it |
| 22 | How often a word is used again on the page where it has just been used |
| 23 | The share of the manuscript's distinct words that differ by one sign from a commoner word |
Two of these matter most for the verdicts. On measurement 2 the manuscript stands outside every sample. Its next sign is easier to guess than the next letter in any of the thirty-one languages. Measurement 9 is the share of words used only once. It is where the best guess fails.
The best guess is the explanation we favour among the seven. It is a procedure carried out by hand with dice and tables of words. Its tables hold words taken from the manuscript itself.
The best guess gets an easier test and a harder one. On the easier test the tables are filled from the whole manuscript. There the large tables match 21 of 23. The best small ones match all 23. The harder test fills the tables from half the manuscript only. So they hold only the words of those pages. The test then asks the procedure to write the other half.
That is done five times over. Each time is one run. In each run we divide the pages in two a different way. We fill the tables from one half and test on the other. Then we swap the halves and do it again. So each run tests both halves.
In the harder test the small tables write too few words that appear only once. About two thirds of the manuscript's distinct words appear only once. In the pages the small tables write, far fewer do. On most of the five runs the shortfall is a clear miss. That is more than three times the wobble. The third procedure was put through the harder test too. It shares that miss. The large tables were never given the harder test. So how they would do on it is not known.
The miss does not overturn the verdict. The procedure stays the best guess. The miss is why it is not confirmed. The best guess page describes the procedure and both tests.