Beinecke MS 408 · independent examination
Voynich Text Examination

What the text of the Voynich manuscript is like, measured on the page scans and on full transcriptions, and compared with real writing in thirty-one languages.

The findings in brief

A sign is one letter of the manuscript's own alphabet. The manuscript's next sign is more predictable than the next letter in thirty-one languages. Its phrases almost never repeat. So a real language in a made-up alphabet is very unlikely. A simple cipher always writes the same letter or word the same way. It is very unlikely too. A procedure is a fixed set of steps carried out by hand with dice and tables of words. It is the best guess, not confirmed. The page ‘Verdict’ says what each of the five verdicts means.

This report does not translate the manuscript, and the measurements say why a translation has been so hard to find. A natural language written in an invented alphabet is very unlikely. So is a cipher of one under any fixed mapping, that is, under any rule that always turns the same letter or word into the same glyphs, the signs of the script. The reason is repetition: an encoding that always turns the same passage into the same glyphs keeps the repeated phrases of the original. Every one of the 140 language samples, the 92 samples of transcribed medieval texts and the eleven lists, glossaries and rolls of the period compared here has dozens or thousands of repeated phrases, where this text has one. Its words repeat as often as the words of a language do, but its sentences never repeat. The ruling out of those explanations holds for every form tested, and for the family as a whole it rests on one assumption, which is that the original text repeats its phrases as every sampled text does. An encoding whose spelling varies by chance or with position escapes that argument, and the one published example, Greshko's Naibbe card cipher, went through every test here. It reproduces the statistics of the glyph sequence and of repetition, but it fails on the lexicon, the stock of words, and on the page structure. The verdict says what a variant would have to do to pass. The code-breaking solvers, programs that search for a cipher key, score that genuine cipher of Latin as they score a meaningless control text (a control being a text of known origin run through the same test). So when they find nothing in the manuscript, that is no evidence against ciphers of that kind.

What the text does have is a rigid structure driven by position, which no sampled language has. Each word is built from a short template of ordered slots, and the choice of glyph depends on where the word is on the line and on the page. Neighbouring words are near-copies of each other far more often than chance allows. Two writing regimes, two distinct styles of writing, coincide almost exactly with two groups of scribal hands. Because of that, the statistics behind the verdicts were re-measured inside each regime alone, against samples of the same length, and none of them changes sides. The repeated-phrase argument is the one that needs the whole text. The wider set of published statistics splits by range. The long-range statistics of the text are inside the language range. They include the shape of the network of words that occur together, the information held by page-sized blocks of text, the similarity of neighbouring pages, and the way the swings in word length grow with the length of text measured. The short-range statistics, which are the predictability of the glyphs, the word template and the near-copying, are outside it. The tests scored four other kinds of text on the same statistics. Two of them are human gibberish, meaningless text that volunteers wrote by hand, and Enochian, the "angelic" language of John Dee and Edward Kelley. The other two are a genuine eighteenth-century cipher and non-linguistic symbol systems, sets of signs that are not writing. Each of the four matches the text on one of those two levels and fails on the other.

For each remaining family of explanation this examination built a generator, a program that writes text by that family's rule, fitted it to the text, and scored it against 23 measured properties. None of them reproduces the text. The closest is a hand procedure of the period, built from tables, lots, a ruler and a renewed working sheet, which comes within three tolerances on 21 of the 23 properties (a tolerance being the range a property wanders over between random halves of the text's pages) and misses two. A search over some forty families of smaller procedure, the kinds of smaller procedure only, finds a procedure of 2,870 table cells (a cell being one entry written out in a table) that matches all 23 on the whole text, as one of 21,305 cells does. Neither holds on pages its tables were not built on, where every candidate misses the spread of word lengths and the count of words used once. The derivation of every word from what the scribe had in sight finds no ruler and less copying than any imitation. It finds the rare words spelled fresh with their length settled first, a step that only the last two candidates carry, and both fail on other counts. What survives, then, is a list of properties and not a mechanism. The list is a rigid word template, an open vocabulary in which new words keep appearing, and similar words clustered by line and by page. It also includes rules for the first and last word of a line, and a trace of word order at the boundary between words, where a rule about glyphs crosses the space. The procedure described below is the best guess, not confirmed. It organises those measurements, and which process produced the text is not established. The sections below test each family of explanation one after another, and the closing verdict says how far each verdict reaches and how sure it is.

Against earlier work. The findings above rest on measurements that other people made first. In 1976 Bennett measured the entropy of the glyphs, the amount of surprise each glyph holds once the one before it is known, and found it unusually low. Lindemann and Bowern set that value below several hundred languages in 2021, in a preprint not yet peer reviewed. Tiltman in 1967, Currier in 1976 and Stolfi in 2000 described the rigid word template, while Zandbergen recorded the absence of repeated phrases and Timm counted it in 2014. What this examination adds comes to three things. Every test runs in one pipeline, a single chain of code, and has its own control. The repetition result is recast in a form that does not depend on where the word breaks fall, so that no fixed encoding can remove it. And each explanation is built as a generator and scored against the whole set of measurements together. The examination also finds that some published results do not survive a wider comparison set. Among them are the claim of Amancio and colleagues in 2013 about how the frequent words are spaced through the text, and the reading-direction asymmetry that Parisel reported in a 2025 preprint. The two-language reading of Arutyunov and colleagues, a 2016 preprint, is another. One published cipher, Greshko's Naibbe cipher of 2025, reproduces the entropy and repetition statistics, and this examination rules it out on the lexicon and the page structure instead. Each section below says who measured what before, whether it held here, and what is new, and the works referred to are listed at the end.

What is new in the method. This examination used 211 techniques. A literature check made after the analysis finds 32 of them with no earlier application to the Voynich text, 102 that adapt published methods with new controls or validation, and 77 that are established. A second count sets aside four unreviewed preprints from 2025 and 2026 and the forum and student projects, and against peer-reviewed and long-established work only the split is 55 new, 103 adapted and 53 established. The new techniques fall into a few groups. Several are tests of what a cipher could and could not hide. A repeated-string test strips out every space and counts the repeated runs of glyphs, which any fixed cipher would keep, and it finds that the text has no long one where nearly every real text has many. An order-invariant repetition test counts repeated runs of words by the letters they contain, so that an anagram cannot hide a repeat, and it finds that anagrammed Latin keeps its repeats while the text has almost none. A first-glyph test measures how much the first glyph of a word tells about the later glyphs. In a nested category code, where the first sign names a broad class and later signs narrow it, that would be a great deal, and in the text it is almost nothing. A homophone test looks for pairs of glyphs that stand for the same letter, recovers planted homophones, and finds none in the text. A grid of simulated codebook ciphers, built from five real texts, shows that no codebook matches the text on word order, vocabulary and near-copying together. Others look for content or a hidden message. A picture-to-word test asks whether the words on a herbal page follow what the picture shows, and it found a link planted as a control and nothing in the manuscript. A residual-channel test looks for a message in the part of each word that the best predictive model cannot predict, and it too found a planted message and nothing in the text. Others measure the text against new kinds of comparison. A measurement on the scans compares the space each line leaves at the right margin with a simulated copyist filling the same lines. A measurement of the information across line breaks shows that the end of a line tells almost nothing about the start of the next. Languages keep a third to two thirds of what they have inside a line. A genuine historical ciphertext and several non-linguistic symbol systems were scored beside the text as reference classes, known kinds of text to compare against. A word-pattern fingerprint records which glyphs in a word are the same glyph, a pattern that no substitution can change, and it places the text nearest to scripts written without vowels. A word-space recovery test gives a program the glyphs with every space removed and asks it to find the word breaks, which it finds better than it finds the breaks in Latin. A recurrent neural model, scored against models that predict each glyph from the few before it, gains a little, and the gain comes from inside the word and the few words before it. A calibrated score of closeness to known languages ranks no candidate language. Once the preprints are set aside, three more count as new: the verbose-cipher solver, the key-transfer test between the two writing regimes, Currier A and B (the two writing styles that Prescott Currier identified in 1976), and the controlled re-runs of the preprints' own measurements. The chart below gives every technique its status under both views, and a table lists what goes beyond the preprints.

The sections below give the measurements behind each sentence of this summary, and for each one they say what it rules out, what remains undecided, and what evidence would be needed to go further. Every number in this report was produced by code run on the primary data during this examination, and separate code derived each number a second time before it was used here.

Verdicts on the main explanations

This examination turned each family of explanation into measurable predictions and tested those predictions against the whole running text, so each verdict below reflects only what the measurements support. Some of the verdicts rest on a property that no encoding can change, and a verdict of that kind has two premises. The first is that the encoding keeps the property, which can be shown. The second is that the original text had the property in the first place, and that is an assumption about language, genre and spelling. The comparison texts support that second premise, but they cannot settle it, which is why each row says which premise its verdict rests on.

Ruled out the tests exclude it, and no assumption is needed   Very unlikely every way of doing it fails, tested or untested, if the message behind it repeated its phrases as all real writing does   Unlikely every way we tested fails, but an untested way could pass, so the assumption about repeated phrases does not settle it   Undecided nothing found supports it, and nothing rules it out   The best guess, not confirmed closer to the manuscript than any other explanation, but failing one test
A natural language written in an unknown alphabet
Three properties of a text survive relettering, which is renaming the signs one for one. They are how well each sign predicts the next, the shape of the word-length distribution, and how often phrases repeat. Predictability is measured in bits of surprise, so fewer bits means a more predictable sign. On the first and the third of these properties the text lies far outside the range of every comparison text. That range covers the forty texts in thirty-one languages, 92 samples of transcribed medieval texts, eleven list-like texts of the period, and every language yet measured. Of the medieval samples, 68 are diplomatic, which means transcribed with the scribe's own spelling and abbreviations kept. The nearest languages on predictability are three Polynesian scriptures with alphabets as small as the text's. They come within 0.41 bits of the text, but they still repeat thousands of four-word sequences. On word-length shape the text lies outside the range of the alphabetic texts only, since the abjads, scripts written without vowels, have word lengths as narrow as the text's. The predictability gap also shrinks when the languages are read with as few signs as the text. At a matched sign inventory the margin to the nearest language is about 0.3 bits, and 0.04 when the text's own letters are merged by the same rule. That leg of the case therefore rests on the order information inside the word. Order information is the surprise the signs would have if each word's signs were shuffled, minus the surprise they have in their real order. The text has 1.14 bits of it against at most 0.95 in any language, and no choice of unit removes that difference. A real syllabic script goes the other way, since Sanskrit written as aksharas, one sign per syllable, scores 4.5 bits per sign and so widens the gap. The word-length leg has two parts, which fare differently: the narrowness holds at every inventory against alphabetic texts, but the symmetry does not. Every language in the samples is ruled out, and for the family as a whole the verdict rests on the assumption that an unknown language written in an unknown way would still fall inside that range. No sampled text contradicts the assumption, and no argument establishes it.
Very unlikely
Simple substitution cipher of any of the thirty-one languages sampled, with or without vowels
The same three properties decide this row, because a simple substitution renames the signs and changes nothing else. Writing without vowels makes matters worse for this explanation, since the abjad texts (Hebrew, Arabic, Persian, Urdu) are the least predictable letter streams sampled, at 3.3 to 3.7 bits.
Ruled out
Any fixed, context-independent cipher or code of a natural-language text (simple substitution, fixed verbose, syllabic or abbreviation codes, fixed word codebooks)
A fixed mapping writes a repeated passage the same way each time, so it keeps the repeated phrases of its source. This row therefore reaches any source text that repeats phrases, and every sampled text does. A source with no repeated phrases, such as a list, a table or a catalogue of unique names, belongs instead to the undecided row at the foot of this board. To test that possibility, this examination measured eleven list-like texts of the period at the text's size. They are two English-Latin glossaries, the Corpus Glossary, a Latin plant synonymy, two martyrologies, ten litanies (counted as one text), two freemen rolls, the Durham Liber Vitae and Pliny. Every one of them long enough to sample at that size repeats at least 24 four-word sequences. The Corpus Glossary, at 5,695 tokens, is too short to sample and has none at its own size. So the assumption also covers the genres written one entry per line. Against all of this, the text has one repeated four-word sequence, where every one of the 140 language samples has at least 67. With the spaces removed the text has no repeated string of 30 glyphs, where 131 of the 140 language samples have 28 to 19,637. Four samples of Dante's verse also have none, and five samples of Montaigne, Cervantes, modern Greek prose and Welsh have 4 to 26. That is why the four-word count, which separates the text from every sample, is the one that settles the argument. For a verbose cipher, which writes each letter as several glyphs, the comparison uses equal plaintext length. There the reference texts have hundreds to thousands of repeats and the text still has none. The disputed glyph distinctions do not close the gap either. Merging every one of them raises the text's four-word count to at most 45, which is still below every language sample merged the same way. The 30-glyph count stays at 0 or 1, and no faithfully transcribed medieval text comes near either figure. Every form tested is ruled out, and the untested forms fail too if the text behind them repeated its phrases, as all real writing does as every sampled text does.
Very unlikely
A cipher whose spelling of a unit varies by chance or with position or state (verbose-homophonic card ciphers such as Greshko's Naibbe; polyalphabetic, autokey and rotating-table schemes)
A cipher of this kind can write the same passage differently each time, so repeated phrases need not survive it and the repetition argument does not reach it. The solvers, which are the code-breaking programs, do not reach it either, because the set of keys they search holds no key like that of Greshko's Naibbe cipher, which encodes Latin by drawing cards. As a result they score Naibbe ciphertext as they score the meaningless control. That control, used throughout this report, is text generated by a second-order Markov model of the glyph stream, a process that picks each glyph from the two glyphs before it. Such text keeps the local glyph statistics and has no content. The Naibbe cipher as published reproduces the glyph predictability, the absence of long repeats and the word-order information, which is how much one word tells about the next. Where it fails is on the vocabulary and the page structure, by eight standard deviations or more (a standard deviation is the usual spread of a measurement between runs). Its share of words occurring once is 0.35 to 0.41 against 0.68 in the text. It has no page clustering, the tendency of a page's words to come back on that page, and no near-copies, which are neighbouring words that are almost the same. Its word-order information also lies in the identity of the units, where the text's lies at the glyph boundary between words. Beyond the card cipher, this examination tested 21 polyalphabetic, autokey and position-keyed schemes. In these ciphers the alphabet changes as the message goes on, by a schedule, by the message's own letters, or by the position on the page. They were 18 state-dependent encodings of Latin and three periodic or position-keyed relettering schemes. Those that remove the repeats raise the entropy, the surprise per glyph, to 3 bits or more. The test for repeats can also allow the glyphs to be renamed between the two copies, and even then no window of 25 glyphs recurs anywhere in the text. Line-keyed, Alberti-style and periodic ciphers of Latin, whose alphabet holds still for a line or a few words, leave 70 to 236 such renamed repeats of 30 glyphs. So the repetition argument reaches those schemes after all. One form remains that is not ruled out: a word-level homophonic nomenclator with separate tables for Currier A and B, the two writing styles that Prescott Currier identified. A nomenclator is a code list with a group for each word, and a homophonic one gives each word several groups. In this form the choice of group would persist across a page and depend on the neighbouring words. No one has built it, and it would hold about 9,000 words of Latin.
Unlikely
Fixed verbose or homophonic cipher (each plaintext letter always written as the same string of several glyphs, or written as one of several glyphs chosen at random)
A fixed verbose code keeps the repeated passages of its source, and every fixed verbose encoding of Latin keeps 3,616 or more repeated 30-glyph strings, where the text has none. This examination built two verbose codes to match the text's glyph frequencies, and even those lower the predictability only to 2.35–2.43 bits, where the text is at 2.11. They also leave the end-of-word predictability and the word-length shape far off. Homophones make matters worse, because one chosen at random raises the predictability above the plaintext's, to 2.7 bits or more, and a genuine eighteenth-century homophonic cipher scores 4.8 bits. The languages in the samples and the schemes tested are ruled out, and for the family the verdict rests on the same assumption as the first row. Verbose codes whose unit boundaries or spellings change by chance or with state belong to the row above. The solvers cannot test those, because the keys they search for give each letter one glyph or one glyph pair.
Very unlikely
Constructed or category language (each word a code built from nested classes)
This examination fitted a nested category code to the text, with a closed inventory of 13,704 branch points and, at each branch, a preference for the commoner choices. That code reaches a hapax share, the share of word types that occur once, of at most 0.34, where the text has 0.68. Even so, a closed inventory with a long Zipfian tail, that is, a long tail of rare codes, can keep producing new codes at this sample size. Codebooks built from real Latin, Italian and Finnish word stocks do match the vocabulary shape, but they fail instead on the near-copy structure and on word order. A second test looks at the first glyph of a word. In a category code the first glyph names a class, so it should narrow the later glyphs, but in the text it adds little beyond what the adjacent glyph tells. By the held-out estimate, measured on pages kept out of the fitting, the first glyph adds at most 0.05 bits, against 0.6 in Latin, 0.4 in Italian and 0.3 in Hebrew. By a plug-in estimate, taken straight from the counts, it adds 0.12 bits, against 0.41 to 0.47 in those languages. The test has little power, though, since a category code built from the text's own statistics scores only 0.08 held out. What the test does reach are codes in which the class chosen first narrows the later choices beyond what the adjacent glyph tells. That is the shape of a classification tree in the manner of John Wilkins's philosophical language, with sub-class labels re-used across branches. It does not reach a complete product code, in which every class has the same sub-classes, because such a code is, statistically, a template of ordered slots and is judged as one. Under that heading the form with independent slots is ruled out, and the form in which each slot depends on its neighbour survives inside the word and fails beyond it. A real invented vocabulary, Hildegard of Bingen's Lingua Ignota, passes the first-glyph test exactly as the text does, so that test does not separate the text from a constructed vocabulary. Two other measures do separate them. The glyph predictability of the Lingua Ignota is 3.0 bits. Its one-edit density, the share of its words that are one letter away from another word, is 18 to 23 percent against 50 to 87 in the text. A constructed language with a grammar faces the same evidence as a natural one, because the text has no repeated phrases and almost no word-order structure. That argument assumes that a text with syntax has both. A code without syntax, a list of items each written as a code, belongs to the undecided row at the foot of this board.
Very unlikely
Text produced word by word from a template, with reference to nearby words (the four copying generators tested)
The copying generators reproduce the local similarity of neighbouring words and much of the page clustering, and no phrase repetition is expected of such a process. Three things count against them, however. The fitted copying kernel, the rule for which earlier word a new word is copied from, is flat, which means it does not favour the word just written. The edits also degrade the word grammar, and none of the four copying schemes tested produces the small but real word-order information the text has, even when fitted to it.
Unlikely
Transposition or anagram cipher of a real text, with or without a substitution on top
A transposition re-orders the letters of a text without changing them, and it can have one of two effects, both of which count against this family. Conditional entropy is the surprise in the next letter once the previous one is known. If a re-ordering changes each letter's neighbours, it raises that entropy to the level set by the letter frequencies alone, which for Latin is a rise from 3.30 to 3.94–3.98 bits. The text is at 2.11, and no reading geometry, which is an order of reading the page, changes that. If instead the re-ordering keeps each letter's neighbours, as reversing the text or moving intact blocks does, the statistics inside the word stay those of the original language, and then the first row's assumption applies. Blocks of four words or more also keep the repeated phrases. Even anagramming would turn repeated phrases into repeated sequences of letter-bags, a letter-bag being the letters of a word with their order ignored. Anagrammed Latin has 111 such repeats, where the text has 1 to 5. The one re-ordering that could lower the entropy is a rule that creates order, such as sorting the letters of each word. Sorting is ruled out, though, because 15% of the sign orders inside the text's words break any single sorting order, and a sorted text has no such violations. An anagram-invariant solver, which ignores the order of letters within each word, recovers five known keys, yet on this text it reaches only wrong-language values. That adds nothing for other languages or for a randomised substitution, so the properties that no re-ordering changes are what settle the family. The nine reading geometries and the anagram forms tested are ruled out. The tests did not cover transposition combined with a randomised substitution or with nulls, or transposition of a non-linguistic plaintext.
Very unlikely
Word codebook (a nomenclator, in which each word of a real text has its own code group), with homophones, nulls and split words
A plain codebook replaces each word of the source with a code group, so the source's word order comes through, and word order is the first thing to compare. The sampled prose in Latin, Italian, German and English has 0.26–0.99 bits of word-order information against 0.11 here, and the whole language set spans 0.11 to 1.23. Latin and Sanskrit epic verse come as low as 0.11, though, and would pass that test. This examination built 620 diluted codebooks, with homophones (several groups for one word), nulls (meaningless filler groups) and split words. Only one corner of that grid, Italian with two homophones and one null in eight, matches both the vocabulary and the word order. Even that one fails the near-copy structure by five to ten times. The 620 forms tested are ruled out, and for the family the verdict rests on the assumption that the source text has the word order of prose or, failing that, cannot produce the near-copy structure. Plain codebooks are therefore ruled out, and with homophones and nulls the codebook is unlikely. One case was not built: a homophonic codebook of a source that repeats and varies its own words within lines, such as a litany or a list. No solver could settle this row, because a key with one group per word needs 44,000 to 47,000 words of ciphertext to be pinned down, and only 34,780 are available. So the case rests on the unchanging properties alone.
Very unlikely
Words as numbers written one unit per digit in a positional system, as Roman-style numerals up to 3,000, or as ordered tables; or a syllabic or letter-pair code of an ordinary text, one glyph per syllable or pair
In a number system the digit slots are independent, so the last unit of a number tells only 0.00–0.22 bits about the one before it. In a text word the last two units share 1.37 bits. Ordered tables are detectable in another way: an order statistic, which scores how strictly the entries follow a fixed order, gives 5–84 for such tables and 1.5 for the text. Syllabic and letter-pair codes need 250–1,400 symbols at 3.0–4.8 bits each, where the text, read with its finest alphabet, uses 158 glyphs at 2.55 bits. Such a code would in any case keep the source's repeated phrases. The forms named are ruled out, but a number system with dependent digits, or an unordered table, is not reached.
Ruled out
A message hidden in a sub-channel of the visible text (first glyphs, the tall "gallows" glyphs, word lengths, stroke counts, one slot)
A derived channel is a stream read off the visible text, such as the first glyph of every word or its length, and this examination tested twenty-nine such channels. Each has 0.98 to 1.01 times the entropy it has after shuffling, so each holds no more order than a shuffle, which is within 0.01 of the meaningless controls. The test has the power to find such a message. It detects a Latin text written into any of these channels from a few hundred symbols, at 0.66 to 0.88 of the shuffled entropy. A language-like message of 3,000 words or more in any of the 29 channels tested is ruled out, in the default transcription. The test does not reach shorter or diluted payloads, channels not tested, or a message compressed or keyed to randomness, which no statistical test detects.
Ruled out
Meaningless text with no memory beyond a few lines (a purely local generator)
The choice of words depends on something that lasts across a page and a leaf. Word types recur above chance across the whole page and on the other side of the leaf, at the strength found in natural language. The illustration sections also have their own vocabularies, beyond what scribe and writing regime explain. No local generator can produce either effect, because every one of them has no memory beyond three lines.
Ruled out
Meaningful content with no repeated phrases (lists, names, numbers written in some other way than those tested) in a non-linguistic notation
This explanation is not ruled out, but nothing positive supports it either. Real lists of the period (glossaries, martyrologies, litanies, rolls) repeat four-word sequences 24 to 1,294 times at the text's size. They also put 9 to 16 percent of their entries on one initial letter, where the labels, the words written beside the drawings, put 54. So what remains is a list unlike any sampled. The drawings do not help to pin it down, because words do not follow the pictures they stand next to. The commonest words are not a grammatical class, zodiac labels recur across pages only at the chance rate, and ordered number tables are ruled out. A known-plaintext attack searches for a key that maps the labels onto a known list of names. Run against Pliny's plant names, Latin star names, eight general vocabularies and a published set of plant identifications, it finds no key. The labels also put half or more of their tokens on one initial glyph, which no name list in any sampled language does and no relettering removes. What remains is content whose words never repeat as phrases and never attach to anything drawn beside them.
Undecided