Beinecke MS 408 · independent examination
Voynich Text Examination

What the text of the Voynich manuscript is like, measured on the page scans and on full transcriptions, and compared with real writing in thirty-one languages.

Is it a language, or a cipher of one?

Is the text a real language in a made-up alphabet, or a cipher of one? A real language is very unlikely. No sample of writing in thirty-one languages repeats as few phrases as the manuscript. A simple cipher always writes the same letter or word the same way. It is very unlikely too. It would keep the repeated phrases the manuscript lacks. Every one tested fails. A changing cipher changes its spellings as it goes. It is unlikely. All 22 tested fail. But one kind was not tested. That kind could hide the repeated phrases.

This part of the report tests the explanations that read the text as a real language, either written plainly or enciphered, and it takes them family by family. Two kinds of argument appear below, and they differ in how far they reach. The first kind rests on invariants, which are properties that no relettering of the signs can change. Three of them matter here: how predictable each glyph is, the shape of the word-length distribution, and how often phrases repeat. Because no relettering touches those properties, an argument from an invariant reaches a whole family of ciphers, on the one assumption that the original text shared the properties with the texts sampled here. The second kind rests on solvers and fitted models, and those reach only the ciphers that were tried, so each section below says which kind of argument it rests on.

Substitution ciphers: decipherment attempts with controls

This examination tried in earnest to decipher the text as a substitution cipher, which is a cipher that replaces each letter of the original with a symbol, following a fixed key. Every code-breaking program, or solver, ran on four kinds of input, which is what stops a false "solution" from being claimed. It ran on the Voynich text, on the same text with the glyphs shuffled inside each word, and on a meaningless text generated from a second-order Markov model of the Voynich glyph stream. Such a model is a random process that picks each glyph from the two before it, so its output keeps the local glyph statistics and has no content. The fourth input was Latin enciphered with known keys, which shows whether the solver can find a key when there is one. A solution would count only under three conditions. It had to score clearly better on the Voynich text than on the meaningless controls (texts of known origin run through the same test), and it had to come near the score of real text in the target language. The third condition was that a key found on the Currier-A pages had to work on the Currier-B pages, the two writing styles Prescott Currier told apart.

The solvers do find a cipher when there is one to find. Each of them searches by simulated annealing, which means that it changes the key a little at a time, keeps a change when the score improves, and sometimes accepts a worse change to avoid getting stuck. The score comes from a character n-gram language model, a table built from a reference text that says how likely each letter is after the few letters before it. On Latin enciphered with a known key, the solvers recovered the plaintext exactly. For a monoalphabetic cipher, which writes each letter as one fixed symbol, they recovered 100% of characters in every restart. For a homophonic cipher, where each letter has two to three symbols to choose from, they recovered 100% once the score charged for the size of the key. And for two verbose ciphers, which write each letter as a string of one to two or one to three symbols, the decoded text matched the plaintext. After those checks the solvers ran on the Voynich text in four symbol views. One view is the plain EVA letters, where EVA is the transcription alphabet that names each common sign with a Latin letter. Two more merge the conventional glyph groups into single symbols, and the fourth is the v101 alphabet, Claston's finer transcription, which splits several shapes the other views merge. In each view they ran against 21 language models, including vowel-stripped Latin, Italian, Spanish, German and English and two Hebrew texts, with and without word boundaries.

Test (EVA view, word boundaries kept)Voynich textMeaningless Markov controlReal text in that language
Best monoalphabetic key, Latin (Caesar) model, bits per character4.754.782.16
Best monoalphabetic key, Italian (Dante) model4.734.752.27
Best monoalphabetic key, Hebrew model4.274.293.30
Best many-to-one (homophonic) key, Latin (Caesar) model4.554.592.16
Best many-to-one key, vowel-stripped Latin model3.943.992.60
Decoded tokens that are real Latin words (≥ 5 letters), best key0.7–4.1%0.7–4.2%–
Decoded tokens that are real Italian words (≥ 5 letters), best key1.9–11.1%4.3–11.5%–

No solver found a key. The grid had 276 cells, four symbol views by three solvers, with and without word boundaries, against 21 language models, and filling it took 1,205 solver jobs. The score in each cell is the number of bits per character that the language model charges for the decoded text, so a lower score means the decoded text looks more like the language. In every cell the best Voynich key scored 0.4 to 3.7 bits per character worse than real text. In every cell it also stayed within −0.28 and +0.25 bits of the best key the same solver found on the meaningless control. The median gap per view was between −0.06 and +0.05. Eight cells beat the control by more than 0.15 bits and nine fell short of it by as much. Margins of that size are of the order of the solver's own spread from run to run. That spread reaches 0.08 bits for the monoalphabetic solver and 0.24 for the many-to-one solver, the one that allows several symbols per letter. So no single cell counts as evidence either way, and the decoded output is pseudo-language of the same kind the control produces. The best Latin decoding of the first line of folio 1r reads nupeiv icus um ugullo veas veami pgemrv i cam veasti, and the best Hebrew decoding contains no Hebrew words beyond chance. Under every model, the share of decoded words that are dictionary words is the same for the Voynich text as for the meaningless control. The vowel-stripped dictionaries show how little that share can mean: they "recognise" half of any consonant string, and they accept 27% of the shuffled control. Without word boundaries the results are the same.

The key-transfer test removes the remaining doubt. In this test the solver fits a key on the Currier-A pages, scores that key on the Currier-B pages, which it never saw, and then does the same in reverse. A key fitted on A scores 0.15 to 0.72 bits per character worse on B than a key fitted on B does, and the reverse transfer loses the same. A cipher of one language would give a gap near zero, and the enciphered Latin does give one. The test is sensitive enough to see a single error, since one wrong entry in a transferred key of 23 letters costs 0.92 ± 0.44 bits per character on that Latin. What the solver calls the "best key" differs from section to section because the two writing styles have different glyph statistics, whereas a real key would work across the whole text.

Loss when a key fitted on Currier A is used on B, and back

A substitution key is fitted on the pages of one Currier language (one of the two writing styles, A or B, that Prescott Currier identified) and then used on the pages of the other. The chart gives the cost in bits per character on those pages with the transferred key, above the cost with a key fitted there. It shows the text, its two meaningless controls and a genuine substitution cipher of Latin, under a one-to-one substitution solver in the EVA view and twelve language models.

Voynich text Markov control shuffled control genuine cipher of Latin -0.4 -0.2 0.0 0.2 0.4 0.6 0.8 1.0 one wrong key entry on Latin: 0.92 no loss English (King James), no spaces English (King James), spaces kept German (Goethe), no spaces German (Goethe), spaces kept Hebrew (responsa), no spaces Hebrew (responsa), spaces kept Italian (Dante), no spaces Italian (Dante), spaces kept Latin, vowels removed, no spaces Latin, vowels removed, spaces kept Latin (Caesar), no spaces Latin (Caesar), spaces kept Genuine Latin cipher, Latin model transfer loss, bits per character
The text's best key does not carry over. Under every language model and in both directions, the transferred key costs 0.15 to 0.72 bits per character more than the key fitted on the held-out pages, the pages the transferred key was not fitted on. The meaningless Markov control, a text generated from the glyph statistics alone, loses 0.11 to 0.70 in the same way. A genuine substitution cipher of Latin transfers with no loss at all. The two writing regimes have different glyph statistics, which is the opposite of what a key that generalises would show. The two markers of each pair are the two directions of transfer.

Two facts about the script frame these results. The first is that the size of the alphabet depends on the transcription. In EVA, 16 letters cover 99% of the text, which is as many as vowel-less Latin uses, and with the glyph groups merged the count is 22 to 24, like English or Hebrew. In v101, though, it takes 38 glyphs to cover 99% of the text and 85 to cover 99.9%, more than any alphabet has. So a one-to-one cipher fits the size of the first two views, but in the v101 view it would need at least 130 of the 158 glyphs to be variants of other glyphs. The second fact concerns glyph predictability, the number of bits of surprise the next glyph brings once the one before it is known. The text scores 2.1 to 2.5 bits and the languages score 3.0 to 3.6, so no language matches the text, and the kind of cipher decides whether that gap could be closed. A monoalphabetic substitution cannot change the number at all, a homophonic cipher raises it, and only a verbose cipher can lower it, which is why this examination checked the verbose solver with particular care. That check covers fixed codes of up to three glyphs per letter, and the Naibbe test below shows that a verbose code with random spellings lies beyond the solver's reach.

The v101 view, with its 158 glyphs, gives the same picture under both solvers for six languages. The Voynich scores differ from the meaningless control by −0.10 to +0.14 bits, and they stay 1.2 to 3.4 bits above real text. The many-to-one solver spends 0.7 to 1.2 bits per character on choosing among homophones, and it spends that much on the Voynich text and on the control alike. One cell of the grid does stand out: the one-to-one solver under the Hebrew model, with word boundaries, scores 0.275 bits better on the Voynich text than on the control. That cell is the −0.28 end of the range given above, rounded, and nothing else supports it. The many-to-one solver on the same model scores the Voynich text worse by 0.09, and the same model without boundaries gives +0.02. The decoded text stays 1.24 bits per character away from real Hebrew and holds no Hebrew vocabulary beyond the control level. Since a search with two restarts produces one deviation of this kind in a grid of this size, the report records the cell as such a deviation and does not count it as evidence.

The verbose solver may delete symbols and merge pairs of symbols, and it recovers the plaintext of a genuine Latin cipher of one to three symbols per letter exactly. On the Voynich text it behaves as it does on the meaningless control. On both it deletes 13 to 17 of the 26 EVA symbols (the 25 letters plus the space), merges about twenty pairs and shrinks the text to half its length. Even so, its output costs 4.0 to 5.8 bits per decoded letter, against 2.5 for the genuine verbose cipher and 2.1 to 3.3 for real text. Its score on the Voynich text is 0.09 to 0.15 bits better than on the control, which is the size of its own spread between restarts, and it moves nowhere near the level of real text. The merged-glyph views count ch, sh, the gallows-with-bench glyphs and the final i-groups each as one symbol. The gallows are the tall glyphs of the script, the bench is the shape written ch, and a gallows-with-bench glyph is a gallows drawn through a bench. These views give the same picture for six languages under both solvers, with every cell within −0.21 and +0.05 bits of its meaningless control and 0.7 to 3.3 bits above real text. The second-largest margin in the whole grid is 0.206 bits, under the Hebrew model with the many-to-one solver in one merged view. It vanishes in the other merged view (−0.07) and in the EVA view (−0.03), and its decoded text contains Hebrew words at the control rate.

Two classical tests that need no language model. The first is the index of coincidence, which is the chance that two glyphs drawn at random from the text are the same glyph. For the glyph stream it is 0.077, which is inside the range of letter streams in natural languages (0.063 to 0.084) and outside the range of polyalphabetic ciphertext (0.045 to 0.053). A polyalphabetic cipher is one that changes the substitution from position to position, following a key, so the value already tells against that kind. The second test looks for a key that repeats at a fixed period, and it does so in two ways. The periodic index compares the text with a copy of it shifted by a fixed number of places, and the Kasiski spectrum measures the distances between repeated strings. Neither shows a periodic key at any period from 2 to 200, in any alphabet view. The highest peak is 1.004 times the baseline, with z at most 5, where z measures the height of a peak in units of its chance spread. For comparison, a Vigenère cipher of Latin, which cycles through alphabets under a repeating key, gives z of 3,687 or more at a 37-symbol period. Its Kasiski value is 2.9 or more, against 1.03 for the text. So a polyalphabetic cipher with a character-periodic key of period up to 200, applied to the glyph stream, is ruled out. Keys that are not periodic in the characters (word-keyed, autokey and state-dependent keys) give no peak in the controls either, so these two tests cannot judge them. The entropy and repetition tests below (entropy being the bits of surprise in the next glyph once the previous one is known) do that instead. A cipher-type classifier is a program that guesses the kind of cipher from the statistics of the ciphertext. One of the kind Nuhn and Knight (2014) described, trained on Latin ciphers, places the text with its own Markov imitation. The spread of the index of coincidence across pages is 3.3 times that of Latin, one more sign of the page-level variation. Lehofer (2022) showed that the homophones of a letter can be recovered by grouping the symbols that occur in the same contexts, and the method as rebuilt here recovers artificial homophone groups perfectly at this sample size. On the text it finds no exchangeable pair among the frequent glyphs, that is, no two glyphs that could stand for the same letter. Once glyph groups are merged it finds three rare pairs whose status cannot be decided, and in Claston's fine alphabet it finds only the plume variants of sh and the borderline pairs p/f and m/iiin. The plume variants score 0.9 to 1.2 on exchangeability, where a value near one means the two behave as one glyph, so the distinctions that behave as one glyph are the ones EVA does not make. That absence of exchangeable pairs matters for one kind of cipher in particular. A homophonic substitution in which the writer chose the homophone freely, by chance or by a card draw, would show such pairs. One in which the neighbouring symbols govern the choice would not show them, and that kind stays in the state-dependent family below.

Luo, Hartmann, Santus, Barzilay and Cao (2021) deciphered Ugaritic and Gothic with a model that also ranks known languages by how close each comes to reading the text, and for Iberian their model found no such language. Their model needs phonetic transcriptions and a training run beyond the means here, so this examination rebuilt their closeness score on the homophonic solver tested above. The score is the solver's gain on the text compared with its gain on the text's own Markov imitation. On the scale used, real text of the candidate language reads 1 and a meaningless text with the same glyph statistics reads 0, and the score behaves as it should on known cases. A monoalphabetic Latin cipher reads as Latin at 1.00 and as Latin's relative Italian at 0.14, which is the same-family signal Luo and colleagues report for Gothic. A Basque ciphertext, from a language with no relative among the candidates, scores 0.05 or less against every one. The Voynich text scores between −0.04 and +0.05 against Latin, Italian, Hebrew, Arabic, Persian, Greek and Turkish with the word spaces kept. With the spaces removed, which is the under-segmented reading the method was built for, it scores between −0.01 and +0.06. That is the pattern of an isolate, a language with no known relative. Read as any of the seven, the text costs 1.4 to 3.0 bits per character more than real text of that language. The share of the decoded text that matches the language's vocabulary at high confidence is 0.02 to 0.10, which equals the share for the text's Markov imitation, against 0.47 for the Latin control. So as a substitution cipher of any of these languages, the text is no closer to one than to another.

The text is long enough for the solvers' failure to mean something. Shannon (1949) defined the unicity distance as the length of ciphertext beyond which the statistics of the language fix the key. It equals the entropy of the key divided by the redundancy of the plaintext, where redundancy is the part of each character that the language makes predictable. Latin has a redundancy of 2.2 bits per character, which is the alphabet's 4.6 bits less the highest of four entropy estimates, and across the nine candidate languages the figure is 1.5 to 2.7 bits. On that basis a simple substitution over the 25 EVA letters is fixed by 37 characters, one over the 38 merged glyphs by 49, and one over the 158 v101 glyphs by 75. A homophonic substitution with 36, 50, 88 or 158 symbols is fixed by 76, 112, 195 or 337 characters. A fixed verbose code of one to three glyphs per letter is fixed by 47, or by 167 when each letter has three alternative strings. Against those distances the running text offers 208,278 characters with spaces, of which 64,836 are in Currier A alone and 7,267 in the labels. That is 530 to 6,200 times the unicity distance for the whole text and 22 to 195 times for the labels. The solvers come within a factor of two of the theoretical limit, since they recover known Latin keys at 44 characters for simple substitution and at 100 to 176 characters for 36 and 50 homophones. So their failure on the manuscript is not a matter of length.

Two families come out differently, each for its own reason. A codebook with one code group per word has a key of 80,000 to 96,000 bits, against about 2 bits of word-level redundancy per token. So its unicity distance is 44,000 to 47,000 words for Latin, or 18,000 under the most optimistic word entropy, which is of the order of the 34,780 words available and four times the length of Currier A. No solver could determine such a key from this text, and that is why the case against codebooks below rests on lexicon and word-order statistics that need no key. The Naibbe cipher, a published card-draw cipher tested in a later section, has a large key too: 2,360 bits with the 414 table strings permuted within their kind, or 14,345 bits with the strings chosen freely. But it expands the plaintext 3.4-fold, so each plaintext letter costs 18 bits of ciphertext, against 2.7 bits of Latin entropy and 3.0 bits of randomness from the card draws and the respacing. Its unicity distance is therefore only 192 to 1,167 plaintext letters, and the 53,000 letters the text would hold exceed that 46 to 277 times. The solvers cannot reach that family for a different reason, which is that their key space contains no table cipher, and text length has nothing to do with it.

Unicity distance of each cipher family against the text's length

The unicity distance of a cipher family is the length of ciphertext after which the redundancy of the plaintext, here Latin, pins down a random key of that family. Each row sets that distance against the running text available to the family, in the same unit. The unit is characters with spaces unless the label says letters or words. The scale is logarithmic.

unicity distance, Latin plaintext running text available to the family red line: the text is shorter than the distance 10 100 1,000 10,000 100,000 1,000,000 Simple substitution, 25 EVA letters 5,597× Simple substitution, 38 merged glyphs 3,583× Simple substitution, 158 v101 glyphs 2,397× Homophonic, 36 symbols 2,355× Homophonic, 50 symbols 1,860× Homophonic, 88 symbols 1,069× Homophonic, 158 v101 symbols 531× Naibbe cipher, smallest key (letters) 277× Naibbe cipher, largest key (letters) 46× Verbose code, 1 to 3 glyphs (letters) 2,257× Verbose code, 3 spellings (letters) 641× Polyalphabetic, period 10 622× Polyalphabetic, period 100 62× Polyalphabetic, period 1000 6.2× Codebook, one group per word (words) 0.74× Codebook, 6,983 groups (words) 0.79× Columnar transposition, width 20 5,648× Columnar transposition, width 500 92× text ÷ dist. characters with spaces, letters or words, as each label says (log scale)
For every substitution, homophonic, verbose, polyalphabetic and transposition family, the running text is 6 to 5,648 times longer than the unicity distance. A solver that fails on the text therefore does not fail for want of text. The two codebooks are the exception. Their keys need more words than the text contains, 47,160 and 44,204 against 34,780. The distances use Shannon's formula with the redundancy of Latin, 2.2 bits per character. No solver was built for the Naibbe family, so its distance shows only that the text would be long enough for one.

One alphabet or several. The whole text uses one alphabet. If different parts had used different alphabets, a key found on one part would not fit another, and the classical same-alphabet test would show it. That test is the chi statistic (Kullback 1938), which compares the glyph counts of two stretches of text and comes out high when both use the same alphabet. The tests computed chi between the five scribal hands, the two Currier regimes, the sections, the paragraph openers, the line openers, the stretches that open with a gallows glyph and the pages. In every case it is 0.92 to 0.99, whereas in Latin controls keyed by page a second alphabet gives 0.37 to 0.55. Even the lowest page pairs (0.61 to 0.67), which are herbal and pharmaceutical pages with unusual word stocks, stay above that range. A search for a relettering that maps Currier A onto Currier B returns the identity, which is no change at all, so the five hands and the two Currier regimes write one alphabet. Some glyph tables and word columns in the manuscript have been read as keys, so the tests treated them as keys, and none of them behaves like one. The order of the glyph ring on folio 57v is unrelated to glyph frequency or to word position, with rank correlations of 0.30 and 0.03, where a rank correlation measures how far two orderings agree. Nor does the ring work as an Alberti disk, a pair of alphabet rings that turn against each other. Used as one, it does not decipher the text (9.62 bits per glyph against 9.59 for random orders). The marginal column on folio 49v is a cyclic sequence with one trigram, a run of three glyphs, repeated three times, which does not make a key. Of the words in the column on folio 66r, 60 percent occur elsewhere in the text, which is the rate for labels, while for running text the rate is 86 to 93 percent. Finally, the tests measured the solvers' reach on a ladder of sizes and inventories, to check that a key could be found from the amount of text available. The solvers recover a monoalphabetic Latin key and a 57-symbol homophonic key from 500 tokens, and they recover a key across a change of genre to 99.4 percent at 1,000 tokens. They recover a 144-symbol homophonic key with the frequency profile of the 158-glyph view to 99 percent of letters at the text's size. The verbose solver recovers 96 to 97 percent of words at every size. Their failure on the text is therefore not a matter of sample size or of inventory.

Earlier work. The argument at the centre of this section has been made before. Zandbergen sets it out on his entropy pages, and Bowern and Lindemann (2021) make it in the Annual Review of Linguistics. A one-to-one substitution cannot lower glyph predictability, homophones raise it, and only a verbose encoding lowers it. This examination ran that argument as a test, on Latin enciphered with known keys, and it holds. Hauer and Kondrak (2016) ran a decipherment search that ranked 380 languages and suggested Hebrew. The solver grid here extends their search by scoring every configuration against two meaningless controls, and the Hebrew runs score no better than those controls. Rozanova and Temerev's 2026 preprint rules the same ciphers out from an attack calibrated in a different way. Three techniques here have no earlier application to this text. The first is the set of index-of-coincidence, periodic-index and Kasiski spectra behind Nuhn and Knight's (2014) cipher-type detection, which rule out a character-periodic polyalphabetic key of period up to 200. The second is the glyph-level homophone detection of Lehofer (2022), of which only the abstract could be consulted, and it recovers artificial homophone groups perfectly and finds no exchangeable glyph pair in the manuscript. The third is the key-transfer test, which fits a key on the A pages and scores it on the B pages, and the transfer loses 0.15 to 0.72 bits where a genuine Latin cipher loses none. The language-closeness score of Luo and colleagues (2021), rebuilt here on the solver tested on known keys, ranks no candidate language, although it does pick out a Latin cipher and its relative Italian together.

The thirty-one languages, and how far each is from the text

Sixteen of the reference texts are European or Hebrew, and nineteen more, in fifteen further languages, cover the language families and scripts that the first sixteen leave out. Together they are thirty-five texts in twenty-six languages, of which twenty-eight are long enough to sample at full size, and the 140 language samples are five windows in each of those twenty-eight. With the five texts at the extremes of the world's language types, described below, the set reaches forty texts in thirty-one languages. The nineteen include the abjads most often proposed for the text, an abjad being a script that writes the consonants and leaves most vowels unwritten. Those texts are Biblical Hebrew at full size, Classical Arabic in the Quran and in a fifteenth-century commentary, Persian and Urdu. The others are Turkish and Uzbek, Sanskrit, Homeric and nineteenth-century Greek, Welsh and Irish, Basque, Russian, and five languages of Africa and Asia (Tagalog, Indonesian, Swahili, Hausa and Somali). Classical Chinese, Japanese, Old Church Slavonic and Quechua could not be sampled at a usable size, and Nahuatl could be sampled only as the modern Huasteca Bible. The same code recomputed every statistic at the size of the Voynich running text, with the repetition counts taken from five positions in each corpus, 140 samples in all. The two substitution solvers ran with a language model for each new language. Before any model was used on the text, it had to recover a known-key cipher of a held-out part of its own text in full, a part kept out of the model's training.

None of the new languages comes close on the two hard constraints. The first is glyph predictability, measured as conditional entropy, the bits of surprise the next glyph brings once the one before it is known. The text scores 2.11 bits, while the new languages score from 2.70 (Hausa) to 3.69 (Arabic). The abjads are the farthest of all (Persian 3.40, Hebrew 3.45, the Quran 3.60), because writing without vowels removes the most predictable letters. The second constraint is the count of repeated four-word sequences, of which the text has one, whereas every one of the 140 samples has at least 67 and every new language has at least 92. Repeated 30-glyph strings separate the text less cleanly. The text has none, but four samples of Dante's verse reach zero as well, and five samples of Montaigne, Cervantes, modern Greek prose and Welsh fall to 4 to 26. That is why the report leans on the four-word count. The new languages do move two softer statistics. The Hebrew Bible comes closer than any of the sixteen core texts on the shape of the vocabulary, as measured by one-edit neighbours. One word is one edit from another when adding, dropping or changing a single glyph turns it into the other. Of the Hebrew Bible's word types, the distinct words, 0.65 to 0.70 are one edit from a commoner type, against 0.757 for the text. But its word-length spread, the standard deviation of its word lengths, is 1.2 to 1.3 against 1.79 for the text, so it misses from the other side. Latin and Sanskrit epic verse match the text's weak word-order information (0.11 to 0.12 bits against 0.113), where word-order information is how much one word tells about the next. No language, though, is close on both fronts together. The languages nearest on glyph predictability (Maori, Hawaiian, Tongan, Hausa, Indonesian, Swahili, Tagalog) are the farthest on repetition and word order, and the abjads, nearest on word shape, are the farthest on predictability. Ranked by distance over fourteen statistics, the five closest texts are still European, with Dante's Italian first.

The typological extremes. The languages nearest the text on predictability all have small alphabets, so this examination extended the comparison set to the two extremes of the world's languages, where the alphabets are smallest and where the words are longest. At one extreme are three Polynesian languages, which have the smallest alphabets and the simplest syllables in use. The texts are the Hawaiian Bible of 1868, the Tongan Bible and the Maori New Testament, each in its nineteenth-century spelling without vowel-length marks, with Tongan and Maori also sampled with the marks. At the other extreme is a polysynthetic language, one whose words are whole sentences, represented by the North Alaskan Inupiatun New Testament. The examination also added Huasteca Nahuatl at full size, and all five texts come from the same verse-per-line scripture files, cleaned and matched in size as before.

The Polynesian texts lower the language floor of glyph predictability, the lowest value any language reaches. The floor had been 2.70 bits (Hausa), and Maori gives 2.52 (2.68 with vowel length marked), Hawaiian 2.65 and Tongan 2.65. They need 14 to 16 letters to cover 99 percent of their text, as the Voynich text does with 16, whereas every other sampled language needs 17 or more. Maori and Hawaiian also match the text's fourth-order predictability, the surprise of the next glyph once the four before it are known (1.78 and 1.77 bits against 1.77). At the level of words the resemblance fades. Maori's words are short, 3.35 letters against the text's 4.99, but their lengths vary more than the text's (standard deviation 2.18 against 1.79). Its one-edit density, 0.37 against the text's 0.76, is inside the range of the other alphabetic texts. So the margin on predictability at the letter level is 0.41 bits, and not the 0.6 that the Hausa floor gave. In the v101 reading (2.55 bits) a Polynesian text without vowel marks is below the text.

Everything else goes the other way, and by a wide margin. Hawaiian and Maori repeat 2,299 to 4,097 four-word sequences in every one of ten windows (97 to 270 with the words shuffled), and they also repeat hundreds to thousands of 30-letter strings. Tongan repeats 3,411 four-word sequences and 2,243 such strings. Their once-used share of types, the share of distinct words that occur only once, is 0.38 to 0.41 (the text 0.68). Their ten commonest words make up 33 to 43 percent of the tokens (the text 13). Their adjacent-word information is 1.01 to 1.66 bits (the text 0.11), and their vocabularies hold 1,413 to 1,570 types (the text 6,983). Inupiatun has the shape a polysynthetic language predicts: its words average 10.5 letters, it has 15,149 types and a once-used share of 0.80, it repeats 91 to 455 four-word sequences, and its predictability is 3.26 bits. On the fourteen-statistic distance none of the five comes closer than Hausa (2.34): Nahuatl 2.36, Maori 2.59, Tongan 2.78, Inupiatun 2.84, Hawaiian 2.87. The closeness score reads the text as none of them, between −0.15 and +0.02, even though a substitution cipher of each language scores 0.93 to 1.01 and is solved to the last token. Read as any of them, the text costs 2.8 to 4.0 bits per character more than real text of that language, farther than any earlier candidate. So a language of the Polynesian type is the one kind that comes near the text at the glyph level, and it is the farthest of all on repetition and on the shape of the vocabulary.

Table: the twenty-eight reference texts that were sampled at full size, ranked by their distance from the Voynich profile
RankTextDistance from the Voynich profile (sd units, 14 statistics)Conditional entropy (bits)Repeated 4-word sequencesRepeated 30-glyph stringsTypes one edit from a commoner typeWord-length spreadWord-order information (bits)
–Voynich (Zandbergen–Landini reading)02.11100.7571.790.113
1Italian, Dante1.773.1111300.4322.250.384
2German, Nibelungenlied (Simrock translation)1.993.11390530.4492.160.528
3Latin, Caesar2.113.3067400.2702.930.255
4Spanish, Cervantes2.123.034861100.3002.570.434
5French, Montaigne2.123.17261420.3012.670.458
6Somali, Quran translation2.183.127841,0410.3562.580.587
7Latin, Virgil2.233.36971840.3272.200.114
8Biblical Hebrew, Tanakh (consonantal)2.243.455445240.6861.250.453
9Swahili, Quran translation2.262.931,4101,5430.2492.520.803
10English, Shakespeare2.263.323251540.3432.000.411
11Tagalog, prose2.302.941,3072,3790.3052.660.420
12Hausa, Quran translation2.342.702,4461,8340.4462.051.226
13Greek, 19th-century prose2.413.23701850.2282.840.355
14Indonesian, Quran translation2.422.891,8032,5370.1632.390.979
15Welsh, prose2.463.35844820.4092.120.724
16Persian, Quran translation2.463.401,0024770.4321.910.654
17Irish, prose2.503.141,5844640.3662.640.958
18Basque, encyclopedia articles2.563.301,2231,5580.1823.431.009
19Finnish, Kalevala2.583.201,1734,0190.2812.590.481
20Homeric Greek, Iliad and Odyssey2.623.379712,2530.3122.570.496
21Uzbek, Quran translation2.633.384017120.2332.790.386
22English, King James Bible2.653.132,6013,3750.2732.070.960
23Classical Arabic, 15th-century commentary2.703.696169840.4681.660.307
24Classical Arabic, Quran2.723.601,4621,5540.4911.610.608
25Russian, Quran translation2.773.431,6222,2570.2852.950.806
26Urdu, Quran translation2.793.261,8387860.5081.291.060
27Turkish, Quran translation3.053.364,17119,6370.2343.001.033
28Sanskrit, Ramayana3.203.382025190.1833.980.121

The solvers behave on the new languages as they did in the substitution-cipher tests above. There were 100 runs, covering 19 corpora, two solvers, with and without word boundaries, and a second symbol view for six languages. The best key found for the Voynich text scores between 0.19 bits per character worse and 0.14 bits better than the best key found for the meaningless control, with a median of 0.01 bits worse. Seven runs beat the control by more than 0.10 bits, so the tests ran those seven again with new random seeds (a seed being the starting point of the random numbers), and every margin shrank (Sanskrit from 0.19 to 0.02, Indonesian from 0.15 to 0.01). None stayed beyond 0.12, which is inside the solver's own spread between restarts. The decoded text stays 1.4 to 3.5 bits per character above real text in every language, and its dictionary hit rate equals the control's.

Ruled out: simple and homophonic substitution of any of the twenty-six languages tested. The confidence is high for the six languages run on both symbol views, Hebrew, Arabic, Persian, Turkish, Sanskrit and Homeric Greek, and medium for the thirteen further texts run on one view. Chinese, Japanese, Old Church Slavonic and Quechua were not tested, so they are not ruled out, and the verbose-cipher solver and the key-transfer test were not run for the new languages.

One fingerprint of a text survives any substitution, and that is the pattern of repeated letters inside a word. The pattern, or "isomorph", of daiin is 1-2-3-3-4, because its third and fourth glyphs are the same. Once the conventional glyph groups are merged, Voynich words contain fewer repeated glyphs than the words of every alphabetic language sample: 15 to 16% of tokens hold a repeat, against 24 to 72% in the languages. The distribution of patterns is nearest to the consonantal scripts (Urdu, Arabic, Hebrew, Persian), at twice the sampling floor, the distance two samples of one text show by chance. That statistic alone rules out a one-to-one substitution of Latin, Italian, German, English, French, Spanish or Greek, unless the intended units are the single EVA letters. In that case the repeats are all doubled e and i strokes, and the distribution is far from every language. The fingerprint does not identify a language, however, because the Markov imitation reproduces it, and so does the slot generator, a text generator that builds each word from a template of slots. Four abjad corpora and the Aeneid differ from the text by 0.037 to 0.052, which does not separate them, so the fingerprint is consistent with a vowel-less plaintext and equally with no plaintext at all.

Mixed plaintexts do not escape either. This examination mixed Latin with Middle High German or Italian in four ways. The four were blocks, word-by-word mixing, macaronic style (the two languages mixed within sentences) and coined nouns (invented words in place of the nouns). All four raise the conditional entropy to 3.3 to 3.4 bits, which moves them away from the text, and only word-by-word alternation removes the repeated four-word sequences (0 to 12). Even that comes at a price, because the adjacent-word information then falls to 0.07 to 0.10 bits, below the text's 0.11.

Earlier work. Hauer and Kondrak (2016) tested the text against a large sample of languages, 380 of them, instead of a single favourite. Lindemann and Bowern did the same in their 2021 preprint, which compared character entropy across 294 languages and eighteen historical texts. The survey here is smaller, but it runs every statistic where they ran one, at a matched size, with a solver and a known-key positive control for each added language. A positive control is a cipher with a known answer, which the solver must break before its result on the text counts. The survey's result confirms their position from the other side: none of the nineteen added corpora, including the abjads most often proposed, comes close on glyph predictability and phrase repetition together. The abjads are the farthest of all on predictability, because writing without vowels removes the most predictable letters. Hauer and Kondrak used the word-pattern fingerprint to identify languages, and any one-to-one substitution preserves it exactly, which is what makes it usable here. This examination applies it to the Voynich text for the first time. It rules out a substitution of seven European languages with no language model at all, and it places the text nearest the consonantal scripts. That placement is an observation about repeat rates, though, and does not identify a language. Bennett (1976) made the Polynesian comparison first, when he set the text beside Hawaiian. Lindemann and Bowern's 2021 preprint lists Hawaiian, Maori and Tongan at 2.77 to 2.95 bits, with Nahuatl and Inupiaq at 3.23 and 3.35. The values depend on the spelling, as Stallings (1998) noted of Bennett's Hawaiian text. Without vowel-length marks, as in the nineteenth-century printings sampled here, they come out lower (2.52 to 2.65), so the margin below the text is narrower than the published comparisons make it. No precedent was found for the mixture and coined-noun plaintexts.

Transposition and anagram ciphers

A transposition cipher keeps the letters of a text and changes their order. Suppose the letters of a real text were re-ordered inside each word, inside each line, or along a route across the page, and then relabelled. Every statistic that depends on order would change, but the set of letters in each word, line or page would not. So the tests here use quantities that do not depend on order, and three of them settle the family.

First, any transposition that changes a letter's neighbours pushes the conditional entropy up to the unigram entropy. Conditional entropy is the surprise the next letter brings once the letter before it is known, and unigram entropy is the surprise when nothing before it is known. A transposition that keeps a letter's neighbours, such as reversal or the moving of whole blocks, leaves the within-word statistics of the original language in place, and the sections above judge those. Latin goes from 3.30 bits to 3.94 when its letters are shuffled inside words, and to 3.98 when lines or pages are transposed. The Voynich text, by contrast, is at 2.11, with 1.3 to 1.5 bits more structure between neighbours than any language has. None of nine reading geometries lowers its entropy or creates repeats. The geometries are reversed lines, boustrophedon (lines read alternately in each direction), columns across a page or across a block of lines, and diagonals. Each of them does to the text what it does to the meaningless control. Only a fixed sorting rule can lower entropy the way the text requires (sorted Latin: 2.50). But sorting leaves no precedence violations at all, where a precedence violation is a pair of letters inside a word written in the less usual of its two orders. The Voynich text has 15% and the languages 31 to 41%, so the text is not sorted either.

Second, repeated phrases survive anagramming as repeated sequences of letter-bags, a letter-bag being the set of letters in a word with their order ignored. Anagrammed and relabelled Latin has zero repeated four-word sequences but 111 repeated four-bag sequences, so the phrases are still there to be counted. The Voynich text has 1 to 5, its control 1, and the languages 33 to 2,414. Third, a solver that ignores letter order recovers the true key for anagrammed Latin, Italian, English, German and Hebrew in every case. It scores each candidate key by how well the letter-bags of the decoded words match a held-out word list, and it searches by simulated annealing. On the Voynich text it reaches only the values of a wrong language: 24 to 29 bits per token, against 9 to 19 for the controls it solved. That is within half a bit of what it reaches on the meaningless control.

TestVoynich textMeaningless controlLanguagesPositive control (known key)
Conditional entropy after neighbour-changing transposition (bits)2.11 as read; 3.8+ under every routesame behaviour3.0–3.6 → 3.9–4.0Latin 3.30 → 3.94/3.98
Precedence violations inside words (sorted rule gives 0%)15.3–15.9%–31–41%sorted Latin 0%; random anagram 50%
Repeated four-word sequences counted on letter-bags1–5133–2,414anagrammed Latin 111
Anagram-invariant solver, bits per token (lower is more language-like)24.2–29.1within ±0.5 of the text–key recovered exactly in 5 of 5 languages, 8.9–18.6

Ruled out for random and sorted anagramming inside words, and for line-level, page-level and reversed reading orders (high confidence). The tests found no evidence for a remaining transposition inside words combined with a fixed substitution of the five languages tested. Transpositions combined with randomised homophonic substitution were not run.

Parisel (2025) reported that the text is easier to predict read forwards than read backwards, by 0.065 bits per character, and that the languages are not. The tests here do not reproduce that. With a definition that treats the two directions the same way, every text, the Voynich included, gives a difference below 0.001 bits, which is what theory expects. For a stationary sequence, one whose statistics do not change along its length, the forward and backward conditional entropies are equal in the limit. The small residuals here come from the smoothing of the estimate, and the shuffled controls show the same residuals. A difference of the published size can only come from treating the two ends of a unit differently. Under such a reading the text would score as reported, because its word-final and line-final glyphs are far more predictable (2.5 to 2.6 bits) than its initial ones (3.2), which is the ordered-slot template described above. Even under that reading, the published signs for the languages do not come back, so reading direction is not an independent line of evidence.

Earlier work. D'Imperio (1978) listed transposition among the possibilities, and it has stayed on the list since. Hauer and Kondrak (2016) made it concrete when they treated Voynich words as anagrams sorted into alphabetical order and searched for a source language. This examination reproduced their order-blind solver, which recovers the true key for anagrammed Latin, Italian, English, German and Hebrew. On the manuscript it reaches only control values, so a re-run of their own kind of test does not support their Hebrew suggestion. Two techniques in this section have no earlier application to this text. One re-reads the glyph stream under nine page geometries and scores entropy and repetition against known-key controls. The other counts repeated sequences of letter bags, which extends the phrase-repetition fact to a statistic that anagramming inside words cannot hide. The reading-direction asymmetry reported in Parisel's 2025 preprint does not reproduce: under a direction-symmetric definition every text here, the manuscript included, gives a difference two orders of magnitude below his published value. His pattern of signs across languages does not appear either.

Codebooks, number systems and syllabic codes

Codebooks. A codebook gives each word of the plaintext its own code group, so if each Voynich word stood for a word of a real text, the word order of that text would come through. Adjacent words in Latin, Italian, German and English share 0.26, 0.39, 0.53 and 0.99 bits of information, on the same measure that gives the Voynich text 0.11, and a codebook would have to close that gap. This examination built six hundred and twenty synthetic codebook texts to see whether three devices could bring a real text down to the Voynich values on every statistic together. The devices were homophones (several code groups for one word), null groups (code groups that mean nothing) and word splitting, and each has a cost. Homophones dilute the word-order information about as fast as they inflate the vocabulary, since Latin with two groups per word gives 0.11 bits, but 11,400 word types against 6,983. Nulls lower both, but they push up the share of the ten commonest words, and splitting long words raises the word-order information. One narrow corner matches vocabulary shape and word order together, Italian with two homophones and 12 to 15% nulls, but that corner fails the near-copy statistics by a factor of five to ten. Its adjacent identical words run at 0.1 to 0.2% against 0.9% in the text, and its adjacent one-edit neighbours at 1.6 to 1.8% against 3.6%.

Number systems. Words that are numbers leave signatures that no relabelling can hide. The first is that the digit slots of a number are nearly independent of each other. A positional table gives 0.00 bits of information between the last two units of a word and Roman numerals give 0.2, whereas Voynich words give 1.37 and Latin words 0.8 to 1.4. The second is that number tables are ordered. The order statistic is how often neighbouring entries share their first two units, as a multiple of chance. Every ordered numeral table gives 5 to 84, whereas the Voynich text gives 1.5 and a shuffled text 1.3. A third test, the leading-unit test, compares the spread of the first unit of each word with Benford's law, the skewed distribution that the leading digits of real numbers follow. It decides nothing, because the text's distance from that law is 0.086 and the meaningless Markov control's is 0.087. Roman-style numerals do reproduce the low glyph entropy, but they reproduce nothing about the vocabulary.

Syllabic and letter-pair codes. In such a code one glyph stands for a syllable or for a pair of letters, so each symbol holds more information than a letter does. Syllabified Latin, Italian, Spanish and Finnish need 900 to 1,400 symbols, with a conditional entropy of 3.8 to 4.4 bits, and letter-pair codes need 250 to 490 symbols at 3.0 to 4.8 bits. The finest Voynich transcription, by contrast, uses 158 glyphs at 2.55 bits, and 38 of its glyphs cover 99% of the text. So a Voynich glyph holds less information than a letter, which is the opposite of what such a code produces. And because such a code always writes the same syllable the same way, the 83 to 2,113 repeated four-word sequences of the source texts would survive it, whereas the Voynich text has one.

TestVoynich textMeaningless controlLanguages or plaintextsBest codebook corner (Italian, 2 homophones, 12–15% nulls)
Adjacent-word information (bits)0.113 ± 0.016 (page halves; ± 0.03 by quire, a gathering of leaves)0.03–0.040.26–0.990.117–0.142
Word types / hapax share6,983 / 0.687,700 / 0.732,500–8,700 / 0.41–0.627,300–7,500 / 0.66
Adjacent identical words / adjacent one-edit neighbours0.94% / 3.6%0.53% / 2.3%0.04–0.06% / 2.1–2.7%0.08–0.19% / 1.6–1.8%
Information between the last two units of a word (bits)1.371.380.82–1.44positional numerals 0.00; Roman 0.17–0.22
Symbols needed for a syllabic or letter-pair code / their conditional entropy158 (38 cover 99%) / 2.55–900–1,400 / 3.8–4.4 (syllables); 250–490 / 3.0–4.8 (pairs)–

Ruled out for plain codebooks of the five plaintexts (Caesar, Virgil, Dante, the Nibelungenlied and the King James Bible). The same verdict covers positional, Roman-style and ordered number tables, and syllabic or letter-pair codes of any normal text (high confidence). A nomenclator is a codebook of the classical kind, one that gives words their own code groups and allows homophones and nulls. A diluted nomenclator with about two homophones and one null in eight passes on vocabulary shape and word order, so those two statistics do not rule it out. It fails the near-copy structure, though, and it offers no test that would separate it from a meaningless text with weak word order.

Hermes (2022) proposed Trithemius's Polygraphia III as the model of a cipher that pretends to be an artificial language. It is a table cipher that writes one invented word for each letter of the plaintext, and Hermes showed that its output has low glyph entropy, narrow word lengths and dense families of near-identical words. The tests here reproduce all three, but they also rule out the class as a model of the text, on the same grounds as the nomenclators above. Such a cipher inherits the letter-pair structure of the plaintext as word order, at five to twenty-five times the text's value. It has a closed vocabulary, with as many types as the table has entries, no words occurring once, and a vocabulary that stops growing within 10,000 words. Above all, it repeats a passage word for word wherever the plaintext repeats a string of letters. Any letter-to-word table of 6,600 entries or fewer, applied to 34,780 letters of Latin, produces at least 241 repeated four-word sequences and 1,357 repeated 30-glyph strings, no matter what the table contains. The text has 1 and 0.

Earlier work. Blog and forum writers have discussed the codebook and nomenclator possibilities without measuring them, and the syllabic reading goes back to Stolfi's (2002) argument from the size of a syllable inventory. The simulation grid used here, 620 synthetic codebook texts with homophones, nulls and word splitting, scored on every statistic together, has no earlier application to this text. Neither have the leading-unit test, which looks at the first unit of each word, and the ordered-table test used against the numeral reading. Hermes (2022) proposed the Polygraphia III construction, and the tests above rebuild it and rule it out on phrase repetition. The entropy literature implies the general point that a syllabic or letter-pair code makes each symbol hold more information than a letter does. This section spells it out with the symbol counts and entropies of four syllabified languages.

Ciphers whose spelling varies: the Naibbe card cipher and state-dependent schemes

The repetition argument made above, that every real text repeats its phrases and that a fixed encoding would keep those repeats, reaches only the encodings that always write the same passage the same way. Some ciphers spell the same unit differently from one occurrence to the next, whether by chance, by its position or by some internal state of the cipher, and a cipher of that kind need not keep repeated phrases. So other statistics have to judge it. One scheme of this kind has been published with working code, and this examination put it through every test.

That scheme is Greshko's Naibbe cipher, published in Cryptologia in 2025, and it enciphers Latin or Italian. It first cuts the stream of letters into units of one or two letters, choosing each cut at random, and for each unit it draws a card from a shuffled deck. The card picks one of six tables, the table gives the unit's spelling as a Voynich-like string, and a two-letter unit also gets a prefix string and a suffix string. Because both the cuts and the cards are random, the same passage of plaintext can come out differently each time it occurs, which is how the scheme escapes the repetition argument. This examination rebuilt the cipher from the code and tables in Greshko's repository, with a fixed random seed, and checked the code's constants against the original. The rebuilt cipher then enciphered Caesar and Dante at the size of the running text, with both deck sizes, and Greshko's own ciphertext of Pliny made a fourth text. All four texts were poured into the manuscript's own layout of pages and lines, the page skeleton that every generator below also uses, and each then went through every statistic in this report. The four texts also went through the solvers, the code-breaking programs of the substitution-cipher section above, and through the key-transfer test. That test fits a key on one half of a text and scores it on the other half.

The result narrows what is ruled out in one place and confirms it in another. It narrows them because the Naibbe texts reproduce the statistics that rule out the fixed verbose ciphers above (a fixed verbose cipher writes each letter as the same string of several glyphs every time). Conditional entropy, which is how many bits of surprise the next glyph carries once the previous one is known, is 2.09 to 2.11 bits for the Naibbe texts against 2.11 for the text. They have one repeated four-word sequence, no repeated string of 30 glyphs and a longest repeat of 21 to 25 glyphs, so the long repeats are missing from them as they are from the text. Their word-order information, the amount that one word tells about the next, is 0.09 to 0.10 bits against 0.113 for the text. Their held-out bigram gain, which is how much a model that has learned word pairs improves on text it has not seen, is at the text's value too. The substitution solvers cannot tell them apart either, and give Naibbe-of-Latin the same score as the text and as the meaningless control. That control, used throughout this report, is a text generated from the glyph statistics of the Voynich stream, and it has no content. The scores are 4.55 and 4.46 bits per character for Naibbe, 4.56 and 4.45 for the text, and 4.58 and 4.49 for the control. The solvers decode only 0.4 percent of the Naibbe units. They fail because their search covers only keys that give each letter one glyph or one glyph pair under a fixed rule, and no key like Naibbe's is in that space. So their failure on the text is no evidence against ciphers of this kind, and the amount of text is not the reason either. The unicity distance of a cipher family is the amount of ciphertext a solver needs before only one key can fit. For the Naibbe family it is 192 to 1,167 plaintext letters, where the text would hold 53,000.

On the other statistics, those of the vocabulary and the page, the Naibbe texts lie eight standard deviations or more from the text, and this is where the ruling out is confirmed. Their hapax share, the share of words that occur only once, is 0.35 to 0.41 against 0.68. Their Zipf slope, which measures how fast word frequencies fall away from the commonest word down the ranks, is −0.90 against −1.04. Adjacent identical words are 0.13 percent of their word pairs, against 0.94. Their page burstiness, the tendency of a word to cluster on one page, is 0.98 against 1.96. Their recurrence profile, which tracks how often the words of one line come back at each distance in lines, is flat at 0.8 to 1.1 at every distance. The text's is 4.6 at the neighbouring line and decays to 0.9. Words one edit apart, which differ by a single glyph, are neighbours in context no more often than chance in the Naibbe texts, at 0.02 against 0.086 in the text. Even the word-order information, which the Naibbe texts reproduce in amount, is in a different place from the text's. It lives in the identity of the units and the word boundary holds almost none of it, so the last glyph of a word tells 0.003 bits about the first glyph of the next, against 0.186 in the text.

The verbose solver, the program that searches for keys in which one letter becomes a string of several glyphs, behaves the same way. It scores 4.59 bits per decoded character on Naibbe-of-Latin, 4.90 on the text and 5.06 on the control, against 2.28 for real text. The key-transfer test separates the text from a fixed-table cipher. A fixed-table cipher passes its best key from one half of the text to the other with no systematic loss, so the gaps scatter around zero and half of them are exactly zero. The text, however, loses 0.02 to 0.07 bits between halves, always in the same direction, and 0.15 to 0.72 between Currier A and B, the two writing styles that Prescott Currier identified. One test was built for this family in particular, homophone exchangeability. Homophones are several signs that stand for the same letter, and the test asks whether frequent word types behave as homophones would, as interchangeable in context. It recognises the Naibbe homophone classes. Naibbe units that stand for the same letter score a median z of 0.17, and 92 percent of such pairs are within noise. A z score is the distance from the chance value, measured in standard deviations. The text's frequent word types are far more distinct in context than Naibbe units or the Markov control, with a median z of 4.8 and only 14 percent of pairs within noise. The one caveat is that part of that distinctness may come from the separate vocabularies of Currier A and B.

Ruled out as published, with high confidence, because the Naibbe texts are eight standard deviations or more from the text on a dozen statistics and all four ciphertexts agree. Unlikely as a family, because the gaps say what a Naibbe-like scheme would have to be to pass and nobody has built or tested such a scheme. It would need units the size of words or morphemes (the smallest meaningful parts of words), each with several spellings, and an open vocabulary needs an open set of such units. A scheme built that way is a homophonic nomenclator, a code list that gives each word several code groups, and it has left the family of letter ciphers. Its spelling choices would have to persist within a page and depend on the neighbouring unit, which is a state-dependent rule. It would also need two sets of tables, one for Currier A and one for Currier B, and tables that change with the position in the line. At Naibbe's rate of 1.5 plaintext letters per cipher word, the whole running text would hold about 53,000 letters, some 9,000 Latin words or twenty printed pages. Because this examination did not search that space, confidence that no such variant exists is moderate to low.

A further test built 18 state-dependent encodings of Latin, a state-dependent cipher being one whose alphabet, table or spelling depends on the position in the text or on an internal state. All of these were deterministic, with no chance in them, which is what sets them apart from the card cipher. Some chose the alphabet by the index of the word or the line, with periods of 2 to 20. Some were autokeys, in which the previous plaintext or cipher letter shifts the alphabet, either running over the whole text or restarting at each word. The rest were verbose codes with rotating tables, homophones chosen by a counter, word-level state machines with 8 or 32 states, and null affixes chosen by a counter. A null affix is a meaningless addition at the start or end of a word, and the null affixes here give every word type 4 or 16 spellings one edit apart. None of the schemes reaches the text's combination of properties, because in these schemes the two halves of that combination trade against each other. Every scheme that removes the long repeats has a conditional entropy of 3.2 bits or more, and every scheme at 2.5 bits or less (these are the fixed verbose codes) keeps 3,616 or more repeated 30-glyph strings. The one-edit family share, the share of word types that are one edit from a commoner type, never exceeds 0.38, even where every type has sixteen spellings. Where a scheme does create information at the word boundary, that information is an artefact of a visible state, and the plaintext's word order stays in the core of the word, whereas the text shows the reverse. The substitution solver flagged the two-alphabet scheme, and half of its words came out in clear. But it scored the autokey and state-machine ciphers exactly as it scored their Markov imitations, so the solver protocol, the set of steps the solvers follow, does not detect ciphers of that kind.

Deterministic state-dependent ciphers are therefore unlikely, with moderate confidence. What counts against them is the two constraints taken together, low glyph predictability and no long repeats, because the repetition argument on its own does not reach them. The confidence is moderate because the test did not exhaust the combinations. The Naibbe card cipher, which is not deterministic, is the one scheme that passes the joint constraint, and the paragraphs above record where it fails.

Repeats that a changing alphabet would leave. A polyalphabetic cipher, which switches between several alphabets as it goes, hides a repeated passage from an exact search, although a search for isomorphs can still find it. Two windows of text are isomorphs when a relettering turns one into the other, and searching for them was the tool of the classical cryptanalysts, as Friedman and Callimahos (1959) describe it. The test here reduced every window of 12 to 30 glyphs in the spaceless stream (the glyphs with the word spaces removed) to its pattern of internal coincidences. That pattern records which positions in the window hold the same glyph, and a one-to-one relettering cannot change it, which is why two relettered copies of a passage share it. The test then counted the patterns shared by windows that differ in glyph content, against ten Markov controls and ten shuffled copies of the text. It did so in the three transcription views, three conventions for which marks count as one sign. A variant of the test tolerated mismatches, to catch a sparse homophonic cipher. No window of 25 or 30 glyphs recurs in the text under any relettering: the count of such pairs is 0 in all three views and in every control. A Latin text of the same size under a line-keyed or Alberti-style polyalphabetic cipher, one that takes a fresh alphabet every 3 to 10 words or one of 20 alphabets chosen by the line, does keep such repeats. It leaves 70 to 236 relettered repeats of 30 glyphs and 200 to 405 of 25. A periodic relettering leaves 21 to 82 exact repeats, and a sparse homophonic substitution leaves 35 near repeats with one to four mismatches. The text does show an excess at 15 to 24 glyphs, with 57 relettered pairs of 20 glyphs against 1 to 6 in the Markov control. That excess, though, consists of runs of one word and its variants set against another such run, qokedy.qokedy.qokedy against olkain.olkain.olkain, and it is gone by 25 glyphs. This extends the repetition argument to polyalphabetic schemes whose alphabet holds still for 30 glyphs, whether the alphabet changes by line, by paragraph or by an Alberti-style change of table, and to periodic schemes too. The extension rests on the same assumption as the fixed-mapping row of the verdict board, that the original repeats phrases. It also stops where the isomorphs stop, since autokey schemes and per-occurrence random spelling leave none, and so the test does not reach them. A second test asked whether the units of the slot grammar (the prefix, core and suffix parts of the word template) behave as homophones would, as interchangeable in context. Of 3,081 prefix, core and suffix pairs, 15 are near the exchangeability floor, the value a fully interchangeable pair would give, against 28 to 34 in the Markov control. The nearest pairs are the k and t alternation and the ch and sh alternation. The limit of this test is that it recovers planted Naibbe homophones only when it is given the cipher's own units, with a separation of 0.93 to 1.00, and it does not recover them through an inferred segmentation.

Earlier work. The literature acknowledges that an encoding whose spelling varies escapes the repetition argument. Bowern and Lindemann (2021) singled out verbose encodings as the family that lowers entropy in the right direction, and Rozanova and Temerev's 2026 preprint likewise leaves respaced verbose encodings unrejected. Greshko's Naibbe cipher (2025, Cryptologia) is the one published scheme of this kind with working code, which is why this examination rebuilt it from that code and put it through every test here. Greshko's published comparison with the manuscript, on symbol frequencies, word length, entropy, bigrams and Zipf's law, is confirmed here. The cipher is then ruled out on the lexicon and the page structure, where its four ciphertexts lie eight standard deviations or more from the text. Two things are new: the eighteen deterministic state-dependent encodings built here, and the homophone-exchangeability test applied to frequent word types. Because the exchangeability test recognises Naibbe's own homophone classes, its negative result on the manuscript has weight. The section stops short of searching the family: the gaps say what a Naibbe-like scheme would have to be to pass, and nobody has built such a scheme.

Scribal spelling, abbreviation and transcription: could they explain the difference?

The reference texts used so far are modern editions with regularised spelling, which is not how a medieval scribe wrote. A scribe abbreviated words and spelled them inconsistently, and in principle that could lower the predictability of a text or break its repeated phrases, so this examination added diplomatic transcriptions, which keep the scribe's own spelling and marks. Thirty of them are early English and Latin prints of 1481 to 1550 on medicine, surgery, distillation, regimen and natural history, taken from the Text Creation Partnership files with the abbreviation marks kept as single symbols. Four are Old Norse manuscripts of about 1250 to 1350 from the Menota archive: a didactic encyclopaedia, the Codex Wormianus, a law code, and a medical fragment too short to use. The third source is the Middle High German reference corpus of 1050 to 1350, which holds medical, herbal and lapidary texts (texts on the virtues of stones), recipes and charms, charters, and a verse chronicle. Each text comes in two or three versions, from normalised spelling to facsimile level with the scribe's own word spacing. To go beyond what any of these transcriptions shows, the tests also simulated scribal abbreviation and spelling variation on the Latin references and on two of the real texts. They respelled 10 and 30 percent of the tokens, the running words. That is more inconsistent than any of the real manuscripts, in which 1.4 to 7.6 percent of tokens depart from their word's usual spelling. In all, 92 samples of the medieval texts and 40 simulated ones were profiled at the size of the running text. The medieval samples are three windows from each of some thirty versions of eleven texts, and 68 of them are diplomatic.

Every step towards faithfulness moves the conditional entropy (how many bits of surprise the next sign carries once the previous one is known) away from the text. Keeping the abbreviation marks, collapsing the expansions, facsimile-level letter forms, manuscript word spacing and simulated variation all raise it, by 0.02 to 0.19 bits. No faithfully transcribed sample comes within 0.85 bits of the text: the samples score 2.96 to 3.58 against 2.11, and the closest is a formulaic recipe book of 1528. The one-edit family share, the share of word types that are one edit from a commoner type, stays far above every sample, at 0.757 against at most 0.55. That highest value comes from a law code with a small inflected vocabulary, and simulated variation adds at most 0.07. Adjacent identical words stay far above every sample too, at 0.94 percent against at most 0.2. Phrase repetition survives every real transcription, each of which keeps 89 or more repeated four-word sequences, 9 or more repeated 30-glyph strings, and a longest repeat of 35 glyphs or more, although it does not survive every simulation. Respelling 30 percent of the tokens, which is four to ten times the inconsistency of the real manuscripts, drives the 30-glyph count of two real texts to zero and the four-word count of Caesar to 6 or 7. Respelling 10 percent alone does not: the 30-glyph count stays at 10 or more and the four-word count at 31 or more. Respelling 10 percent with half the words abbreviated leaves Caesar 8 to 12 repeated four-word sequences and no repeated 30-glyph string. The text's abundance of short repeats also tells against the scribal-noise reading, because noise thins repeats at every length, whereas the text has 5,596 repeated 12-glyph strings, more than any heavily respelled simulation, and no long ones. Manuscript word spacing turns out to be the largest perturbation of any word-level statistic, but it leaves the space-free substring counts unchanged, which is what those counts were designed for. Two statistics are not anomalous in this comparison. One is the fourth-order entropy, the conditional entropy taken with more of the preceding signs known, which a formulaic recipe book takes lower than the text's value, and no verdict here rests on it. The other is the word-order information, which heavily varied language brings down to the text's value.

The verdicts from entropy and from one-edit neighbours therefore stand with high confidence against faithfully transcribed medieval text. The verdict from phrase repetition stands against every real transcription obtained, since only a scribe four to ten times more inconsistent than any measured could reach it. The coverage has one limit: the English and Latin material is early print, the manuscripts are Norse and German, and no transcription of a fifteenth-century Latin or Romance manuscript was obtained.

Three further questions were put to the repetition argument. The first is how much spelling variation removes a language text's repeated phrases, and the natural yardstick is the disagreement between the two transcriptions of the manuscript, which is 2 characters per 100. Random substitution of confusable letters at that rate leaves 34 to 60 percent of the repeated four-word sequences of Latin, Italian and Middle High German in place, and 34 to 53 percent of their repeated 20-letter strings. Removing the repeats takes 10 to 30 percent substitution, which is five to fifteen times the disagreement rate, or else two to four spellings for most or all word types plus 5 to 10 percent substitution. At those levels, though, the conditional entropy rises to 3.5 to 3.7 bits, the hapax share (the share of words used once) rises to 0.78 to 0.85, and the one-edit family share falls to 0.36 to 0.47. So no level of variation moves a real text towards the Voynich profile on every statistic together, a conclusion drawn from 65 noised texts in all. The second question is the reverse, how fragile the text's own count is. The test merged the disputed glyph distinctions one group after another: first the stroke counts, then a/o, r/s and ch/sh, then k/t and o/y. The merges raise the repeated four-word count from 1 to 8, 22 and 45, where chance gives 0.7 to 1.0. Language samples counted the same way have 38 to 151 unmerged and 38 to 67 after the same merges, and at matched vocabulary size the text still has 1.8 to 3.0 times fewer. The long-string counts hardly move: under any merge the text has at most 25 repeated 25-glyph strings and 1 repeated 30-glyph string. Latin has 173 and 46 under the same merges, and Italian 33 and 28. The third question is what the 30-glyph threshold means under a verbose code. At expansion factors of 1.5, 2, 2.4 and 3, thirty glyphs stand for 20, 15, 12 and 10 plaintext letters. Plaintexts of the size that fits into the text's 173,498 glyphs have 39 to 2,578 repeated strings of those lengths. The verbose ciphertexts made from them have 53 to 6,291 repeated 30-symbol strings, 131 to 11,385 at 25 and 565 to 19,696 at 20, where the text has 0, 1 and 73. Allowing for the expansion therefore makes the deficit larger, because a verbose code stretches every repeated passage.

This is why the report quotes the four-word count with its convention, 1 unmerged and 45 with every disputed distinction merged, against 38 to 151 and 38 to 67 for the language samples treated the same way. It is also why the report leans on the two repetition statistics together. The four-word count is sensitive to transcription conventions and stable across samples, the 30-glyph count is the reverse, and so each covers the other's weakness. The argument survives the plausible range of scribal and transcription variation, and it fails only under variation five to fifteen times what was measured.

The tests also simulated abbreviation at the density of a fifteenth-century Latin hand, because the diplomatic prints and manuscripts above have fewer marks than a heavily abbreviated codex. Medieval abbreviation worked by suspension, which cuts the end off a word and marks the cut, and by contraction, which drops letters from the middle of a word and marks the gap. The simulation applied both consistently, at 177 to 229 marks per 1,000 letters, which is the density of Cappelli's repertoire, the standard dictionary of medieval Latin abbreviations, applied wherever it can be. The abbreviation lowers the mean word length of Latin to 4.2 to 5.0 symbols. But it leaves the variance of word length, how spread out the lengths are, at 6.9 to 7.9, against 3.2 for the text. It raises the skew, the lopsidedness of the length distribution, to 0.7 to 1.4, and it keeps the conditional entropy at 3.3 to 3.6 bits. Dense abbreviation, then, shortens words without making them uniform.

Earlier work. The idea that scribal abbreviation and unstable spelling might explain the entropy has been raised repeatedly, and Lindemann and Bowern tested it. Their 2021 preprint included eighteen transcribed historical texts and examined the effects of abbreviation and transcription convention, and Bowern and Lindemann (2021) tested it again in the Annual Review of Linguistics. The finding here, that every step towards faithfulness moves a reference text away from the manuscript, agrees with theirs and extends their result to 92 samples of medieval material and 40 simulated ones. The latest work on the reliability of the word spaces is Rozanova and Temerev's 2026 preprint. Their internality index measures how much of the glyph-to-glyph association found inside words survives across a space that the transcribers marked as uncertain. The preprint leaves the index open to two readings, and this examination reproduces their result under the ratio reading and does not reproduce it under the literal one. Their finding from the glyph coordinates, that uncertain spaces are physically narrower, holds here in direction and falls short in magnitude. The area under the curve, which measures how well the width of a gap sorts uncertain spaces from certain ones, is 0.68 here against their 0.905. Their quire bootstrap redraws the quires (the gatherings of leaves) at random many times to see how much a statistic varies. It is reproduced here as a method and shown not to apply to repetition counts, and quire-level intervals for the headline statistics are reported here for the first time. Unsupervised segmentation of the spaceless glyph stream, which asks whether the transcribed spaces are the ones the glyph statistics support, is also applied to this text for the first time. It finds the spaces as recoverable as those of Latin, Italian or Hebrew, and finds the uncertain ones supported at about four fifths the rate of the certain ones. On abbreviation in particular, Bowern and Lindemann (2021) measured an abbreviated Secreta at 3.4 bits against 3.2 for its plain text. Lindemann and Bowern's preprint set parallel diplomatic and normalised versions of the Codex Wormianus, the Necrologium Lundense and Napier's casebooks (a private early-modern hand) beside the text. The Paston letters below put a fifteenth-century family's hand in the same role. Edwards (2025b) abbreviated Latin by rule and reported the mean word length, and the abbreviation simulation here extends that to density, variance, skew and repetition.

Other kinds of text: human gibberish, glossolalia, a genuine ciphertext and symbol systems

The literature has proposed four reference classes as the text's nearest relatives, and this examination scored all four on the same statistics as everything else. Each turns out to match the text at one level and to fail it at the other. Three further artefacts whose status is known were then scored the same way, to calibrate what those kinds of production look like. They are a text generated by a table rule, an invented vocabulary of the twelfth century, and a set of cover words that hid a message.

Human-produced gibberish. Gaskell and Bowern (2022) collected 38 meaningless texts written by volunteers and trained a classifier, a program that sorts texts into classes, on 42 statistics to tell those texts from meaningful ones. The classifier calls the Voynich text gibberish, and they concluded that the text is statistically similar to human meaningless writing. Their result is reproduced here on their own table, on their files recomputed here, and on the default transcription. The classifier's probability of "meaningful" is 0.35 to 0.40 for the Voynich text, against 0.83 to 0.98 for meaningful texts. But the same classifier calls three scrambled or imitation texts gibberish as well. The within-word shuffle, the text with its glyphs shuffled inside each word, scores 0.39, and the word-order shuffle, the text with its words shuffled, scores 0.46. The Markov imitation, a meaningless text generated from the glyph statistics of the Voynich text, scores 0.49. So the classifier's verdict means "not a meaningful text of the kinds it was trained on". Features of the lexicon and of line position drive that verdict, and the classifier cannot tell the text from a scrambled copy of the text. The full battery of tests was then run at the 200-word scale of their samples, and it shows where the resemblance lies. Human gibberish reproduces the text's word-level statistics, which are word length, Zipf slope, hapax share, adjacent-word information, adjacent identical words and the letter distributions at line edges, and it reproduces none of the glyph-level ones. No gibberish text reaches the Voynich glyph predictability: the gibberish scores 2.41 to 4.00 bits against 1.91 at that scale, where the Markov imitation scores 1.95. None reaches the end-anchored positional information, which is how much a glyph's identity depends on its position counted from the end of the word, and there the gibberish reaches at most 0.60 against 0.85. None reaches the share of one-edit neighbours either, at most 0.25 against 0.37. Pooled, the gibberish scores 3.96 bits, higher than every language and near the within-word shuffle at 3.57, and it ranks last of 37 corpora by overall distance. People inventing nonsense, then, produce letter strings with high entropy and no rigid word template, so the resemblance to the text is at the level of the lexicon.

Enochian. Boxer (2022) compared the text with Enochian, the "angelic" language of John Dee and Edward Kelley recorded in Sloane MS 3188, and found that Enochian's vocabulary is almost all words used once: 86% of its 2,716 types (distinct words). The running share of once-used words never falls below half in Enochian, whereas the Voynich share falls as a language's does, and this examination reproduces that exactly. Enochian has 2,715 types, 86.0% of them used once, and its running share falls from 0.74 to 0.57. The Voynich curve falls from 0.50 to 0.25 at 4,000 words, inside the band of 29 languages at every point. Two things are added here. First, Enochian's curve coincides with the curve of its own word-order shuffle, so it has no lexical drift, no change of vocabulary as the text goes on. The Voynich curve, by contrast, runs 0.01 to 0.03 below its shuffle and 1.1 to 1.4 standard deviations below a random page order, which is the signature of vocabularies that belong to particular pages. Second, at matched size Enochian is unlike the text on every glyph-level statistic, with 3.25 bits against 1.8 to 2.2 and one-edit families at 0.33 against 0.62 to 0.72, and in this it resembles the within-word shuffle. Boxer's second point was that both texts repeat a word immediately, but that fails for Enochian in this transcription. Enochian has 6 immediate repeats against 5.1 expected, where the text has 2.4 times chance and eleven triple repeats. Kelley's language, then, is close to random letter strings with a language-like word-length distribution, which the Voynich text is not.

A genuine period ciphertext. Nobody had compared the text statistically with a real historical ciphertext, so this examination used the Copiale cipher, an eighteenth-century homophonic cipher with a nomenclator, which Knight, Megyesi and Schaefer (2011) deciphered. A homophonic cipher has several symbols for each letter, and a nomenclator is a list of code groups for whole words. This examination scored it as a symbol stream at matched size next to the Voynich glyph stream without spaces. Its conditional entropy is 4.81 bits, which is the value of the synthetic homophonic control, 4.75. That is 2.4 bits above the Voynich stream at the same size, which scores 2.36, and 1.4 bits above its own German plaintext, which scores 3.38. It has no repeated 20- or 30-symbol strings, where its plaintext has 405 and 110. So a homophonic cipher does destroy long repeats, as the Voynich requires, but it does so at the price of a high entropy, which the Voynich does not have. Wallis's nomenclator, the other specimen available, is too short for the entropy statistics, at 1,858 groups. One complete specimen cannot characterise the class, but this one lies on the side of the synthetic ciphers, away from the text.

Non-linguistic symbol systems. Sproat (2014) assembled corpora of symbol systems that are not writing, among them Pictish stones, Mesopotamian deity symbols, Pennsylvania barn stars and weather icons, and showed that simple statistics do not separate them from writing. This examination reproduces that on his corpora with the six languages used here. The symbol systems have a conditional entropy of 1.33 to 4.04, around the language span of 2.95 to 3.63. Their positional rigidity, which measures how firmly each sign is tied to the start, middle or end of a unit, is 0.026 to 0.306, around the languages' 0.105 to 0.157. Sproat's four statistics make a scorecard, and where the Voynich text falls on it depends on which level is taken as the sign. With the glyph as the sign and the word as the text, the Voynich text is extreme with respect to the languages on three of the four axes and inside the non-linguistic spread on all of them. Its positional rigidity is 0.249, against 0.105 to 0.157 for the languages, and only a corpus of emoticons, at 0.306, is more rigid. Its conditional entropy is 2.07, near Pictish at 2.08 and weather icons at 2.24, and its Zipf slope is −2.33, against −1.02 to −1.48 for the languages. Its top-three successor coverage, the share of what follows a sign that its three commonest successors account for, is 0.81 against 0.50 to 0.63. The within-word shuffle removes all of that, and the Markov imitation reproduces it. So at this level the scorecard is measuring the slot template, and the template scores like an emblem system, which letters in words do not. With the word as the sign and the line as the text, the text looks like the languages on the lexicon axes. Its hapax share is 0.71 against 0.48 to 0.76, and its Zipf slope −0.76 against −0.51 to −1.04. On sequence structure, though, it is at the bottom: its word order within lines is barely distinguishable from its own shuffle, with a gap of 0.03 bits over all 4,129 lines. That is the scorecard form of the word-order result above.

Reference classProposed byMatches the text atFails at
Human-produced gibberish (38 texts)Gaskell and Bowern (2022)the lexicon: word length, Zipf slope, hapax share, adjacent-word informationthe glyph level: predictability 2.41 to 4.00 bits against 1.91, positional information, one-edit families
Enochian (Sloane MS 3188)Boxer (2022)nothing beyond a language-like word-length distribution, and its hapax curve is the one the text does not followthe glyph level (3.25 bits) and the lexicon (86% hapax, no drift), like the within-word shuffle
Copiale ciphertext (homophonic, with a nomenclator)Knight, Megyesi and Schaefer (2011), as a specimenthe absence of long repeatsthe glyph level: 4.81 bits against 2.36 at matched size
Non-linguistic symbol systems (eight corpora)Sproat (2014)the glyph-in-word level: the rigidity, predictability and successor coverage of an emblem systemthe word-in-line level, where the text has a language-like lexicon and less order than any system or language
Book of Soyga tables (generated by the rule Reeds recovered)Reeds (2006), as a specimen of table-generated textnothing: it is the mirror image of the textthe glyph level (4.3 to 4.4 bits against 2.3, the least predictable stream in the set) and the repeats (353 repeated 30-letter strings against none)
Lingua Ignota (Hildegard of Bingen's invented vocabulary, 760 entries)Steinmeyer and Sievers (1895), as a specimen of a constructed vocabularythe ending system (five or six final letters cover 90 percent of words) and the first-glyph test, which it passes as the text doesthe glyph level (3.0 bits) and the one-edit density (18 to 23 percent of words a single letter from another, against 50 to 87)
Steganographia conjurations (Trithemius, 1,216 cover words)Reeds (1998), as a specimen of invented words carrying a hidden messagenothing at the glyph level, though the channel test catches its concealment rule when that rule is planted herethe glyph level (3.1 bits, one-edit share 0.13, positional information 0.33) and the lexicon (once-used words throughout)

A text known to be generated by a rule. The Book of Soyga is a sixteenth-century treatise with thirty-six tables of letters that John Dee could not read, and Reeds (2006) showed that a recurrence generates the tables. Each cell is the sum, modulo 23, of the letter above it and a fixed value of the letter to its left, from a six-letter code word written down the margin. Daruka (2021) compared the manuscript with the tables of Dee's Liber Loagaeth and reported statistical matches between the two. The Soyga tables give the sharper comparison, because their generating rule is known, so this examination regenerated them from that rule, which gives a meaningless text of known algorithmic origin. Its glyph stream is the least predictable in the reference set, at 4.3 to 4.4 bits at matched size. The manuscript scores 2.3, the languages 3.4 to 3.8 and the within-word shuffle 3.6. At the same time it is the most repetitive, with 353 repeated 30-letter strings and a 107-letter repeat, where the manuscript has none and the languages have up to twelve. Cut into words of the manuscript's lengths, it has no word structure: its one-edit share is 0.17 against 0.73, and its end-anchored positional information 0.001 against 0.71. A deterministic recurrence, then, is regular at long range and disordered at short range, the mirror image of the manuscript. So low glyph predictability is not something generated text has as such. It is a property of the particular procedure that the generative-model section below fits.

An invented vocabulary of the period. Hildegard of Bingen's Lingua Ignota is a glossary of about a thousand invented nouns, ordered by semantic class (God, angels, people, parts of the body, plants, and so on). It is the one constructed vocabulary of the Middle Ages whose meaning is known. This examination took 760 of its entries in manuscript order from the edition of Steinmeyer and Sievers (1895). It scored them as a list of types against Voynich, Latin, Italian, Hebrew and Enochian type lists of the same size. The glossary shares two of the manuscript's word-level properties. One is a narrow ending system: five or six final letters cover 90 percent of its words, and its last-letter entropy is 2.4 to 2.5 bits, against 2.5 to 2.8 for the manuscript. The other is a first letter that tells nothing about later letters beyond its neighbour. The measure for that is held-out gain, how much the first letter improves the prediction of the later letters on words the test has not seen. The glossary's held-out gain is at or below zero, as in the manuscript, against 0.21 for frequent Latin types. What the glossary does not share are the two properties this examination treats as distinctive. Its letter stream is nearly as unpredictable as a language's, at 3.0 bits against 2.1 to 2.5 for the manuscript's types and 3.1 to 3.6 for the languages. And its words rarely differ from one another by a single letter: 18 to 23 percent do, against 50 to 87 percent for the lists of word types it was scored beside. The manuscript's 0.74 to 0.76 given in the data section is for a sample of its own window size. Its semantic order leaves a small trace in word form, since adjacent entries are 3.5 percent more alike than random pairs, and compounds such as nilz-peueriz, stepfather, account for that. The trace is the size of the one found for the manuscript's word types in order of first appearance. The first-glyph test therefore does not separate the manuscript from an invented vocabulary. The glyph predictability and the one-edit density do separate them, and the constructed-language row of the verdict board leans on those.

Invented words that carried a message. Books I and II of Trithemius's Steganographia (1499, printed 1606) contain conjurations, which are lists of invented spirit names. Those names hid instructions, as the printed key of 1606 explains and as Reeds (1998) recounts. This examination extracted the conjurations from a digital edition, 1,216 words, and scored them at their own size. They read as a language on every glyph statistic, scoring 3.1 bits against 1.8 to 2.2 for the manuscript at that size. Their one-edit share is 0.13 against 0.52 to 0.68, their positional information 0.33 against 0.69 to 0.95, and their vocabulary is made of once-used words, with no phrase repetition. The concealment rule described for them puts the message in the alternate letters of alternate words. The tests planted a Latin text by that rule into two carriers of the manuscript's length, one made of invented cover words and the other the manuscript. The channel test of the hidden-message section below looks at the letters a concealment rule would use, called the channel. It reports their entropy as a share of the shuffled entropy, so a hidden message lowers the value. It registers the planted message at 0.66 against the baseline of 0.89 that the Markov imitation sets for this channel, and it registers a message of a thousand channel letters at 0.79. The manuscript's own alternate-letter channels score 0.82 to 0.89 in all four phases, which is the value of its meaningless imitation. The conjurations show no text under this rule at their size, scoring 0.91 to 0.97, and their chapter keys could not be rebuilt from the sources at hand. Three further artefacts of known status could not be tested, because no machine-readable text of them exists. The Codex Seraphinianus is in copyright and has never been transcribed, and the Rohonc Codex has a transcription from 2018 that is not distributed. Nor could any corpus of transcribed glossolalia, the fluent speech in made-up syllables that some people produce in religious ecstasy, be found.

Lists, glossaries and rolls. Every repetition argument above leaves one escape open: a plaintext that never repeats a phrase because it is a list. So this examination measured eleven texts of the period with one entry per line, at the text's size and with the same code as the language samples. Two are English-Latin glossaries, the Promptorium Parvulorum and the Catholicon Anglicum. The others are the Corpus Glossary, the Sinonoma Bartholomei (a Latin plant synonymy), two martyrologies, the ten litanies counted as one text, the York freemen rolls of 1272 to 1759 and the Durham Liber Vitae. The last is Pliny's Natural History, with the duplicated blocks of the available file removed. Every bare list measured at 34,780 tokens repeats at least 24 four-word sequences. The York roll repeats 24 to 40, the glossaries 30 to 110, the later roll 386 to 1,294, the martyrologies 25 to 342 and the litanies 139 to 260. Pliny repeats 10 to 90 over ten windows, where the text has 1. One glossary, the Corpus Glossary, is too short for that size. At its own 5,695 tokens it repeats no four-word sequence, but the text cut to the same 5,695 tokens repeats none either. The glossary reaches that zero by having no frequent words: its ten commonest words make up 6.5 percent of its tokens, against 21 percent for the text's at that size. And 91 percent of its types occur once, against 69 for the text. The fourteen-statistic distance is a single figure that adds up the gaps on fourteen statistics. On it, the York roll (1.79) and Pliny (1.91) come closer to the text than every language sample but Italian, which scores 1.77. What keeps them apart is the glyph level and the phrase repetition they keep. Their glyph predictability is 3.32 to 3.36 bits against 2.11, and their one-edit density 0.28 to 0.40 against 0.76. A list of the period, then, either repeats its phrases or lacks the frequent-word head that the text has, and no sampled genre does what the text does. The same holds of a text written by non-professional hands, since the Paston letters, fifteenth-century family correspondence, keep 763 to 1,053 repeated four-word sequences and a predictability of 3.35 bits.

Earlier work. Gaskell and Bowern (2022) supplied the human reference class when they collected gibberish from volunteers and trained a classifier that calls all five Voynich samples meaningless. Their classification is reproduced here on their own data and on the manuscript, and two findings qualify it. The same classifier calls a scrambled Voynich text meaningless. And human gibberish matches the manuscript's word-level statistics and none of its glyph-level ones, so volunteers invent letter strings and do not invent a rigid word template. Boxer (2022) set the manuscript beside the Enochian of Sloane MS 3188, and his hapax fingerprint is reproduced exactly. His observation that Enochian shares the manuscript's immediate repetitions is not reproduced, because in the transcription used here Enochian repeats a word immediately no more often than chance. The Copiale cipher of Knight, Megyesi and Schaefer (2011) is used as a genuine period ciphertext. It shows that homophonic substitution raises glyph entropy by 1.4 to 2.2 bits, which is the wrong direction. Hermes (2022) analysed the Polygraphia III construction, Trithemius's cipher that turns each letter into an invented word. That construction is rebuilt as a reference text and ruled out on phrase repetition. The non-linguistic symbol corpora of Sproat (2014) are applied to this text for the first time. On Sproat's own scorecard the manuscript's glyph level falls among the symbol systems and its word level among the languages. The list-genre explanation of the missing phrases was raised on Zandbergen's site. The Voynich Project (2026), a self-published site, tested it with other statistics on one recipe compendium and one Vulgate register. The eleven lists above test it on four-word repeats at matched size.