Beinecke MS 408 · independent examination
Voynich Text Examination

What the text of the Voynich manuscript is like, measured on the page scans and on full transcriptions, and compared with real writing in thirty-one languages.

Does the text carry content?

Does the text carry content? Whether the text is about something stays undecided. A text about something leaves two traces. Its words go with its pictures. Its commonest words, the grammar words, fall evenly across the pages. In the manuscript, pages with similar pictures do not use similar words. And the commonest words crowd onto some pages and stay off others. That counts against prose about a subject, not against a catalogue of names used once. A catalogue has no sentences and no grammar words. Words in the text do tend to be used again on their page. But meaningless words drawn from a slowly changing pool do that too.

Whatever produced the text, a text about something leaves traces. The first section below looks for five of them in the running text and beside the pictures, and the second asks what modern statistical learners find in the text beyond a few glyphs of context. Neither approach assumes any reading.

Traces of subject matter: pictures, grammar, labels and recurrence

A text about something leaves traces in where its words fall, because the words follow the subject. Words recur on the pages that concern them, pages about the same subject share words, and words stand next to the pictures of what they name. A small closed class of grammatical words, the function words such as and and of in English, behaves differently from the rest, because it is used whatever the subject is. And names recur across the diagrams that use them. This examination measured five such signatures on the running text. To know what each looks like with no subject behind it, it measured them as well on the same page skeleton, the manuscript's own layout of pages and lines, filled with the three meaningless generators (programs that write text by a rule) used throughout this report. It measured them on Latin, Italian, German and English too, which show what a text with a subject gives. The first generator is a glyph-level Markov imitation, meaningless text generated from the text's own glyph statistics, like the one that controls the code-breaking solvers above. The other two are the local-copying and page-level vocabulary generators described under generative models. Every comparison held the scribal hand and the Currier regime fixed, the regime being Currier's writing style A or B, and that way neither hand nor regime could produce the result.

The result has two halves. The referential signatures, the ones that tie words to things, are absent. Image descriptors of 119 herbal pages (pigment shares, plant shape, roots, flowers) do not correlate with page vocabulary. The test is a Mantel test, which correlates two distance matrices, here the distances between pages by picture and by vocabulary. It gives r = −0.02, p = 0.66, where a planted picture-linked vocabulary is detected at r = +0.26. There is no sign of a class of grammatical words either. The thirty commonest words are as bursty across pages as mid-frequency words, 1.43 against 1.65, a bursty word being one that crowds onto some pages and stays off others. In every language, by contrast, the function words are flat: Latin 1.08 against 1.49, Italian 1.02 against 1.21, German 1.08 against 1.33. And zodiac label types recur across pages exactly at the chance rate, 37 observed against 38.4 expected, so the labels show no re-used names.

The other half of the result is that the text is not without memory. Word types, the distinct words, recur above chance at every range up to the leaf. After removing what hand and regime explain, a word comes back at 2.9 times the random rate in the same or the adjacent line. It comes back at 2.2 times the random rate two to three lines away, 1.7 times elsewhere on the page, and 1.5 times on the other side of the same leaf. Within the quire, the gathering of leaves sewn together, the rate is 1.05, which is close to the random rate. That is the profile of Latin (2.2, 2.8, 1.9, 1.4, 1.0) and, beyond the adjacent line, of German (1.0, 2.3, 1.8, 1.5, 1.0), which has no excess in the same or the adjacent line. Of the four real texts poured into the same pages, two match the manuscript on the other side of the leaf, the Latin at 1.4 and the German at 1.5, and the Italian sits near chance there, at 1.08. For the English no leaf figure is recorded, and in the same or the adjacent line it comes back at six to nine times the random rate, which the formulaic repetition of the King James Bible accounts for. None of the generators reproduces the profile. The local-copying generator has nothing beyond three lines (1.00 on the page, 1.03 on the leaf), the page-vocabulary generator has nothing beyond the page (1.04 on the leaf), and the Markov imitation has nothing at all. The illustration sections also have their own vocabularies beyond scribe and regime. The test here is PERMANOVA, a permutation test of vocabulary profiles, which asks whether the sections differ more than a random regrouping of the pages would. It gives F = 2.91 against a null of 2.05 from regrouping within hand and regime, p = 0.001. The controls score at their nulls, while contiguous natural prose gives 1.05 to 1.2, and the English Bible 2.45.

How far away a word comes back, against chance

For each distance, the number of pairs of the same word, divided by the number expected if the words were shuffled within each scribe and Currier language. The text is compared with Latin poured into the same page layout, with a generator that copies nearby words, and with the drifting-state generator.

1× 2× 5× 10× 20× chance recurrence, times chance (log scale) same or adjacent line 2 to 3 lines away same page, farther other side of the leaf same quire farther away Voynich text Latin (Caesar) in the same pages Drifting-state generator (fitted to this profile) Copying generator (Timm-type, not fitted) Voynich text 2.91 2.22 1.71 1.51 1.05 0.96 Latin 2.24 2.83 1.94 1.38 1.01 0.98 Drifting state 2.78 1.89 1.66 1.35 0.97 0.99 Copying 20.31 3.00 0.93 0.97 0.92 0.99
With scribe and Currier language held fixed, a word of the text comes back 2.9 times more often than chance in the neighbouring lines and 2.2 times more often two to three lines away. Elsewhere on the page it comes back 1.7 times more often than chance, and on the other side of the leaf 1.5 times. Within the quire and beyond, it comes back at the chance rate. Latin poured into the same page layout gives a profile like it. The copying generator shows nothing beyond three lines. The drifting-state generator, whose parameters were fitted to this profile, reproduces it. The count uses words that occur 2 to 50 times, and the generator's values are the mean of 5 runs.
SignatureVoynich textMeaningless generators (Markov / local copy / page vocabulary)Latin / Italian / GermanResult
Recurrence excess, same page ≥ 4 lines away / other side of the leaf (within hand × regime)1.71 / 1.511.01 / 0.97 · 0.93 / 0.97 · 1.38 / 1.011.94 / 1.38 · 1.29 / 1.08 · 1.80 / 1.47present, language-sized
Section vocabulary beyond hand × regime (PERMANOVA F against its null)2.91 vs 2.05, p = 0.0012.01 vs 2.02 · 1.01 · 1.001.10 · 1.05 · 1.21 (English 2.45)present
Picture–word association on herbal pages (Mantel r; planted control +0.26)−0.017, p = 0.66−0.04 / −0.05 / +0.02−0.05 / −0.03 / −0.01absent
Function-word class: page burstiness of the top 30 vs ranks 31–3001.43 vs 1.651.01 vs 0.98 · 1.23 vs 1.49 · 1.48 vs 1.441.08 vs 1.49 · 1.02 vs 1.21 · 1.08 vs 1.33absent
Zodiac label types recurring on two or more pages37 (chance 38.4)––no re-used names

Topics against the drifting generator. A text about things has topics that persist through a paragraph and return across the book, and a codebook of such a text, which swaps each word for a code word, keeps them. A topic model describes each paragraph as a mixture of a few topics, where a topic is a set of words that tend to occur together. This examination fitted paragraph topic models of the standard kind (latent Dirichlet allocation, Blei, Ng and Jordan 2003). It scored them on the second half of each paragraph, predicted from its first half, in three strata that hold hand and regime fixed. The baseline was a cache of the paragraph's own words, which expects the words of the first half to come back, and the topics add nothing beyond it: −0.13 to −0.03 bits per token in the three strata. That is within 1.6 standard deviations of the runs of the drifting-state generator, the meaningless generator whose working vocabulary drifts slowly. In the same way, the share of a word's recurrences that fall more than seven pages away is at the level of a page permutation in every stratum, as it is for the generator. The tests can catch a text with topics, since a codebook of English narrative poured into the same pages is caught by them, but only through that recurrence share. It is caught in the Currier B strata, at 2.3 to 5.5 standard deviations below the permutation level, and the topic gain does not catch it. So the null rules out topics of that strength in those strata, and nothing weaker.

Ruled out: meaningless text with no memory beyond a few lines, with high confidence, because the recurrence profile reaches the leaf and no generator without memory reproduces it. Undecided: meaningful text, which is not ruled out, although no signature that ties words to things supports it. The most economical reading of all five tests is a production process whose working vocabulary drifts slowly, that is, a scribe working in batches. That reading fits because it predicts coherence at the scale of the page and the leaf, together with no picture link, no class of grammatical words, and labels that are each used once. Resets of the vocabulary alone, however, do not reproduce what gives each section its own vocabulary. The confidence in this reading is medium, and three findings would change it. One is a picture–word association above r = 0.1 with better botanical descriptors, and another is a shared-topic gain inside a single hand and regime that a matched natural-language control does not show. The third is labels that track a shared referent, one thing named in more than one place.

The published measures of content, re-run. Montemurro and Zanette (2013) measured how much the identity of a word tells about which block of the text it falls in, as the block size varies. They found that the information peaks at a block of about 800 words, at a height a little above English, and read this as evidence of content organised by topic. This examination reproduces the curve, which rises to 0.305 bits per word at 512 words and 0.306 at 1,024, and falls beyond that. A block bootstrap, which re-samples the text in blocks to see how stable the peak is, puts the peak at 512 or 256. The height ranks eighth of 29 texts, next to the King James Bible (0.326, with Caesar at 0.133 and Dante at 0.055). The new finding is where the information lives. A shuffle that keeps hand and regime fixed keeps 0.148 of the 0.306, so half of the peak is the block structure of hands and regimes. A shuffle within pages keeps 0.303 and a shuffle within paragraphs 0.306, so nothing lives below the page, and the values at small block sizes are the page-scale clustering leaking into small blocks. Inside Currier A alone the peak is 0.170 and inside B 0.230, still inside the language range. The controls show what produces such a curve. The generators without page structure give at most 0.03, the local-copying generator gives 0.16 (peak 128), and the Latin codebook gives 0.13, which is the narrative's own topics passing through. But the drifting-state generator gives 0.36, which it was fitted to give, so the statistic does not separate content from a drifting working vocabulary, since both produce it.

The same authors ranked keywords, the words most concentrated in particular blocks, and the same ranking here makes 95 of the 100 commonest words significant, against their 30. That count means less than it seems, because the same pipeline gives 25 to 98 for the languages: Virgil 25, Dante 30, Homer 41, Caesar 46, the King James Bible 85, the Hebrew Bible 88, and a Turkish scripture 98. It gives 8 to 10 for structure-free generators, where 5 are expected by chance, and 98 for the drifting-state generator. The published number is what a narrative gives under a stricter criterion. The top of the list is the set of regime-marker words, the words that mark one Currier regime (shedy, qokain, qokeedy, daiin, qokedy, chedy). Inside a single regime the count falls, to 85 inside Currier B alone and to 58 inside A. The keywords are modestly over-represented in the biological section (21.6% of their tokens against 17.9% of all tokens) and under-represented in the herbal section (30.7 against 32.9). That is a difference of three to four points, small but not produced by a null that holds hand and regime fixed.

de Arruda, Marinho, Costa and Amancio (2018) built a network of paragraphs, joining paragraphs with similar vocabularies, and found it far more modular than a paragraph-shuffled text, with two dominant groups. Modularity measures how strongly a network splits into groups whose links stay inside the group. This examination reproduces the modularity, 0.462 against 0.244 ± 0.004 for the shuffle, and the excess of 0.218 is inside the range of 28 languages poured into the same paragraphs, 0.022 to 0.547, with seven above it. Its source, though, is the two regimes and the hands. A shuffle within hand-and-regime blocks keeps 0.414, so 78% of the excess is that block structure. The five communities found are hand-and-regime blocks, since only 24 of 665 labelled paragraphs fall in a community dominated by the other regime. The 0.048 left over, the residual, is the paragraph-level counterpart of the page recurrence above. The generators without page structure give no excess, the drifting-state generator gives 0.16, and the test recovers exactly a positive control of Latin and Italian paragraphs alternating in runs.

Reddy and Knight's adjacent-page result is described above. Their index of page topicality has no published definition that could be matched. Two dispersion indices built here measure how unevenly words spread over the pages, and they put the text with Caesar, the Quran, Homer and Persian, below the King James and Hebrew Bibles. Their induction of word classes, which groups words by the company they keep, gives classes that are orthographic, which means they follow word shape and regime. The standard low-data test for syntax is the gain of a class-bigram model over unigrams on held-out text (text the model was not trained on). That gain is how much predicting each word from the class of the word before it beats predicting it from its frequency alone. On this text it is 0.016 ± 0.013 bits per word, which is 0.06 to 0.12 above its shuffle and below every language sample except Latin prose (Italian 0.17 to 0.33, English 0.76 to 1.36). Coarsening words to 10 to 50 classes, which is how sparse syntax would appear, does not make the order signal grow.

Earlier work. Two published results anchor this section, and this examination reproduces both. Montemurro and Zanette (2013) measured how much a word tells about which block of the book it falls in, and found a peak near 800 words at a height just above English. The curve here peaks between 512 and 1,024 words, at a height in the upper third of the language range. This examination adds that the clustering lives entirely at the page scale and above, and that a drifting generator with no content produces the same curve. Reddy and Knight (2011) measured how often a page's most similar page is the page beside it, at 15.6 percent over all pages; the figure here is 17.9 percent, and the frequent words account for it. Montemurro and Zanette's keyword ranking is reproduced as a phenomenon but not as a number, since here 95 of the 100 most frequent types come out significant, against their 30. The count does not separate the text from a drifting generator. Reddy and Knight's word-class induction is reproduced as well, but the induced classes are orthographic, which means they follow spelling, and not syntactic. The paragraph-network modularity that de Arruda, Marinho, Costa and Amancio (2018, a preprint) reported is reproduced too, and most of it turns out to be the block structure of hand and Currier regime. The picture-to-word association test, with automated image descriptors and a planted picture-linked vocabulary as its positive control, had not been applied to this text before. It finds nothing on the herbal pages, where the planted link is found at r = +0.26. Sterneck, Polish and Bowern (2021, a preprint) fitted topic models to the text and judged them by their agreement with sections and hands, without scoring them on held-out text. Held-out gains of a cache model at several scopes, from the line to the bifolium (a sheet folded once, giving two leaves), appear in Yoshida's unreviewed 2026 audit. The paragraph-completion score above is the document-completion evaluation of Wallach, Murray, Salakhutdinov and Mimno (2009). It uses a cache model (Kuhn and De Mori 1990) as the baseline, a generator without topics as the null, and a planted English codebook as the positive control. None of those had been applied to this text before.

Language-like structure through modern methods

This examination measured three signatures that natural language shows to statistical learners, on matched samples of 34,780 words. The first is predictive information, which is how much a model gains from a longer context and from more training data. Held-out cross-entropy is the surprise per symbol of a model scored on text it did not train on, and the order of a model is how many previous symbols it looks at. For languages the held-out cross-entropy keeps falling as the context grows: 0.44 to 0.85 bits gained beyond order 3, with the best order at 6 to 8. It also keeps falling as the training data grows, by 0.11 to 0.30 bits from half to full size. The Voynich glyph stream behaves differently, gaining 0.023 bits beyond order 3 and 0.05 from more data. Its Markov imitations, meaningless texts generated from its own glyph statistics, gain the same (0.00 to 0.05, and 0.06), and its best order is 4 in every fold, that is, in each of the held-out splits. At the word level a bigram model predicts each word from the one before it, and a unigram model uses word frequencies alone. In every language the bigram model gains 0.2 to 1.5 bits over the unigram model, whereas on the Voynich text the gain is −0.05 ± 0.03, against −0.09 to −0.14 for the controls. That is the same 0.1-bit word-order trace measured above, and nothing more.

The second signature is the geometry of embeddings. An embedding gives each word a position in a space where words used in similar contexts lie close together, and nearest-neighbour structure and clustering in that space do not separate the text from anything. The hubness of its embedding space, which is how often a few words are everyone's nearest neighbour, is 3.0, which equals its Markov imitations (3.1 to 3.3) and differs from the languages (0.2 to 1.9). One statistic is language-sized. Word pairs one edit apart are distributional neighbours, words used in similar contexts, far above chance: 8.6% of them are in each other's ten nearest neighbours against 2.0% expected. That holds in split halves, in Currier B alone (one of the two writing styles), and with only line-internal contexts, and while languages give 4.3 to 8.0%, every copying or Markov control is at chance. It is what inflected variants of one lemma, the forms of one dictionary word, do in a language. But it is also what the page-and-line vocabulary generator, which draws the words of each page and line from a subset of the vocabulary, does without meaning or copying, at 5.4 to 13.7% depending on its subset size. So the statistic is the diffuse clustering of word families seen through embeddings, and not a new fact.

The third signature, unsupervised alignment of the Voynich embeddings onto each language, could not be assessed. Alignment tries to match the word space of one text onto that of another without a dictionary. The positive controls fail at this corpus size: Latin onto Latin recovers 0% of a dictionary, and even supervised alignment with the true dictionary recovers at most 3%. A method that cannot align Latin with Latin cannot say anything about the Voynich text.

SignatureVoynich textMarkov and copying controlsLanguagesResult
Cross-entropy gain beyond order 3 (bits per symbol)+0.023 ± 0.0040.000 to +0.050+0.44 to +0.85no evidence
Gain from doubling the training data at order 5+0.051+0.056 to +0.064+0.106 to +0.300no evidence
Recurrent-model gain over the best n-gram, a model that predicts each glyph from a fixed number of previous glyphs (bits per glyph)+0.129 ± 0.017+0.017 to +0.063−0.009 ± 0.007 (Latin, at this size)real, within the word and the few words before it
Word bigram over unigram, held out (bits)−0.048 ± 0.030−0.094 to −0.143+0.21 to +1.490.1-bit trace only
Hubness of the embedding space3.03.1–3.30.2–1.9like the controls
One-edit pairs among ten nearest neighbours (observed / expected)8.6% / 2.0%2.6–4.0% / 2.0–3.5% (copy, Markov); 5.4–13.7% / 1.1–1.7% (page-vocabulary generator)4.3–8.0% / 0.4–3.5%real, reproduced by a template

An n-gram, a model that predicts each symbol from a fixed number of previous symbols, is not the only kind of model, and a different kind sees a little more. A recurrent network reads the text one glyph at a time and keeps a memory state as it goes. The one used here is a two-layer LSTM, a long short-term memory network, trained and scored on exactly the same held-out folds as the n-gram models. It beats the best Kneser-Ney n-gram (an n-gram with a standard smoothing method) on the text by 0.129 ± 0.017 bits per glyph, or 7.5 percent. On the text's second-order Markov imitation it gains 0.035 ± 0.015, on the third-order one 0.017 ± 0.032, and on the copying generator 0.063 ± 0.007, so the text gives it more than its imitations do. On Latin at this sample size, however, it gains nothing (−0.009 ± 0.007), which means the positive control fails, and the gain cannot be read as a mark of language.

What the network learns can be seen from the within-word shuffle. That shuffle has no sequential structure at all, and yet the network gains 0.555 ± 0.018 bits on it. The gain comes from tracking which glyphs of the current word have already been written, a constraint that no n-gram represents. A scrambled-context test locates the rest. In it the network sees either the true preceding text or an unrelated stretch of the same length before the last few glyphs. On the text, the true context beyond the last 32 glyphs is worth 0.089 ± 0.014 bits. That is the same as on the Markov imitation (0.077 ± 0.023), where it can only be the network recognising which section's statistics it is in. The text's excess over its imitation therefore lies within the last 32 glyphs, about six words. It is 0.05 to 0.07 bits from the eight to thirty-two glyphs before the target, at two standard deviations, and 0.075 within the last eight. So the finding that nothing beyond a fourth-order glyph process is learnable holds for n-gram models. A model that keeps state finds about a tenth of a bit per glyph more, all of it within the current word and the few words before it. That is where the near-copies of neighbouring words described above would put it, and the copying generator at its fitted parameters does not reproduce it.

The tests then took that gain apart, to see where it lands and what it draws on. It lands on the interior glyphs of the word (54 percent), the first glyph (20 percent), the second glyph (11 percent), the last glyph (12 percent) and the space (3 percent). It appears at every word length, with 5 to 16 percent of it in words of eight or more glyphs. Of that gain, 77 percent disappears when the words are shuffled before training, so it draws on the preceding words, though not on their identities. A word-level model over word identities gains 0.06 ± 0.12 bits per token on the text, the same as on a Markov-3 control, where a Latin word bigram gains 0.55. The cross-word information comes instead from the neighbouring words' glyphs and from the page's regime. Linear probes, which are simple classifiers that read the network's memory state, decode the page's Currier regime at 0.97 and its hand at 0.91, against a rate of 0.50 for always guessing the commonest answer. They decode the previous word's prefix and suffix only weakly (0.34 against 0.32, and 0.47 against 0.42). None of this is in the generators. The examination also trained a network on the output of the drifting-state generator, the generator whose working vocabulary drifts slowly. That network loses 0.163 bits per glyph on the text and gains only 0.037 on its own output. The network trained on the text, in the other direction, predicts the drift, copying and Markov texts worse than a glyph 4-gram does. Its probes decode language and hand from the generator's text at only 0.66 and 0.57. The ratio of the two networks' surprisal separates real lines from generated lines with an area under the curve of 0.96, on a scale where a half is guessing and one is perfect. Three quarters of the recurrent model's gain, then, about a tenth of a bit per glyph, crosses word boundaries and lies below the level of whole words, and every generator built here lacks it.

Taken together, these tests give no evidence of predictive or semantic structure beyond four things. Those are a fourth-order glyph process, the composition of the current word, the re-use of the few words before it, and the clustering of word families already described. The alignment test is not usable at this size, so it adds nothing either way.

Earlier work. Word embeddings of the Voynich text exist in blog posts and unpublished student work, beginning with Perone (2016), and Sterneck, Polish and Bowern fitted topic models in a 2021 preprint. None of that work used matched meaningless controls, and the controls are where the results here come from. With them, the hubness of the embedding space turns out to equal that of a Markov imitation and to differ from every language. The one language-sized signature, that words one glyph edit apart are distributional neighbours far above chance, is also produced by a page-and-line vocabulary generator with no meaning. So it is the clustering of word families seen through embeddings, and no new fact about the text. The long-context neural comparison, for which no peer-reviewed application to this text was found, is the one test in this section that separates the text from its imitations. It does so by about a tenth of a bit per glyph, located within the word and the few words before it, but because the positive control fails at this sample size, the gain is not a mark of language.