Beinecke MS 408 · independent examination
Voynich Text Examination

What the text of the Voynich manuscript is like, measured on the page scans and on full transcriptions, and compared with real writing in thirty-one languages.

Is there a hidden message?

Is there a hidden message? A message could hide in some feature of the writing, not in its words. The length of every word is one such feature. We wrote a program to search for such messages. We tried it on 29 features of that kind. To test the program, we planted a Latin message in each feature, written into a copy of the text. The program found every message of 3,000 words or more. It found none of a thousand words or fewer. We did not try sizes between. In the manuscript the program found none. So a message of 3,000 words or more is ruled out. A shorter one could be missed.

A text can hide a message without being a cipher, because the message can live in some feature of the visible text. It could be in a subset of the glyphs, in the word lengths, in the shapes of the letters, or in whatever the visible structure leaves unpredicted. Hiding a message in a feature of an innocent-looking text is called steganography, and each such feature is a channel. A search that finds nothing in a channel means little unless the search is known to work, so this examination tested each channel with a planted message as its positive control. A planted message is a known text written into the channel on purpose, and the test has to find it before its silence on the real text counts for anything. When it does find it, that shows the test can detect a channel of the strength it looks for.

A message in a channel of the visible text

Suppose the words were only a carrier and the message lived in some sequence derived from them. That sequence would then show what any message shows: it would be more predictable than the same sequence in random order, and it would repeat. The test therefore read twenty-nine such channels off the running text, each by a fixed rule, and looked in each for those two marks. Some channels take one glyph from each word: its first, second or last glyph, or one of its first two units. A unit is a group of glyphs that behaves as one sign, such as ch or aiin. Some take a count from each word instead: its length in glyphs, its length in units, whether that length is odd or even, and its numbers of i-strokes and e-strokes. Some take one part of each word, its prefix, core or suffix slot, which are the three parts of the word template. Some record a yes or no for each word, according to whether it holds a q, a gallows (one of the tall glyphs) or a bench (one of the shapes written ch and sh). The rest are the sequence of gallows glyphs, every second glyph and every third glyph. The last channels work at the level of the line: the number of words and of glyphs in each line, and the first and last glyph of each line.

Each channel got three scores. The first is held-out cross-entropy: a model learns the channel from most of the text, and the score is how many bits of surprise per symbol it has on the part held back. To give that number a baseline, the test compares it with the score for the same channel shuffled, over the whole text, within each page, and within each line. The second score is the channel's size after compression, and the third is its number of repeated runs of 8 to 16 symbols, since a message repeats where random order does not. The tests took all three scores on five texts. Two are the Voynich text and its within-word shuffle, the text with the glyphs of each word in random order. Two more are meaningless generators (programs that write text by a rule), the Markov imitation, which is generated from the text's own statistics of which glyph follows which, and the copying generator, which makes each new word with reference to nearby words. The fifth is Latin, which shows what a language gives in the same channels. The positive controls were planted messages, for which the examination wrote a Latin text into one channel at a time of the real word sequence. The channels used for this were the first glyphs, the word lengths, the presence of q, the gallows, the words per line and the line-initial glyphs.

The test found every implant. The score reported is the ratio of the held-out entropy to the shuffled entropy. A value near one therefore means the channel is no more predictable than random order, and a value well below one means that something in it can be learnt. A Latin text in the first glyphs brings that ratio down to 0.66, in the word lengths to 0.70, and in the q-channel to 0.88. The test detects the implant once it has 250 to 500 channel symbols to work with. The Voynich channels show nothing of the kind: every one scores 0.98 to 1.01, within 0.01 of the meaningless controls. Two channels do score lower, every second glyph at 0.86 and words per line at 0.95, but both are just as low in the Markov imitation or in every text. They inherit the glyph trigram structure, which is how glyphs run in threes, or the page layout, and that is where their low scores come from. Some Voynich channels do have more repeated runs than a shuffle. The last-glyph channel, for instance, has 223 repeated runs of 12 symbols against 20 for a shuffle of the whole text. That excess is not the mark of a message either, because shuffling within lines keeps it (194 runs) and the copying generator reproduces it. The implanted channels behave differently, since their repeats survive no shuffle at all. To see how short a message the test can find, it was also run on ten blocks of 3,478 words, each scored against its own shuffle. A Latin implant filling one block gives 0.80 in that block against 1.00 to 1.02 in the others, and a 3,000-word stretch gives 0.89, but stretches of 1,000 or 300 words go undetected. Every Voynich block of every word-level channel scores 0.98 or above.

The planted-message check

A Latin message written into one channel of a copy of the text is found by the detector. The same channels of the manuscript show nothing.

1 Plant a Latin message into one channel in a copy of the text: the first glyph of each word key, message letter to first glyph: n > l s > a t > d i > ch u > q nine consecutive words of the planted copy, each first glyph replaced through the key: lchsy aydaiin dl ch d qchey dar char am n s t i t u t i s read back: n s t i t u t i s, part of Caesar's "lingua institutis legibus" 2 Run the detector how unpredictable the channel is on held-out pages, divided by the same measure for the channel shuffled. A score of 1.0 means no structure beyond the carrier 0.6 0.7 0.8 0.9 1.0 first glyph of a word 0.66 word length 0.70 presence of q 0.88 gallows sequence 0.88 first glyph of a line 0.77 planted Latin: found from 250 to 500 symbols the manuscript: 0.98 to 1.01 in all 29 channels 3 How large a plant the detector finds (first-glyph channel, ten blocks of 3,478 words) block of 3,478 words found: 0.80 stretch of 3,000 words found: 0.89 stretch of 1,000 words not found: 1.00 stretch of 300 words not found: 1.00 Ruled out: a language-like message of 3,000 consecutive words or more, about a tenth of the text, in any of the 29 channels tested. Not reached: a message shorter than about 1,000 words, a message diluted below one word in three, a message with no language-like statistics, or a channel the transcription does not record.
The check used for every hidden-channel test. A channel is one feature of the visible text that could carry a hidden message, here the first glyph of each word. A Latin message from Caesar is written into that channel in a copy of the text, through a fixed key. Nine consecutive words of the planted copy are shown with the letters they carry. The detector measures how unpredictable a channel is on held-out pages and divides that by the same measure for the channel shuffled, so a score of 1.0 means no structure beyond the carrier. The planted channels score 0.66 to 0.88 and are found from 250 to 500 channel symbols. Every channel of the manuscript scores 0.98 to 1.01. A plant covering a whole block of 3,478 words, or a stretch of 3,000 words, is found. Stretches of 1,000 or 300 words are not.

The same split of words into slots shows where the text's small word-order information lives, and the answer is at the word boundary. The measurement is how many bits one part of a word tells about a part of the next word, with the same correction for small-sample bias as before. The last glyph of a word tells 0.19 bits about the first glyph of the next, the suffix tells 0.08 bits about the next prefix, and the whole word tells 0.15. Once the next word's prefix is known, though, its core adds 0.001 bits and its suffix adds nothing, so almost everything one word says about the next is said across the space between them. A glyph chain of order two, which predicts each glyph from the two before it and treats the space as a glyph, reproduces 86% of the boundary value with no meaning in it. Latin shows the reverse pattern: 0.04 bits at the boundary, 0.30 from the next word's core, and 0.23 from its last letter. In a language, then, the information is in the words and not at their edges. Some correlation remains between the same slot in neighbouring words (first glyph to first glyph 0.05, core to core 0.07), and that is the size of what the copying process produces. So the word-order trace is a glyph rule that crosses the space, and not a channel in the choice of words.

ChannelVoynich text (held-out ÷ shuffled)Markov / copying controlsLatin in the same channelLatin implanted in the Voynich carrier
First glyph of each word0.9971.002 / 1.0000.9990.657
Word length0.998≈1.00≈1.000.703
Presence of q in the word0.998≈1.00–0.881
Gallows sequence0.978 (0.995 within lines)1.004 / 0.992–0.881
Line-initial glyph0.999≈1.00–0.771
Second glyph, last glyph, slots, stroke counts, yes-or-no markers0.991–1.012within 0.01––
Information across the word boundary: last glyph → first glyph of the next word (bits)0.1860.160 (glyph chain with spaces)0.036–
Next word's core given its prefix (bits)0.001–0.30–

Ruled out for a language-like message that covers 3,000 consecutive words or more, about a tenth of the text, in any of the 29 channels. This verdict holds with high confidence, because planted messages of that size were found every time. Four things lie outside it, since the planted controls also set its limits. A message shorter than about 1,000 words could be missed, and so could a message diluted below one word in three. A payload whose statistics are not language-like would not show the predictability and repetition the test looks for. And a channel that the transcription does not record, such as the glyph shapes, is out of reach of a test that reads the transcription, which is why the next section measures the shapes directly.

The examination added one historical concealment rule to calibrate the test. Trithemius hid messages in the alternate letters of alternate words of invented conjurations, as the printed key to the Steganographia explains and as Reeds (1998) recounts. This examination planted a Latin message by that rule into a carrier of the manuscript's length, and the channel test registers it at 0.66 of the shuffled entropy, where the meaningless imitation gives 0.89. A message of only a thousand channel letters still registers at 0.79. The rule has four phases, which are its four starting points: an odd or an even word, and an odd or an even letter. The manuscript's own alternate-letter channels score 0.82 to 0.89 in all four. That is the value of its imitation, as the section on other kinds of text above describes, so those channels look as they would with no message in them.

Earlier work. The steganographic reading, the idea that a message hides inside the visible text, has one substantial published treatment, by Matlach, Janeckova and Dostal (2022), who proposed a Trithemius-style code for the text in PLoS ONE. Their diagnostic was the autocorrelation of symbol reuse, which measures whether a symbol tends to come back at a fixed distance, and they also simulated the encoding of English fiction into Voynich-like text. The tests here widen that single diagnostic into 29 derived channels of the running text, each checked for any learnable sequential structure against held-out models and three shuffle nulls. The positive controls also run the other way from theirs. They built a Voynich-like text from a message, whereas the examination here plants a Latin message into a channel of the real word sequence and measures the size at which the test detects it. That reversal is what lets the negative result be stated as a size, no language-like message of 3,000 consecutive words or more. The planting design adapts their simulation, so it is not a new idea, but it had not been applied to this text before. The split of the word-boundary information by slot extends the coupling of edge glyphs that Smith (2019) measured and that Rozanova and Temerev measured in their 2026 preprint. It places the text's small word-order signal at the word boundary and not in the choice of words.

The glyph shapes, read from the scans

Every statistic so far rests on a transcription, and a transcription records only the distinctions its alphabet has. If the scribes varied the shape of a glyph on purpose, that variation would be a channel that no transcription records, and the only way to look for it is to go back to the scans. This examination therefore cut the full-resolution scans of 43 paragraph pages, in all five scribal hands, into lines, words and glyph-sized pieces without looking at any transcription, and then matched the pieces to the Zandbergen–Landini transcription. The one hand-4 page among the 43 (hand 4 being one of the five scribal hands Lisa Fagin Davis identified), the foldout f67r2, aligned no line, so hand 4 dropped out at that step. In all, 566 lines matched glyph by glyph, which gave 19,744 glyph occurrences. The main limit of this test should be stated first: the cropped glyph images are only 50 to 80 percent correct for each glyph, and the rest are fragments and misaligned neighbours. That is why grouping the images by shape recovers the transcription alphabet only weakly, with 30 clusters at 45% purity, purity being the share of each cluster that belongs to its commonest glyph. Most of the mixing comes from that segmentation noise.

For the 14 glyphs with 300 occurrences or more, the sequence of shape variants along the text is slightly more predictable than the same sequence shuffled within each page. Its held-out cross-entropy, the surprise per symbol of a model scored on pages it did not train on, is 0.4 to 4.4 percent lower for 11 of the 14. That small predictability, however, has the signature of scribal habit. The variant depends on the hand for 13 of the 14 glyphs and on the page for all 14. For half of them it depends on the neighbouring glyph and on the position in the word as well, and for 13 of the 14 it is not consistent between repeated words. The positive control was a Latin stream of vowels and consonants written into the same shape labels. It gives a 21 to 23 percent reduction with no dependence on hand, page or position, so a channel of that strength would have been seen. A control made weak on purpose gives 0.1 to 0.4 percent, which is where the measured effects fall.

One glyph, y, has a genuine split into two shapes, by size and tail, and that split does not depend on the hand. It has a 1.5 percent effect, and it depends on the page and on the next glyph, so it is a lead, compatible with pen drift and the joining of letters, and fifteen times weaker than the control. The clusters allow one more check. Conditional entropy is the surprise the next sign carries once the previous one is known, and transcribing the same glyphs by shape cluster instead of by the alphabet raises it from 2.46 to 3.89 bits. That rise is what a 55 percent noisy relabelling of a 2.4-bit stream looks like, and it is not extra information.

Twenty glyph shapes cut from the scans, each with its EVA letter

Twenty glyph shapes cut from the scans, each with its EVA letter

q=qokeedy(f108v.52) o=lol(f106r.46) a=al(f111r.47) y=shey(f103r.11) e=qokeey(f113r.8) d=qokeedy(f108v.45) l=l(f108v.9) r=okar(f108r.7) s=saiin(f58r.7) n=qoiin(f106r.7) i=aiin(f111r.53) m=qokedam(f111v.2) ch=chedy(f106r.2) sh=shet(f111r.16) k=okain(f111v.8) t=otal(f108r.9) p=peshol(f103r.2) f=chef(f43v.8) cth=cthar(f76r.2) ckh=checkhey(f111r.22)

Twenty glyph shapes cut from the scans, each with the letter it is given in EVA, the transcription alphabet. Each tile shows the ink that the alignment of scan to transcription assigned to that glyph, on pages written by hands 2 and 3. The four tall shapes k, t, p and f are the gallows. The shapes ch and sh are the benches. The shapes cth and ckh combine a gallows with a bench. Yale University Library, Beinecke MS 408, ff. 43v, 58r, 76r, 103r, 106r, 108r, 108v, 111r, 111v, 113r (details)

Undecided, with no evidence. The glyphs tested were e, o, y, a, d, l, k, ch, i, n, r, q, t and sh. For them the shape variation has the signature of scribal habit and crop noise, and it is a fifteenth of the strength the test measures on a planted stream. The verdict stays undecided because of what the test could not see. A variant weaker than the control is not ruled out. Neither are the glyphs not tested, which are s, m, g, p, f, the bench forms, and the plume variants of sh that the v101 alphabet (a finer transcription alphabet) records, nor hand 4, which had no aligned lines. So the confidence that a strong channel would have been detected is high, and the confidence that a weak one would have been detected is low.

The scans also allow the glyph shapes to be compared with the letters of other scripts. Zelinka, Lara, Windsor and Lozi (2023) made that comparison through an autoencoder, a network that learns to squeeze each image down to a few numbers and rebuild it, and they reported a resemblance to Indian scripts. They ran no control, however, so their resemblance had nothing to be measured against. This examination repeated the comparison with one. It used 12,230 clean glyph crops from the scans, in 20 classes, and compared them with the letterforms of fifteen script sets, all rendered from one font. The sets were Latin in two cases, digits, Greek, Hebrew, Arabic, Cyrillic, Armenian, Georgian, Glagolitic, Coptic, Gothic, Ethiopic, Devanagari and Bengali. The comparison ran in three spaces, as raw pixels, by their principal components (the few directions in which the images differ most), and through an autoencoder. Sixty random pen strokes made up a null script, a set of shapes with no script behind them. The positive control was 84 alphabets of known scripts, rendered from other fonts and distorted the way handwriting distorts a letter.

Those alphabets must find their own script if the method works, and they do, 95 percent of the time (93 in the autoencoder space), and with a clear margin. Their distance to their own script is 0.52, against 0.62 to the best other script and 0.71 to the random strokes. The Voynich glyphs have no such margin. Their nearest set is Glagolitic at 0.518, against 0.527 for the random strokes (0.542 against 0.562 once the sets are matched in size). Devanagari and Bengali are within a hundredth of the strokes, and every other script is farther from the glyphs than the random strokes are. A quarter of the glyph classes even have a random stroke as their nearest shape of all, and the three spaces agree on all of this. The comparison does give one thing, which is the shape structure inside the alphabet. The gallows t, p, f and k with s, the two bench-gallows forms, e with o, l with m, and y with q are one another's nearest shapes. Those are the families the clustering above found, and the comparison attaches distances to them. Beyond that, the glyphs resemble no script in the set more than they resemble random strokes.

Earlier work. Others have read glyph shapes from the scans before, for different ends. Fagin Davis (2020) did so for the palaeography of the five hands, the study of their handwriting, and Painter and Bowern (2022) for a phylogeny of glyph forms, a family tree of the shapes. Unpublished work on handwriting recognition has done so as well. Zelinka and colleagues (2023) compared the Voynich letterforms with those of other scripts through an autoencoder and reported a resemblance to Indian scripts. When that comparison is repeated above with a random-stroke null and a positive control of known alphabets, however, it finds the glyphs no closer to any of fifteen scripts than to random strokes. Two things in this section had not been applied to this text before. One is grouping the glyph images by shape, with no labels, and comparing the groups with the transcription alphabet. The other is testing the sequence of shape variants for message-like structure, against shuffle nulls, against the marks of scribal habit such as hand and page, and against a planted stream. Transcribers have long recorded variant forms and used them to tell the hands apart, but what this section measures is whether the variants hold a channel. The answer is that they follow hand and page, which is scribal habit, at a fifteenth of the strength the same test measures on a planted stream.

The information budget, and the residual as a channel

The last test of this kind asks how much of the text stays unexplained once every structure found in this report is used to predict it, and then whether what is left behaves like a message. The amount left over is no measure of content on its own, because meaningless generators leave as much, which is why the second question has to be asked. To measure it, this examination built a predictive model in stages and scored it on held-out pages, leaving out each page once, and charged each stage with what it buys. The first stage is the word frequencies with a spelling model, which prices the words the model has not seen. The second is the dependent-slot template, the word template in which the filler of each slot depends on the slot before. The third is a pair of tables, one for the position in the line and one for the Currier regime, which is whether the page is in Currier's writing style A or B. The fourth is two caches of recently used word families, one short and one long. A cache expects recent words to come back, and these put a quarter of their weight on words one edit away from a recent word. The fifth is the copying geometry of the best copying generator, which says where a copied word comes from, and the sixth is a word bigram, a table of which word follows which. The same pipeline was run on the meaningless controls, poured into the same page skeleton (the manuscript's own layout of pages and lines), and on four languages, which gives the text's budget something to be read against.

The model leaves 11.02 bits per word of the text unpredicted, which is about 93 bits per line and 383 kilobits in all, and about 3.3 of those bits are the cost of spelling words the model has not seen. Each structure buys a known amount of the rest. The slot template buys 1.75 bits. The line-position and regime tables buy 0.30, and the text is the only one of the nine texts scored (the manuscript, four meaningless controls and four languages) that gains from them. The family caches buy 0.15, where the Markov control gets 0.04, and the copying geometry buys nothing. Word order buys 0.02 bits, where the four languages gain 0.54 to 1.28. That 0.02 is the held-out gain of the word-bigram stage once every other structure is in the model, which makes it a different quantity from the 0.11 bits of adjacent-word information measured above. It also differs from the 0.02 bits per word of leftover sequential structure reported at the end of this section, where each word counts as one symbol.

The model's leftover, the residual, is the same size as the residual of the meaningless generators of the same shape, 11.1 to 12.9 bits. So the budget on its own cannot tell content from generator randomness. Nor is it a bound on how much meaning the text could hold, because it measures only what the model does not predict. A Latin word for every Voynich word would fit at the very edge of the budget, so entropy does not rule out that reading, but word order does. A word per line, or a bit per word, is far inside the budget, while a short Latin sentence per line would need every unpredictable choice to be payload.

The same budget by class of model. The examination scored four kinds of model on the same page folds, the same split of pages into training and held-out sets, to see which kind of description predicts the text best. The best procedure model, the staged model above, whose structure is written out as rules, costs 11.02 bits per word token, and a glyph 4-gram, which predicts each glyph from the three before it, costs 11.03. A class-bigram language model, the kind that describes a natural language's word order by grouping words into classes, costs 11.26, and a Latin codebook code, a model in which each Voynich word stands for a Latin word, costs 12.12. The recurrent glyph model described below, a neural network that keeps a memory as it reads, costs 10.23 on the block folds, where the text is split into blocks instead of pages. That is 0.76 bits below any glyph n-gram (a model of that kind with any context length) and 0.96 below any written-out word-level model, both gaps measured on the same block folds. The 11.02 and 11.03 above, by contrast, are page-fold costs. Two of these instruments fail their own controls, which limits what their figures can mean. A rank-matched Latin codebook, which pairs each Latin word with a code word of the same frequency rank, loses 1.1 bits per token on a text that is a Latin codebook (0.9 when it is given the true table). The class-bigram model loses 0.16 to a cache model on Latin. So neither the codebook figure nor the class-bigram figure bears on whether the text is a codebook or a language. The comparison does say two things. The sub-word models, the procedure model and the recurrent glyph model, predict the text best, and the recurrent model finds three quarters of a bit per token that no written-out model in the set represents.

The tests then asked directly whether the residual behaves like a message. To do that they turned the residual into six streams, with one value per word. Two come from the model's prediction, the surprisal of each word, which is how many bits of surprise it caused, and the rank of each word among the model's guesses. Three are the word's prefix, its suffix and its core, and the sixth is the part of the model that produced the word. Each stream was checked for sequential structure against shuffles that keep the page and the position fixed, and also for compressibility, for repeated 5-grams (runs of five symbols that come back), for bursts, and for association with position. The positive control was one Latin letter per word, planted in the suffix slot, and it is easy to see. It shows up as 0.78 bits per symbol of gain at order 3, which is what a model gains when it may look three symbols back, and it doubles the repeated 5-grams. The substitution solver, the program that breaks simple substitution ciphers, recovers its key with 84 percent letter accuracy. The real streams give 0.006 to 0.048 bits per symbol. Only the prefix stream, at 0.048, exceeds every generator control (best 0.029), and its excess of 0.02 bits per symbol, two to three standard deviations, decodes as nothing. The solver run on that stream does no better than on the Markov control's prefix stream, 5.62 against 5.66 bits per symbol, where real Latin gives 2.58. The residual does show the properties described above, seen from the other side. Paragraph-initial words cost 16.6 bits against 10.6 for interior words, the first lines of pages cost 13.8, and novelty comes in bursts. The sections also differ in predictability, with biological at 9.7, herbal at 11.3 and astronomical at 13.4 bits per word.

Undecided, with no evidence of a recoverable channel. The residual holds at most about 0.02 bits per word of sequential structure beyond what meaningless generators of the same shape produce, and a single Latin letter per word gives forty times that. Three things remain outside the reach of this test, and they keep the verdict undecided. No statistical test can see a payload that is compressed or keyed to look random. A channel could live in the variables that the page skeleton fixes, such as line lengths and word placement, and channels could live in glyph distinctions that the transcription collapses. These figures are also relative to this model, so a better model lowers them, and they do not measure how much the text could mean.

Earlier work. The idea of word entropy as a measure of how much a Voynich word could hold comes from Zandbergen, who put it near 9.9 bits and set it beside Dante and Pliny. The information-theoretic framing is standard in the literature that Bowern and Lindemann (2021) survey. The staged budget here, which charges each structure found in this report against the held-out cost of the text, extends that idea without introducing any new statistic. What had not been applied to this text before is treating the residual of a fitted predictive model as a channel, stream by stream, against generators of the same shape. That test has a clear positive control, since one Latin letter per word planted in the suffix slot shows up at 0.78 bits per symbol and the solver recovers it at 84 percent letter accuracy. The real residual streams give at most 0.048 bits, and the one stream that exceeds every generator control decodes as nothing. The planted low-rate control is adapted from the simulation of Matlach, Janeckova and Dostal (2022), and it also follows ordinary known-key practice in cryptanalysis, where a method is tried first on a cipher whose key is known.