Beinecke MS 408 · independent examination
Voynich Text Examination

What the text of the Voynich manuscript is like, measured on the page scans and on full transcriptions, and compared with real writing in thirty-one languages.

What the text is like

What is the Voynich text like, measured sign by sign and word by word? A sign is one letter of the manuscript's own alphabet. Its next sign is easier to guess than the next letter in thirty-one languages. The guess is scored in bits, and fewer bits make an easier guess. The text scores 2.11. Forty comparison texts were scored the same way. They score 2.52 to 3.69. Latin prose of the text's length repeats 67 to 92 four-word phrases. Dante's verse repeats 113 to 155. The manuscript repeats only one. Both findings are established with high confidence.

This section describes the text by five measurements before any explanation is tested, because every explanation tested below has to account for all five together. The first is how predictable the stream of glyphs is. A glyph is one sign of the script as the transcription writes it, and the standard transcription alphabet, EVA, writes each sign as a Latin letter or group of letters. The second measurement is how the words are built, and the third is whether phrases come back. The fourth is how a word's position on the page changes its glyphs, and the fifth is the two writing regimes, the two styles of writing that Prescott Currier told apart.

Glyph level: more predictable than any language

The firmest fact about the text is how easy the next glyph is to guess once the previous one is known. The measure is conditional entropy, which is how many bits of surprise the next glyph carries once the glyph before it is known, with the word space counted as a glyph. A bit is the unit of surprise, and one bit is the uncertainty of a fair coin toss. The Voynich text scores 2.11 bits, while the sixteen reference texts, the core set of real texts used for comparison, span 3.03 to 3.61. Throughout the technical sections "predictability" names this figure in bits, so a lower number means a more predictable text. The unigram entropy, which is the surprise of a glyph when nothing before it is known, is ordinary at 3.87 bits against a range of 3.85 to 4.17, so the text uses its stock of glyphs with normal frequencies. What is constrained is the order of glyphs inside words, because shuffling the glyphs within each word, and changing nothing else, lifts the value to 3.62 bits, above every language.

Where each glyph stands in the word

For each of the 18 commonest glyphs, named by its EVA letter, the share of its occurrences that open a word, stand inside it or close it. The count is over the 34,780 running-text words of the Zandbergen-Landini transcription. One-glyph words are counted apart. The n beside each bar is the number of occurrences.

word-initial inside the word word-final one-glyph word o 31% 63% 4% n = 22,295 e 99% n = 18,038 h 100% n = 16,535 y 9% 87% n = 16,057 a 14% 86% n = 12,837 c 53% 47% n = 12,195 d 27% 68% 4% n = 11,839 i 100% n = 11,016 k 12% 87% n = 10,033 l 13% 30% 55% n = 9,468 r 5% 16% 77% n = 6,536 s 60% 22% 12% 5% n = 6,497 t 16% 83% n = 6,010 n 99% n = 5,797 q 99% n = 5,330 p 35% 63% n = 1,488 m 96% n = 932 f 25% 70% 4% n = 418 0% 25% 50% 75% 100% share
Several glyphs are almost fixed to one position. The glyph q opens a word in 99.4 percent of its occurrences and never closes one. The glyphs n, m and y close a word in 99, 96 and 87 percent of their occurrences. The glyphs i, h and e stand inside the word. This fixed order of glyphs within the word is the slot template described in the next section. The 18 glyphs cover 99.9 percent of all glyph occurrences in the running text.

The first five lines of f. 1r, in two halves, with their EVA transcription

The first five lines of f. 1r, in two halves, with their EVA transcription
The first five lines of the first page, f. 1r (Currier A, the first of the two writing styles Prescott Currier identified, written by hand 1, the first of the five scribal hands Lisa Fagin Davis identified in 2020), shown in two halves. The left half of each line is above and the right half below. EVA is the transcription alphabet, which gives each common glyph shape a Latin letter. In EVA the lines read as follows. First line: fachys ykal ar ataiin shol shory cthres y kor sholdy. Second line: sory ckhar or,y kair chtaiin shar ase cthar cthar,dan. Third line: syaiir sheky or ykaiin shod cthoary cthes daraiin sy. Fourth line: soiin oteey oteos,roloty cthiar,daiin okaiin or okan. Fifth line: sair,y chear cthaiin cphar cfhaiin. The word ydaraishy, which the transcription records as a sixth line, stands alone at the right end of the fifth line. A comma marks a word space the transcribers were unsure of. The reading names shapes. It is not a translation. The dark mark after the first line is not part of the text. The transcription records it as a non-Voynich symbol. Yale University Library, Beinecke MS 408, f. 1r (detail)

Conditional entropy of the next glyph, bits

Given the previous glyph, with the word space counted as a glyph. Each reference text is sampled to the Voynich size (208,278 characters), and each bar shows the mean of 20 sample windows.

Voynich transcriptionsReference languagesControls
Table of values
CorpusbitsNote
The gap to the sixteen core texts is a full bit, and no sample window of any of them comes within 0.9 bits of the Voynich value. The wider survey below narrows the gap to 0.41 bits, with Maori the nearest language. The three Voynich transcriptions agree with one another, and the v101 alphabet, a finer transcription that keeps more shapes apart, gives 2.53. The smaller reference texts (not shown) score 3.15 to 3.61 at a 10,000-word tier.

Three more facts about the glyphs hold in every transcription and in both Currier regimes, the two writing styles:

Tall gallows glyphs on the first line of a paragraph

Tall gallows glyphs on the first line of a paragraph
The first line of a paragraph on f. 106r (Currier B, hand 3), beside the star that marks its start. The words are pshdar shoefy yteedy shal korchy sheky. The tall glyphs are the gallows, the four shapes written p, f, t and k in EVA. Here p opens the first word, f stands in the second, t in the third, and k in the fifth and sixth. On the first line of a paragraph, p and f are about fifteen times as frequent as on other lines. Yale University Library, Beinecke MS 408, f. 106r (detail)

The glyphs also alternate between two classes, the way vowels and consonants alternate in a language. Sukhotin's algorithm sorts the signs of an unknown script into two classes, on the principle that signs of one class tend to stand next to signs of the other. It gives the text an alternation ratio of 1.48. That ratio measures how strictly the two classes take turns, and since the languages range from 1.42 to 1.68, the text is inside the language range. The glyphs a o y e ch form one class. This establishes an alternating structure, but it does not establish that the classes are vowels and consonants.

The EVA letters and where they stand in the word

EVA letterPer 1,000 lettersCountCommonest positionShare there, %
o128.522,295inside the word63
e104.018,038inside the word99
h95.316,535inside the word100
y92.516,057last in the word87
a74.012,837inside the word86
c70.312,195first in the word53
d68.211,839inside the word68
i63.511,016inside the word100
k57.810,033inside the word87
l54.69,468last in the word55
r37.76,536last in the word77
s37.46,497first in the word60
t34.66,010inside the word83
n33.45,797last in the word99
q30.75,330first in the word99
p8.61,488inside the word63
m5.4932last in the word96
f2.4418inside the word70
g0.7121last in the word84
x0.119inside the word42
b0.115last in the word100
j0.110inside the word60
u<0.17last in the word100
z<0.14inside the word75
v<0.11first in the word100
The 25 EVA letters of the transcription as they occur in the running text, 173,498 letters in 34,780 words. For each letter the table gives its frequency per thousand letters, the position in the word where it most often stands, and the share of its occurrences at that position. EVA, the transcription alphabet, gives each common shape a Latin letter. The shapes written ch and sh, and the bench gallows written cth, ckh, cph and cfh, take several letters each, and almost every c and h in the text belongs to one of those groups. John Tiltman in 1967 and Prescott Currier in 1976 described the preference of glyphs for fixed positions inside the word.

Two further views of the glyph stream. The first view merges glyphs into larger units. A byte-pair encoder merges the commonest pair of adjacent units into one new unit and repeats that step by step. At each level of merging the test measures how much the next unit depends on the current one, which gives the U-shaped curve that Rozanova and Temerev reported, with its lowest point at 64 merges. The value here is 1.053 bits against their 1.045, reached at 89 units with a mean span of 2.40 glyphs. Every language has its lowest point at zero merges, but three texts with no meaning reproduce the Voynich U-shape. The first is the meaningless Markov imitation, a random text generated from the manuscript's own counts of which glyph follows which pair of glyphs. The second is the drifting-state generator, a text generator whose stock of words drifts as the pages go by, which the section on what produced the text describes. The third is a Latin codebook, a Latin text with each word replaced by its own code group. Their lowest points are 0.93 to 1.10 bits, at 80 to 96 merges. The reason is that the first sixty-odd merges absorb the fixed glyph groups (ch, dy, aiin, ok, qok). After that the units are word-sized and straddle the word boundary, where the coupling between units is strong. The curve is therefore another face of the slot template, and it tells nothing further about the process that made the text. The second view recodes each glyph as a fixed sequence of pen strokes, a proposal from the forum literature, and that removes the anomaly automatically, whatever the text. The stroke stream scores 1.68 bits, against 1.56 for Latin and 1.67 for Italian, because a code built from three primitive strokes caps the entropy of any text near log2 3 bits. That cap is the most that a three-way choice can give, and permuting the stroke table moves the value more (to 1.47) than changing the text does. Whether the alphabet is a stroke-composition code cannot be tested with a table, because it would need the strokes read from the scans.

The same comparison at matched inventory. Part of the entropy gap could be an artefact of counting. The text is read with 24 or 38 signs while each language is read with its own alphabet, and a stream with fewer signs has less room for surprise. So this examination merged the 28 sampled languages down to the text's 24-sign and 38-sign inventories. The merging used the text's own unit-induction procedure, which joins the pairs of signs that occur together far more often than chance would give. The same procedure built the merged view of the text, the reading in which the conventional glyph groups count as single signs. Under that merging the languages' conditional entropy rises by only 0.0 to 0.4 bits: the lowest language is Hausa, at 2.74 and 2.86 bits, and the medians are 3.33 and 3.41. The text stays the lowest text in the sample at every inventory, but its margin over the nearest language shrinks. At letters the margin is 0.60 bits, with Hausa the nearest language in that set. The three Polynesian scriptures measured in the language section below are not in that set, and they come nearer to the text. For the conventional merged view and the v101 view the margin is 0.31 to 0.36 bits, and it is 0.04 to 0.25 bits when the text's own letters are merged by the same procedure. Mutual information between units is how much one unit tells about the next, and at the merge level where it drops below one bit (97 signs, 3.72 bits) the text is inside the language range. One quantity does not shrink with the inventory, and that is the order information inside the word, which is the entropy of the shuffled words minus the real value. It is 1.14 and 0.94 bits for the text, against at most 0.95 and 0.92 for any language. The predictability margin is therefore quoted together with the sign inventory it was measured on. It is 0.6 bits at letters and about 0.3 bits at a matched inventory, and the within-word order information is the part that no choice of unit removes. A real syllabic rendering goes the other way. An abugida is a script with one sign per consonant-and-vowel unit, and Sanskrit written that way needs 358 to 1,160 signs to cover 99 percent of the text, at 4.5 to 4.6 bits per sign. So the text's v101 view, at 2.53 to 2.55 bits by sample, is 1.9 to 2.0 bits below a syllabic script, twice its distance from Sanskrit letters. The mean word length in signs, 3.2 to 3.9 for the Sanskrit, matches the text's 3.9, but the skew, 1.1 to 1.2 against 0.23, does not.

The figure under other estimators. The 2.11 bits is a plug-in estimate, which means it is computed straight from the observed counts, on 200,000 glyphs. The bias-corrected estimators of Grassberger (2003) and Chao and Shen (2003), which allow for the way a plain count from a limited sample runs low, give 2.108 and 2.109, which leaves the figure almost unchanged. The interval from page half-samples is 0.008 bits, and the spread between quires (the gatherings of leaves that make up the book) is 0.049. The tests also measured longer-range structure as excess entropy, the information that a block of glyphs holds beyond what its parts hold. At this text length that measure is usable only up to blocks of four glyphs. There the text has 2.09 bits, below all 30 language samples (2.28 to 6.82) and equal to the drifting-state generator (2.08 ± 0.05). At block length eight the sample-size bias is as large as the value, so a text of this size tells nothing at that length. Two whole-text measures complete the profile. Lempel-Ziv complexity counts how many new strings a compressor meets as it reads through the text. On that measure the text is more compressible than the drift generator (0.539 against 0.580 ± 0.002, by the measure of Kaspar and Schuster 1987). The Kneser-Ney held-out rate is the bits per glyph that a statistical model of the text needs on pages it was not fitted on. On that measure the text is more predictable (1.88 against 2.00 ± 0.02 bits per glyph, with a spread of 0.10 between quires). On both measures the text lies inside the language range (0.39 to 0.71 and 1.22 to 2.68).

Earlier work. The low conditional entropy of the glyph stream is the oldest quantitative fact about the manuscript. Bennett measured it in 1976, and Zandbergen's entropy pages set it beside Latin, Italian and German. Lindemann and Bowern, in a 2021 preprint that has not been peer reviewed, placed it below all of several hundred comparison texts. They also showed that it does not follow from the transcription alphabet, from abbreviation or from missing vowels. The values here agree with theirs, at 2.11 bits in the EVA reading (EVA being the standard transcription alphabet) and 2.53 in the finer v101 alphabet, so this measurement confirms a result that was already established. What this examination adds is three things: it cuts the reference texts to exactly the text's size, runs all three conventions for word spaces, and runs the within-word shuffle. That shuffle matters because Lindemann and Bowern explained the deficit by positional restriction inside words, and they reached that by inference. The shuffle turns it into a direct control, a comparison text in which that restriction alone has been destroyed. The two-class alternation that Sukhotin's algorithm recovers is Guy's (1991) and Reddy and Knight's (2011) result, and this examination reproduces it, with an alternation ratio inside the language range. Two recent measurements belong here as well. The dependence gap across merge levels, defined in Rozanova and Temerev's 2026 preprint, comes out almost exactly the same here (1.053 against their 1.045). But the tests show that it holds no information beyond the trigram statistics (the counts of which glyph follows which pair), because a meaningless imitation gives the same curve. This examination is also the first to measure the stroke-level reading of the glyphs proposed in Cham's (2014) blog work, and the reading adds nothing. What it does is spread the anomaly over a larger number of strokes per word, in a way that depends on the stroke table chosen. The published conditional-entropy figures for the text, Bennett's, Stallings's, Zandbergen's and Lindemann and Bowern's, are plug-in estimates at one block length without a bias correction. The one published compression measure, Gaskell and Bowern's (2022) ratio on 200-word excerpts, is a single number. The profile above puts the same quantities under the bias-corrected estimators, the excess-entropy curve, a Lempel-Ziv ratio and a quire jackknife, which recomputes the figure with one quire left out at a time. It does so on the same language samples and generator seeds (a seed being the starting point of a program's random numbers) throughout, and no earlier work on this text does that.

Word level: a language-like vocabulary built from a rigid template

At the level of whole words, the text looks at first like a language. Zipf's law describes how quickly word frequency falls as one goes down the ranking from the commonest word, and the Zipf slope measures the steepness of that fall. The text's rank-frequency curve has a Zipf slope of −1.04 over the thousand commonest words (−0.93 over all of them). Heaps' law describes how the number of distinct words grows as a text gets longer, and the Heaps exponent is the rate of that growth, which for the text is 0.72. One word in five is distinct (a type-token ratio of 0.20, where a type is a distinct word and a token is one occurrence of a word), and of the distinct words 68% occur once. The ten commonest words make up 12.8% of the text, a share like Latin's or Finnish's and unlike English's. All of these values fall inside the range of the sixteen reference texts.

4.99 ± 1.79
mean word length in EVA letters, with its standard deviation. The variance is 3.19, against 4.0 or more for every alphabetic reference text
0.03
skewness of the word-length distribution, which measures how lopsided it is. The reference texts are 0.20 to 1.7, skewed to the right with a long tail of long words. The one exception is a Basque sample at 0.03, whose words are a third longer and nearly four times as spread
~60 rules
fillers in an eight-slot ordered template, enough to generate 90% of all word tokens. The same procedure needs about twice as many for Latin and still rejects far fewer scrambled words
3 in 4
distinct words that are one glyph edit (one glyph added, removed or changed) away from a more frequent word (Latin 1 in 4, Italian 2 in 5, Hebrew 1 in 2)

But the words are built differently from words in any of the languages. The spread of word lengths is narrow and symmetric, whereas in every language it is wide, and in all but one it is skewed to the right. The exception, a Basque sample, is symmetric but nearly four times as spread. The symmetry depends on taking the letters as the units. Once the text's glyphs and the languages' letters are merged to the same 24- or 38-sign inventories, the text's skew rises to 0.27 to 0.47, the values for the text's letters merged by the procedure. Some languages then fall to 0.03 to 0.21. The v101 transcription, with its own inventory of 38 signs for 99 percent coverage, has a skew of 0.23, the figure used in the abugida comparison above. The narrowness, on the other hand, does not depend on the unit, because at a matched inventory the text's variance is 2.1 to 2.8, against 3.6 or more in every alphabetic text. An ordered-slot grammar derived from the text alone describes how the words are built, and it has the following form.

The word template

The eight ordered slots of the template, the fillers each slot allows, and four of the commonest words placed on the slots they use.

slot 1 slot 2 slot 3 slot 4 slot 5 slot 6 slot 7 slot 8 q o y a ch sh k t p f l r ch sh e ee d a o e i d l y a o r in iin n m l r s fillers allowed in each slot real words laid on the slots daiin 816 times d a iin chedy 493 times ch e d y chol 378 times ch o l qokeedy 307 times q o k ee d y Each slot is optional and holds at most one filler, and the fillers of a word must fall in increasing slots. About sixty fillers in this fixed order generate 90% of all word tokens. The words shown are the commonest words of three or more glyphs that use every slot between them.
Each word of the text is built from a short template of ordered slots. The top row gives the fillers each slot admits, as derived from the text. Below it, four of the commonest words are placed on the slots their glyphs occupy, each glyph taking the first slot that admits it. The count of each word in the running text is given beside it. Together the four words use every slot.

Word lengths against four languages

The share of words of each length, in glyphs for the text and in letters for the languages. The 34,780 running-text words of the transcription are compared with samples of the same size from Latin (Caesar), Italian (Dante), Hebrew (a collection of Geonic responsa) and Maori (the New Testament).

0 5 10 15 20 25 share of words, % 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15+ word length in glyphs (letters for the languages) Voynich text: 4.99 ± 1.79 Latin (Caesar): 6.27 ± 2.93 Italian (Dante): 3.98 ± 2.25 Hebrew (Geonic responsa): 4.03 ± 1.52 Maori (New Testament): 3.35 ± 2.18 mean ± standard deviation of the word length
The text's distribution of word lengths is narrow and nearly symmetric, with its peak at 5 glyphs. The mean is 4.99, the standard deviation 1.79 and the skewness 0.03, where skewness measures how far a distribution leans to one side. Latin and Italian spread wider and tail off to the right, with skewness 0.35 and 0.72. Hebrew, written without most vowels, is no wider than the text and leans to the right, at 0.20. Maori, with a small alphabet and open syllables, is wider than the text and leans to the right, at 1.37. Each language sample is 34,780 words taken from 40 percent of the way through its text. The Hebrew text has only 23,643 words and is used whole.

q · o / y / a / ch / sh · k / t / p / f / l / r · ch / sh / e / ee / d · a / o / e / i · d / l · y / a / o / r · in / iin / n / m / l / r / s

Each slot is optional, and with about sixty admissible fillers this grammar generates 90% of word tokens. This examination ran the same learning procedure on Latin, Italian and Hebrew text of the same size. There it needs 73 to 116 rules, and its grammars accept 40 to 56% of the same words with their letters scrambled. The Voynich grammar accepts only 14 to 17% of scrambled words, because it encodes order. The set of strings the Voynich grammar admits, relative to the words actually attested, is one to two orders of magnitude smaller than for the languages, that is, ten to a hundred times smaller.

A herbal paragraph in the eight slots of the word template

Slot 1: q. Slot 2: o, y, a, ch, sh. Slot 3: k, t, p, f, l, r. Slot 4: ch, sh, e, ee, d. Slot 5: a, o, e, i. Slot 6: d, l. Slot 7: y, a, o, r. Slot 8: in, iin, n, m, l, r, s.

LineWords as transcribed (EVA)
1fchodaiin shopchey qko shey qoos sheey charochy
2dcheey keo r shor dold dchey kchey otchy cheody
3oeeees cheodaiin sheey ytcheey qotchy chald
4qokcho cho lochey daiin ychey kchos odaiin
5oaiir otaiin
LineWord12345678Left over
1fchodaiinfchodaiin
shopcheyshopchey
qkoqko
sheyshey
qoosqoos
sheeysheey
charochycharochy
2dcheeydcheey
keokeo
rr
shorshor
dolddold
dcheydchey
kcheykchey
otchyotchy
cheodycheody
3oeeeesoeeees
cheodaiincheodaiin
sheeysheey
ytcheeyytcheey
qotchyqotchy
chaldchald
4qokchoqokcho
chocho
locheylochey
daiindaiin
ycheyychey
kchoskchos
odaiinodaiin
5oaiiroaiir
otaiinotaiin
The first paragraph of page f7r (herbal section, Currier A, hand 1): 5 lines and 31 words in the EVA reading. Below it, every word is laid out in the eight ordered slots of the template printed in the Words section, with each unit in the earliest slot that admits it. A word that does not fit is shown up to the unit that fails, and the rest of the word stands in the last column. 21 of the 31 words fit. Over the whole running text, this printed template of 34 fillers fits 74 percent of the words, and the fitted grammar of about sixty fillers fits 90 percent. The template goes back to John Tiltman (1967), Prescott Currier (1976) and Jorge Stolfi (2000).

The vocabulary is correspondingly dense. Three-quarters of distinct words are within one glyph edit of a more frequent word (76% by a strict count, 83% by a looser one), against 27 to 36% for Latin. The three hundred commonest words form a single connected network of one-edit neighbours, where Latin's form 178 separate islands. Half of the words that occur only once (49%) split into two shorter words that both occur in the text, each five times or more, where in Latin 3% do and in Italian 10%. The share moves between 33 and 81% with the cut-off, but the contrast with Latin holds under every cut-off. The variation between words is at the beginning and middle of the word, while the endings stay stable, which is the mirror image of inflected languages, where stems are stable and endings vary.

Neighbouring words one glyph apart

PageLineNeighbouring wordsDifferenceSection
f1r16kshy kchys changed to ctext, Currier A
f17v5ychol choly droppedherbal, Currier A
f31r6qokedy qotedyk changed to therbal, Currier B
f45v4dain dailn changed to lherbal, Currier A
f67r243daiin aiind droppedastronomical, no regime code
f86v610olaiin olkaiink addedtext, Currier B
f102r18oteol qoteolq addedpharmaceutical, Currier A
f116r6lchedy chedyl droppedstars and recipes, Currier B
Eight pairs of neighbouring words that differ by one glyph, whether added, dropped or changed, with the page and line where each pair stands. Each is the first such pair of words of four letters or more on one of eight pages spread through the book. Over the whole running text, 1,105 of the 30,651 pairs of neighbouring words within a line differ by one glyph, which is 3.6 percent. When the same words are shuffled across the text and poured back into the same lines, the share falls to 1.7 percent (the mean of 20 shuffles, range 1.5 to 1.8). Identical neighbours are not counted. Torsten Timm described this excess of near-copies among neighbouring words in 2014.

Four consecutive words each one glyph away from the next

Four consecutive words each one glyph away from the next

okeey qokeey qokeedy qokeey chedal chedy qokeey okeedain otain oolals

The start of a line on f. 108r (Currier B, hand 3). The words are okeey, qokeey, qokeedy, qokeey and then chedal. Each of the first four is one glyph away from the next. Over the whole text, neighbouring words are near-copies of one another far more often than chance allows. Three quarters of the distinct words lie one glyph from a commoner word. Yale University Library, Beinecke MS 408, f. 108r (detail)

A direct test separates two ways of building words: one glyph after another, or by nested categories. If words were codes in which the first element picks a class and later elements pick a subclass, then the first glyph would predict the later glyphs beyond what the glyph just before them does. In the text it adds little. The held-out estimate used here, which fits the model on part of the text and scores it on the rest, puts the addition at no more than 0.05 bits over the whole word. Position by position, the only addition is 0.17 bits at the third glyph, and there is nothing beyond the fifth. Across the transcriptions the figure is −0.01 to +0.05, and the meaningless Markov imitation, a random text generated from the text's own counts of which glyph follows which pair, gives −0.09. A category code built from the text's own statistics gives only 0.08, and a plug-in estimate, computed straight from the counts, puts it at 0.12 bits. In the languages, by contrast, the first letters of a word narrow the word down: Latin 0.63 bits, Italian 0.44 and Hebrew 0.32 held out (0.41 to 0.47 plug-in). Read from the other end, the last glyph adds nothing about the earlier glyphs beyond its neighbour either (−0.03 to +0.01, against 0.44 in Latin). The test can catch codes in which the class chosen first narrows the later choices beyond what the adjacent glyph already tells. A code whose classes combine as independent or neighbour-linked slots is, statistically, a slot template, and the report judges it as one. Uncertainty inside a Voynich word stays high until the last two positions and then collapses, whereas in the languages it falls steadily from the third letter on. So Voynich words are not recognised from their beginnings the way the words of a language are.

Published word grammars, re-run. Reddy and Knight (2011) fitted a two-state hidden Markov model to the letters. Such a model supposes that each letter was produced by one of two hidden states and finds the two states that fit best, and in their solution one state emits the word-final letters. They read that as the a*b grammar of an abjad, a script that writes consonants and leaves most vowels out. In that grammar a word is any run of letters from one class, closed by one letter from the other. The same fit on the EVA letters gives their result, with the second state being the final cluster a i n r l, but on the merged-glyph and v101 alphabets a different solution wins, by 0.12 to 0.13 bits. That solution is a strict alternation of two letter classes, whose switch rate, how often the class changes from one letter to the next, is 0.92 to 0.93, against 0.80 to 0.84 in the languages. And in every view the Markov imitation of the glyph stream gives exactly the same solution as the text, so the result depends on the symbol inventory and adds nothing about meaning. Parisel (2026) proposed four "boundary signatures", statistical marks of the word boundary, and this examination reproduces them in substance. A transition from the word's ending class to its starting class occurs at 67.7% of Currier A boundaries and 64.7% of B boundaries. His repository's calibration gives 71.0 and 64.3, and his paper's 80.6 is not reached. The glyphs at the two ends of the word are strictly asymmetric. In the default reading iin closes a word 335 times and opens one once, while the glyph q opens a word 5,297 times, which is 99.4 percent of its 5,330 occurrences, and never closes one. And 0.22 to 0.31 bits of information cross the space between words, of which 22 to 55% survives a shuffle of the words within lines. The Markov imitation scores three of the four signatures, as the text does, so the signatures are properties of the glyph trigram statistics with spaces, and they are not evidence about how the text was made.

Four regularities that hold for words in languages hold here too, and each has a cause that has nothing to do with language. Alice Kober's method for an unknown script draws up a grid of stems against endings and looks for stems that share whole sets of endings. It finds the text more "paradigmatic" than Latin, meaning that more of its stems share a full set of endings. In the text 48% of stems share a complete ending set with another stem, against 26% in Latin, and the glyph before the ending predicts the ending almost twice as strongly (0.99 bits against 0.54). But the Markov imitation gives 44% and 0.94 bits, and the slot generator (a text generator that builds each word from the slot template) gives 17% and 0.84. So the grid records the small stock of suffixes in the template (in, n, r; iin, l; dy, ey, y), and it does not record an inflection. Layfield and colleagues (2020) found that rarer words become identifiable earlier in the word, as they do in languages, and that holds here, with a rank correlation of +0.40 against +0.19 to +0.39 in the languages. But the Markov imitation gives +0.39, and even the within-word shuffle gives +0.40. The reason is that rarer words are longer everywhere, and the frequency-weighted trie used in the test, a branching index of the words, assigns the shared prefixes to the frequent words by design. The law of abbreviation, which says that shorter words are commoner, holds with a correlation of −0.32 over types, where 13 of 30 languages are weaker and 17 are stronger. The refinement of Piantadosi and colleagues, the length-against-surprisal effect, in which words that are predictable in context are the shorter ones, is absent (+0.03), as it is in most languages at this sample size. A Menzerath relation, in which a word with more units has shorter units, is present (exponent −0.59, the languages −0.71 to −0.47). But every generator and the shuffle reproduce it, so it is a property of the segmentation and the length distribution. None of these regularities separates the text from a meaningless imitation.

Agglutinative languages as controls. An agglutinative language builds its words from ordered strings of suffixes, which makes it the natural-language case nearest to a slot template. So Turkish, Uzbek and Swahili went through the same procedures as controls, which are texts of known origin run through the same tests. Turkish and Uzbek need 114 to 144 slot rules to cover 90 percent of their tokens, where the text needs 60 to 71. They also give the largest first-symbol dependence of any corpus, 0.69 to 0.73 bits against the text's 0.0. Their family density, the share of distinct words within one edit of a commoner word, is 0.25 to 0.36, against the text's 0.83 to 0.86 counted over its whole stock of word types, by letters and by glyphs. In a sample of the text's window size the share is 0.74 to 0.76, the figure the other sections use. Their grammars accept 19 to 24 percent of shuffled words, where the text's accepts 14 to 17. And Swahili, which builds its words with prefixes, varies its words at the beginning as the text does, so those two marks of the template are not unique to the text. A successor-variety segmentation cuts words at the points where many different letters can follow, and it is the kind of method used to discover paradigms. It finds 27 to 79 paradigm signatures (sets of endings shared by a group of stems) in every inflecting language, but in the text it finds 19, which is the number the text's own Markov imitation yields. The vowel and consonant classes that Sukhotin's algorithm recovers imply a syllable canon, a list of allowed syllable shapes, of 35 to 37 patterns for 90 percent of tokens. Of those tokens, 83 percent parse as consonant-vowel-consonant with optional margins, which is the language median, and the Markov imitation of the glyph stream reproduces every statistic of that canon. A second-order chain is a model in which each glyph depends only on the two before it. Vowel harmony, the rule in Turkish that the vowels of a word agree with one another, produces an association between vowels beyond what such a chain gives. Here that association is 0.05 bits, where Turkish gives 0.50.

Earlier work. Tiltman (1967) and Currier (1976) described Voynich words as built on a fixed template of ordered parts, Stolfi (2000) formalised that as the crust-mantle-core grammar, and Zattera (2022) gave it a slot alphabet. The eight-slot grammar fitted here reproduces that picture and adds a measure of how much smaller the admissible set of strings is than in Latin, Italian or Hebrew. The narrow, symmetric spread of word lengths is Stolfi's binomial (the symmetric curve he fitted to the word lengths) and Reddy and Knight's (2011) observation. Reddy and Knight's two-state letter model gave the word grammar a*b, and that has been read as pointing to an abjad, but this examination reproduces it only in the EVA letter reading. Once the conventional glyph groups are single symbols, a strict alternation wins instead, and a meaningless imitation of the glyph stream matches both solutions exactly. That published hidden-state result therefore depends on the symbol inventory and adds nothing about meaning. This examination reproduces the boundary signatures of Parisel's 2026 preprint in substance, though the published numbers depend on which release of his work is taken. The meaningless imitation meets his four-signature criterion, as the text does, and the tests show that the cross-boundary information which survives is mostly finite-sample bias, the excess that a limited sample produces on its own. Layfield and colleagues (2020) found that the uniqueness point of a word, the glyph at which no other word shares its beginning, comes earlier for rarer words. This examination reproduces that, but the relation is uninformative, because a glyph-level imitation and a slot template give the same relation. This examination is the first to measure the brevity law (the law of abbreviation), the length-against-surprisal effect of Piantadosi, Tily and Gibson (2011) and the Menzerath-Altmann relation on this text. Controls with no content reproduce all three. Two tests in this section have no earlier application to the Voynich text. The first asks whether the first glyph of a word constrains later glyphs beyond the preceding one, which separates one kind of nested category code from a left-to-right template. The second is the Kober-Ventris paradigm grid, which can be drawn for this text but comes out the same for a meaningless imitation. A self-published site, The Voynich Project, ran a syllable-shape and vowel-harmony battery (a set of tests applied in the same way) from the Sukhotin classes in 2026, with Turkish and Finnish controls, and reached the same conclusion. Reddy and Knight (2011) extracted affix signatures from the text with Goldsmith's Linguistica, a program that finds affixes without supervision, and Lindemann (2022) ruled out the heavily agglutinating families on type-token statistics. The agglutinative controls above add the slot, family and sequence measurements on Turkish and Uzbek.

Sequence level: words repeat, phrases never

This is the decisive measurement, because it does not depend on the transcription alphabet, on where the word breaks fall, or on any assumption about what a glyph stands for. Natural texts repeat stretches of several words all the time: set formulae, idioms, chains of small grammatical words, names said again. A cipher that always writes the same plaintext as the same glyphs keeps those repeats, whatever its mechanism. A cipher whose rule changes with position or with some internal state need not keep them, and the report judges that family on other grounds below. The Voynich text has practically no repeated phrases.

One repeated four-word sequence, against 67 in a Latin sample

TextFour-word sequenceTimesWhere
Voynich running text, 34,780 wordsol shedy qokedy qokeedy2f75v, line 47, words 5 to 8; f84r, line 22, words 4 to 7
Latin, Caesar, Gallic War, 34,780-word samplequa re nuntiata caesar5the commonest of 67 repeated phrases
proficiscitur eo cum venisset3
qui arma ferre possent3
The Voynich running text, 34,780 words in the Zandbergen–Landini reading, contains one sequence of four words that occurs twice. Both occurrences are on pages of the biological section. No sequence of five words occurs twice. A sample of Latin prose of the same length, from Caesar's Gallic War, contains 67 distinct four-word sequences that occur at least twice. All of them are Latin phrases, and the three commonest are shown. René Zandbergen recorded the absence of repeated phrases in the text, and Torsten Timm counted the repeats in 2014.

Distinct four-word sequences occurring at least twice

Each text is cut to 34,780 words, and the scale is logarithmic. The Voynich value is 1.

Voynich (Zandbergen–Landini reading)Reference languagesControl: Voynich, words shuffled
Table of values
Corpusrepeated 4-word sequenceswords
The one repeated Voynich sequence is ol shedy qokedy qokeedy, which occurs twice, and no five-word sequence repeats anywhere in the Voynich text. Latin prose, the language with the freest word order in the set, has 67 to 92 depending on the sample, and even Dante's verse has 113 to 155.

Word boundaries could be the wrong unit, so this examination ran the test again with every space removed, counting repeated strings of glyphs of a fixed length anywhere in the text. At 25 glyphs the Voynich stream has one repeated string, at 30 glyphs it has none, and its longest repeated string is 25 glyphs long. The reference texts, cut to the same number of characters, have 33 to 5,390 repeated 25-character strings and 0 to 3,899 repeated 30-character strings. Those figures are for the sixteen core texts, one sample each, with Dante at 0 and Montaigne at 4. The wider set has 140 samples of all twenty-six languages, or 131 distinct windows, since short texts allow fewer. Of those samples, 131 have 28 to 19,637 such strings, four samples of Dante's verse have none, and five samples of Montaigne, Cervantes, modern Greek prose and Welsh have 4 to 26. The longest repeats of the sixteen core texts are 37 to 200 characters. Two qualifications belong here. First, a verbose encoding stretches the plaintext, so 30 glyphs might stand for only 10 to 15 letters, and the comparison has to be made at equal plaintext length. Made that way, the comparison gets stronger, because the references have 180 to 9,779 repeated 16-character strings and 1,077 to 15,289 repeated 12-character strings, against none of 30 glyphs here. Second, the length-30 count is the weakest form of the argument on its own. At 20 to 25 glyphs the Voynich stream (73 and 1) is close to Dante's verse (44 and 33), and other samples of that poem reach zero at 30 characters. The four-word count, on the other hand, separates the text from all 140 language samples (1 against at least 67). There is also the question of spelling. The people who transcribed the text disagree on 2% of characters, but that is not the scribes' error rate, and a scribe who spelled the same word several ways would break repeats that the transcribers copy faithfully. So this examination measured the tolerance. Repeated phrases in a language text survive spelling variation up to five times that rate. Removing them takes five to fifteen times the rate, and at that point the entropy and vocabulary statistics move away from the text's values. The text's own four-word count does depend on transcription conventions. It is 1 as transcribed, 8 with stroke counts merged, 22 with a/o, r/s and ch/sh merged as well, and 45 with k/t and o/y merged too. Language samples merged the same way have 38 to 67. The 30-glyph count does not depend on the conventions (0 or 1 under every merge), and so the two counts are read together.

Repeated strings with the spaces removed

A window of 30 glyphs is slid along the stream of glyphs and counts exact repeats, with no word boundaries involved.

Latin (Caesar), as written gallia est omnis divisa in partes tres quarum unam the same glyphs with the spaces removed g a l l i a e s t o m n i s d i v i s a i n p a r t e s t r e s q u a r u m u n a m 30-glyph window, moved one glyph at a time exact repeats of a 30-glyph window in the language samples: 28 to 19,637 (131 of the 140 samples) Voynich, page f75v line 47, as transcribed qokain olshey qokain dar ol shedy qokedy qokeedy the same glyphs with the spaces removed q o k a i n o l s h e y q o k a i n d a r o l s h e d y q o k e d y q o k e e d y 30-glyph window, moved one glyph at a time exact repeats of a 30-glyph window in the text: 0. The longest repeated string is 25 glyphs Repeated four-word sequences (texts cut to 34,780 words): the text 1, every language sample 67 or more Why removing the spaces makes the count blind to where the word breaks fall as transcribed ol shedy qokedy qokeedy the same glyphs broken elsewhere olsh edyq okedyq okeedy spaces removed, either way olshedyqokedyqokeedy The one repeated four-word sequence, ol shedy qokedy qokeedy, is a string of 20 glyphs once the spaces are removed. A search for repeated strings finds it whichever way the breaks fall. No window of 30 glyphs repeats anywhere in the text. So the count does not depend on the transcribers' word divisions, and no fixed encoding of the glyphs could change it.
A line of Latin from Caesar and a line of the Voynich text (page f75v, line 47) are shown as written and with their spaces removed. A window of 30 glyphs is slid along one glyph at a time, and the count is how many windows recur exactly elsewhere in the stream. The count is 28 to 19,637 in 131 of the 140 language samples. In the text it is zero, and the longest repeated string is 25 glyphs. Counted by words, the text has one repeated four-word sequence, against 67 or more in every language sample. The spaces are removed before counting, so the result does not depend on where the word breaks fall.

At the same time, single words repeat more than in any prose. The same word appears twice in a row 0.9% of the time, against 0.0 to 0.3% in the prose references. Every prose text is below the level that a shuffle of its words gives, because writers avoid immediate repetition, while the Voynich text is 2.7 times above that level. Against a shuffle of the words within each line, though, it is at chance, for a reason the next sentences give. Adjacent words one edit apart (one glyph added, removed or changed) make up 3.6% of pairs, where the languages have at most 2.0%. There are also 125 runs of three or more near-identical consecutive words, such as qokeedy qokeedy qokedy qokedy qokeedy. Page vocabularies alone would give 70 such runs, and the references have 2 to 28. The repetition is a property of what a line contains, and it is not consecutive copying. A quarter of all lines hold a repeated word, more than twice the chance rate, but within a line the copies are no more often adjacent than a random arrangement of that line's words would make them. Adjacent words are also more similar to each other than to random words from the same page. The measure is normalised edit distance, the number of glyph changes needed to turn one word into the other, scaled to run from zero, for identical words, to one, for words with nothing in common. Adjacent words are at 0.775, random words from the same page at 0.784, and the corpus average at 0.811. In every reference text measured this way, adjacent words are less similar than random words from the same page. The word directly above is closer still (0.761), and the effect persists two lines up. The word boxes on the scans of 33 pages, the rectangles that mark where each word is written, allow a finer test. Measured from them, the word physically above is no closer than two other candidates, the word at the same index in the line above and the word at the same fraction of the way along that line. Each is 0.016 closer than a shuffled word of the same page, an excess that corresponds to about 2 percent of words being near-copies of the word above. A control with a planted 10 percent rate of physical copying gives 0.077, which the test detects at p 0.0003, where p is the chance of seeing an effect that large by accident. The excess is not tied to the word the eye sees directly above, which a scribe glancing at the line above would produce. It is equally present for the word at the same index, which an inherited layout would give, and at 2 percent it is too small to decide between the two.

Adjacent words tell a little about each other, and the amount is real but small. The measure is mutual information, which is how many bits knowing one word gives about the next. The text has 0.11 bits in excess of a shuffle that keeps the first and last word of every line in place, and that is the figure used throughout this report. Against a shuffle of the whole text the excess is 0.18 bits, but that includes the pairs that straddle line breaks. Most of it comes from the different vocabulary of words at the line edges and not from word order. The excess is at the bottom of the language range on the same measure (Latin epic 0.12, all prose 0.32 or more, the King James Bible 1.1). Between pages the continuity is small in absolute terms but real. Neighbouring pages share vocabulary only 0.002 above chance by overlap, where consecutive chapters of the Bible share 0.047. Yet for 17.9% of pages the most similar page in the whole book is the one before or after it (15.0% on Takahashi's reading, the other main transcription). Chance alone would give 0.9%, and 3.4% when the scribal hand (which scribe wrote the page) and the regime are held fixed. This is the statistic with which Reddy and Knight (2011) first showed that the pages are not in random order (they found 15.6%). It is in the middle of the language range on the same measure (Virgil 1.9%, Caesar 9.7%, the King James Bible 11.1%, the Hebrew Bible 28.5%, the Kalevala 37.2%). Weighting rare words more (the tf-idf weighting) lowers the Voynich figure to 13.0% and raises the languages', so the resemblance between neighbouring pages comes from frequent words, unlike the continuity of topic in a text. Page-to-page vocabulary overlap tracks the Currier regime (p = 0.005) and the scribal hand (p = 0.005), where p is the chance of seeing the pattern with no real link. The illustration section did not reach significance in that test (p = 0.08), which does not show that the section has no effect, and a stratified test reported below finds section vocabularies beyond hand and regime.

A null glyph at the ends of words, a meaningless sign added to each word, does not hide the repeats either. Two parses of the slot grammar strip every word to its core (the middle slots of the template), or to its core with one affix (a slot at either end). Neither stripped text has a repeated 25- or 30-glyph string. As a control, Latin with a null glyph added at each word end recovers its 132 repeated 30-glyph strings once the ends are dropped, which shows that the stripping can recover repeats when they are there.

The findings of this section, then, are three. The text has a stock of words that recur the way a language's words recur, but it has a local repetition that no language shows, and it has none of the repeated phrases that every language shows. So the word sequence is not the word sequence of a text in any of the sampled languages under any fixed mapping.

Earlier work. Zandbergen's survey notes that the text curiously lacks repeated phrases of two or more words, and Timm (2014) counted them. He found 35 sequences of three or more words occurring at least three times, alongside 318 cases of a single word repeated in sequence, and both findings hold here. What is new is the form of the statement. Counting repeated strings of glyphs with all spaces removed makes the fact immune to the word segmentation and to any fixed encoding. The same count on sequences of letter bags (each word reduced to its letters, with their order ignored) makes it immune to anagramming, the reshuffling of letters within words. Timm (2014) and Timm and Schinner (2020) built the self-citation account on local near-copying, an account in which the scribe made each new word by copying and altering words already on the page. The tests confirm that near-copying, with one correction: the excess similarity is flat across the preceding words and the line above, and it is not concentrated on the word just written. Timm (2014) also noted by eye that similar words lie one above the other on some pages, and he counted similar words at the same position in the preceding lines. The physically-above test above separates the physical column from the index position and finds the two equal.

Order at longer range: networks, bursts and drift

Several published studies describe how the words are ordered at distances longer than a phrase. This examination re-ran each of them against the same set of comparison texts, which puts the published claims and the measurements made here on one footing. The first is the network study of Amancio and colleagues (2013). They drew the text as a network in which each distinct word is a point and two points are joined when one word follows the other somewhere in the text. They found that the shape of that network is compatible with languages. Their exception was burstiness, which is how much a word's occurrences bunch together instead of spreading evenly through the text. On that measure the text's frequent words were not compatible with languages. The network half of that finding holds here, as the three measures below show. Each is a ratio to the text's own word shuffle, the same words in random order, which isolates the order of the words from their stock. The clustering coefficient, which is how often two neighbours of a word are also neighbours of each other, is 0.815. The mean path length, the average number of links between two words, is 1.053, and the diameter, the longest of the shortest paths, is 1.14. All three are inside the ranges of 14 languages cut to the same size (0.59 to 1.17, 0.97 to 1.21, 0.87 to 1.91).

The measures that fall outside the language ranges all point the same way, which is that the text's network is closer to its own shuffle than any language's is. The text has the smallest shortfall of distinct word pairs of any text, 0.945 of the count in its shuffle, where the languages have 0.63 to 0.947. Its frequent words are also less choosy about their neighbours, at 0.87 of what the shuffle would give, where the languages have 0.48 to 0.85. The template generators, meaningless generators that fill a fixed word template, give 1.0, so on this measure the text lies between the languages and a text with no word order. The clearest case is the all-mutual triangle, three words that each follow each other, which occurs at 0.52 of the shuffle count. In every language the figure is 0.01 to 0.10, and for Latin passed through a codebook, a code that replaces each Latin word by a fixed code word, it is 0.03. The triangle count reflects a property of the word order that does exist: it is nearly symmetric, since a word that tends to follow another also tends to precede it. A syntax does not predict such symmetry, but the word-boundary rule found by the chain-rule decomposition described below does. That decomposition splits the information between neighbouring words among the parts of the word. It finds the information at the boundary, in a rule about the glyphs on either side of the space.

The burstiness half of the finding does not hold here: the text's value is ordinary. The measure of burstiness is the coefficient of variation of the gaps between one occurrence of a word and the next. That is the spread of those gaps divided by their average, taken over the fifty commonest words. The text scores 1.71, which is 1.72 times its shuffle, and the 28 languages span 1.09 to 2.46. Six of them are above the text. They are Basque encyclopedia prose at 2.46, modern Greek prose at 2.06, Tagalog at 2.00, the Kalevala at 1.96, a Turkish scripture at 1.94 and the Hebrew Bible at 1.74. Narratives of a single genre are at the bottom (Virgil 1.09, Dante 1.09, Homer 1.11, Caesar 1.15). The published study measured its incompatibility against one book in fifteen translations, a uniform set, and that is where the disagreement comes from. Against a set that includes mixed texts the value is ordinary. Half of the text's burstiness comes from its two writing regimes, Currier A and B, the writing styles that Prescott Currier identified: inside Currier B alone the figure is 1.49. A Markov imitation, a meaningless text generated glyph by glyph from the text's own glyph statistics with no word structure at all, makes the point another way. Generated regime by regime, it gives 1.31 for the whole text and exactly 1.00 inside B, which is burstiness produced by the change of regime alone. The drifting-state generator, described below, in which a hidden state that changes slowly from page to page chooses the stock of words, is too bursty, at 2.73. The template, Markov and tree generators, three other meaningless generators described under generative models below, are not bursty at all (1.0).

Two published long-range findings do hold. Schinner (2007) reported correlations in the word sequence that extend beyond about 72 characters, and read them as the mark of a stochastic process, one driven by chance. Gaskell and Bowern (2022) noted that consecutive word lengths are positively correlated in this text, and that in most languages they are not. The test here for Schinner's correlations is detrended fluctuation analysis, which measures how fast the fluctuations of a series grow as longer stretches of it are taken. Its exponent is one half for a series with no memory, and higher when the series has long-range order. On the series of word lengths and of word ranks the text gives an exponent of 0.67 (standard deviation 0.015; 0.50 for the shuffled text). The exponent rises from 0.58 to 0.60 at short scales to 0.77 to 0.87 at long ones. It is inside the language range, in its upper third, so the long-range order Schinner described is there and is of a size that languages also reach. The second finding holds as well: the correlation between the lengths of one word and the next is +0.14, where 22 of 30 languages give a negative value. Both fluctuation statistics are weak instruments for structure at the scale of a page, however. A vocabulary shift planted on half of the pages leaves the exponent unchanged. The intermittency of the frequent words, which is their burstiness, registers a shift only when it covers 14 to 20 percent of the pages.

The correlation depends entirely on which lines lie near which, as shuffling the text at different scales shows. Shuffling the words within each line leaves every exponent unchanged, and so does shuffling the order of the pages. Shuffling the order of the lines, on the other hand, removes almost all of it (0.53). Generators with no memory give 0.48 to 0.53. The local-copying generator, which derives each new word from a recent word by random changes, has the excess only below 256 words. The drifting-state generator, however, reproduces the exponent and its change of slope, with 0.64 at short scales and 0.70 to 0.79 at long ones. So the exponent does not separate a language from a drifting process with no language in it, because both give it.

Two claims concern the distance from a word to the next similar word, which here means a word identical to it or one glyph edit away. Schinner proposed that the distances follow a geometric law, the law of a memoryless process, in which the chance of meeting the next similar word is the same at every step. Timm (2016) proposed that their distribution is incompatible with natural languages. The pooled distribution of the text is not geometric. But neither is any language's, and neither is the shuffled text's, because each of them is a mixture of many words with different rates. The comparison that matters is therefore with the text's own shuffle. The instrument is the hazard ratio, which is the rate at which the next similar word turns up at a given distance, divided by the rate in the shuffled text. Against its shuffle the text has an excess of similar words within the next 20 words. The hazard ratios are 2.06 at distance 1, 1.75 at 2 to 5 and 1.30 at 6 to 20. Beyond that it is exactly at chance (1.02, 0.98), and every one of those summaries lies inside the 30-language range, so what is unusual is where the excess falls. Languages avoid a similar word at distance 1, with a median hazard ratio of 0.56, whereas the text has 2.06, which only Hebrew, Finnish, Indonesian and modern Greek match. Languages also keep an excess out to 21 to 100 words, where all 30 languages give 1.05 to 2.09 and the text gives 1.02, the lowest value. So the near-repeats are more local than in any language sample. That is the finding above, that a word resembles its neighbour and the word above it more than chance allows, restated as a law of distance. Of the generators, the drifting-state generator's profile (1.96, 1.71, 1.16, 0.99) comes closest.

Landini (2001) found periodicities in the glyph stream with the spaces removed, and read them as a sign of linguistic organisation. The periodicity at the length of a word is there, but it has a plainer explanation. A spectral peak marks a glyph that recurs at a regular spacing. The glyphs e, i, y and k show such peaks, and so do the bench glyphs (the shapes transcribed as ch and sh) and the class of word-final glyphs. The peaks are 1.6 to 2.4 times the baseline, at a period of 5.5 to 6.0 glyphs. Since the mean word is 4.99 letters long, that period is close to the length of a word. The within-word shuffle, which scrambles the glyphs inside each word, never exceeds 1.37, and no language exceeds 1.5 on any of its ten commonest letters. The Markov imitation, by contrast, reproduces the peaks at 70 to 90% of their height. There is no component at the scale of a line: that band is at 0.5 to 1.5 times the baseline, within the range of the line shuffles. So the peak is the rigid word shape seen without the spaces, and the line, which shapes the words at its edges, has no fixed period of its own. Zelinka and colleagues (2023) mapped how central each word is in the word network, window by window along the text. Those maps also differ from the shuffled text, by a root-mean-square of 0.023. But the drifting-state generator (0.009) and the Markov imitation (0.012) are closer to the text's map than any language is (0.013 to 0.056). So the maps do not tell a language from a generator.

Earlier work. Amancio and colleagues (2013) reported that the manuscript is compatible with languages in the shape of its word-adjacency network, and incompatible in the intermittency of its frequent words. The network result holds against the reference set used here. The intermittency result does not, because the published incompatibility comes from their reference set and is not a property of the manuscript. The long-range correlations that Schinner (2007) described from a random-walk mapping, and that Landini (2001) read from a spectral analysis, are present. A word-length periodicity can be recovered from the spaceless stream, although no line-scale component can. Schinner's geometric law for the distance to the next similar word holds beyond about two lines, and inside them there is an excess of near repeats. The claim in Timm's (2016) preprint, that the distribution of similarly spelled words is incompatible with natural languages, is not reproduced. The centrality and fractal statistics of Zelinka and colleagues (2023) behave as published. They do not separate a language from a generator, though, because a drifting generator and a Markov imitation lie closer to the manuscript's map than any language does.

Position on the page changes the glyphs

In a written language the first word of a line is an ordinary word that happens to fall there. In this text it comes from a different stock of words. The first letters of line-first words differ from the first letters of interior words by 0.181 bits of Jensen–Shannon divergence. That divergence is a measure of the distance between two distributions, and it is zero when they are the same. Breaking Latin, Italian or English text into lines of the same lengths gives 0.001 to 0.002, and even verse, where the lines are real units, gives only 0.038, in Dante. The effect is exactly one word deep: the second word of a line is already ordinary. The last word of a line shows the same thing at its end (0.082 bits), and of the line-final words, 15.5% end in m, against 1% elsewhere. Line-first words are also longer, at 5.48 letters against 4.97.

Words at the start of a line and inside it

Counts of the commonest words by position, then the share of words (in percent) beginning with each glyph by position

RankFirst in the lineLinesNot first in the lineWords
1daiin161daiin655
2y62ol498
3saiin56chedy487
4dain49aiin458
5dar38shedy417
6sol37chol358
7sain35or333
8or33chey330
9qokeey32ar326
10ol32s294
First glyphAt line startAt paragraph startInside the line
p9.144.70.5
t9.422.81.9
s10.31.62.5
y14.71.83.7
f0.84.20.2
d15.52.28.6
q12.94.215.5
k2.510.43.6
o13.75.121.5
sh4.71.19.3
l1.70.34.4
ch3.50.417.8
a0.30.05.8
The upper table gives the ten commonest words at the start of a line and the ten commonest words elsewhere in the line, counted over the 4,129 lines of running text. The lower table gives the share of words beginning with each glyph, for the 4,129 words that open a line, the 740 that open a paragraph and the 30,651 that stand inside a line. It lists the glyphs with at least two percent in some column, in order of how strongly they favour the line start, and it counts ch, sh and the bench gallows (cth, ckh, cph, cfh) as one glyph each. The report singles out s at line starts and p and f at paragraph starts. Prescott Currier observed in 1976 that the first word of a line is written differently, and Vogt documented it in 2012.

The line as a unit

A schematic page shows the first word of each line, the last word fitted to the margin, and the line rule applied again after a drawing.

margin m m drawing 1 First word of a line Its first glyph is drawn from a different stock from that of the words inside the line. The divergence, which measures how far the two distributions differ, is 0.181 bits. Latin, Italian or English cut into lines of the same lengths give 0.001 to 0.002. The effect is one word deep. Words found mostly at line starts include sain, sor, dshedy, tol. The first word of a paragraph begins with p in 46.5% of cases (the tall mark on line 1). 2 After a drawing Where a stem or leaf interrupts the line, the word written after it starts like a line-first word. On the hand-1 herbal pages the divergence is 0.130 bits, against 0.017 at ordinary word boundaries of the same lines and 0.166 at true line starts. The line-end habit does not cross with it: the word before the drawing does not take the line-end form. 3 Last word of a line It is fitted to the margin by the choice of a shorter word. On pages with a free right margin the space left at the line end is 0.56 of a glyph, against 1.05 to 1.10 for a simulated copyist filling the same lines with the same words. The last word has 4.81 letters against 4.94 to 5.12. 15.5% of line-final words end in m (am, dam, otam), against 1% elsewhere.
The line as a unit, on a schematic page. The first word of every line (dark bars) is written differently from the words inside the line, and the difference is one word deep. The last word (brown bars) is fitted to the margin by the choice of a shorter word, and it ends in m far more often than a word inside the line. Where a drawing interrupts a line, the word written after it starts like a line-first word. The line-end habit stays with the written line. The bars are schematic. The figures are measured on the whole running text and, for the margin, on the scans.

The left ends of four consecutive lines, showing the line-initial words

The left ends of four consecutive lines, showing the line-initial words

daiin,sheeky okeey okeey qaal shedy okeey oteey shedy chcthy lcheol oteeam / dsheedy lkeedy chckhy lchedy qokeey qokear chal qokeear cheokedy sal lokam / saiin oteedy qokeey daiin okedal chedy qokedy l,shedy chctchy okeey lor ar am / saiin sheekshy ol,shedy chokchey lkey otain

The left ends of four consecutive lines on f. 111r (Currier B, hand 3). The lines begin daiin, dsheedy, saiin and saiin. Across the text, the first word of a line is drawn from a different stock of forms from the words inside the line. The difference is one word deep, and the second word of a line is already ordinary. Yale University Library, Beinecke MS 408, f. 111r (detail)

These edge words are mostly ordinary words with a positional element added, which can be shown by taking the element away. Take a line-initial word that begins with s and strip that first letter: the remainder is an attested word, one that occurs elsewhere in the text, 61% of the time. For an interior word beginning with s the figure is only 10%. Replace a line-final m with n, in, iin, l or r, and the result is an attested word 85% of the time. One caveat comes from the hand-by-hand test below: the n replacement alone is attested no more often at the line end than inside the line. Paragraphs are marked in the same way: of the words that open a paragraph, 46.5% begin with p. The tall glyphs p and f make up 4.2% of the letters on the first line of a paragraph, against 0.27% on all other lines. That fifteen-fold difference is spread over the whole first line and is not confined to the first word. This matters for any reading: a glyph whose occurrence depends on whether the line opens a paragraph cannot stand for the same content as a glyph with no such dependence.

The line is a closed unit. The last word of one line tells nothing about the first word of the next. Their mutual information, which is how many bits knowing one word gives about the other, is 0.008 ± 0.013 bits. For neighbouring words inside a line it is 0.129 on the same statistic, whose chance level comes from permuting the second word of each pair among the pairs of the same page. The 0.11 figure used elsewhere in this report takes its chance level instead from a shuffle that keeps the line edges in place. The glyphs on either side of the break show the same thing, at −0.003 bits against 0.185 inside the line. That holds in both transcriptions and in both writing regimes, Currier A and B. The comparison with real writing uses control texts "poured into" the manuscript, which means written out into the manuscript's own line lengths. Every language poured into the same lines keeps 33 to 62% of its in-line word association across the break, and all of its glyph association. A word-level Markov generator picks each word from the probabilities of the word before it. Given a line-start state, so that the first word of a line has its own probabilities, it gives exactly the Voynich profile. The test had an 80 percent chance of detecting a cross-line dependence above 0.2 of the in-line level, and the languages are at 0.33 to 0.62. So the small word-order signal reported above is confined to the line, and the line break is not the wrapping of a continuous text. This measurement has no precedent in the reviewed literature, although a 2026 preprint states that the cross-line information is exactly zero, and the measurement here agrees with it within noise.

Gradients across the line, and habits of the hands. Feaster (2022) observed that when two word forms differ in one glyph, one of them tends to stand further to the right in the line. The ch-forms stand to the right of the sh-forms, the o-forms to the right of the qo-forms, and the t-forms to the right of the k-forms. That holds here for 95.7% of the ch/sh pairs and 86.9% of the o/qo pairs, where he reported about 90% and 83%. It holds for the frequent t/k pairs, five of five with 100 tokens or more, although it is not significant over all such pairs. A shuffle within lines gives 0.48 to 0.51, and the Markov imitation gives 0.51 to 0.53, because its variant forms fall at arbitrary positions. The effect survives the removal of the first and last word of every line (0.89, 0.73 and 0.69). It is therefore a gradient across the whole line, separate from the edge effect described above. It also runs downwards. The ch-forms stand lower in the paragraph than the sh-forms in 83% of pairs, and the t-forms stand higher than the k-forms, which is the known preference of the tall glyphs for top lines. No generator that ignores lines produces it, and the drifting-state generator below is one of those. The scribal hands differ at the start of a paragraph, as Stafford (2022) reported, and the published shares are reproduced here. The glyph p opens 34.8, 58.1 and 53.1% of paragraphs for hands 1, 2 and 3, against the published 34.9, 53.9 and 52.1. Hand 1 differs from hands 2 and 3 beyond a null that holds the section fixed, mainly in opening more paragraphs with k and f, whereas hands 2 and 3 do not differ from each other. The glyphs that open paragraphs are distributed unlike those that open lines or words, with a divergence of 0.5 to 0.7 bits against 0.1 to 0.3. The interaction between a glyph's position in its word and the line's position in the paragraph is small, however. Cramér's V is a measure of how strongly two categories are linked, from zero for no link to one for a complete one. Here it is 0.043, against a permutation 95th percentile of 0.041, and the last line of a paragraph is the only cell of that table with a readable effect.

One pattern proposed on the Cipher Mysteries blog is not there. The proposal is that pairs of single-leg gallows glyphs, the tall glyphs drawn with one leg, stand on the first line of a paragraph about two thirds of the way along, as a kind of key. Had such pairs occurred on more than 10 percent of the paragraph first lines, the test had an 80 percent chance of finding them. What it found runs against the proposal. First lines hold two or more p/f words in 56.6% of paragraphs, which is fewer than random placement at the same rate would give (61.5%). Those words stand left of the line's centre, at a mean relative position of 0.450, below all 200 within-line permutations. The first of them is usually the paragraph's first word, and how often they stand next to each other is at the permutation null (38% against 40%). What reads as a marker, then, is the excess of tall glyphs on first lines described above, combined with the paragraph-initial one.

The line effects are not a scribal convention, a habit of shaping the first and last letters of a line, as a comparison with real medieval scribes shows. The diplomatic manuscripts are medieval manuscripts transcribed sign by sign, with their allographs, which are variant shapes of one letter, and their abbreviation marks kept. Re-parsed at their real line breaks and run through the same code, they show line-initial and line-final symbol effects of 0.002 to 0.036 bits, against the text's 0.176 and 0.083. The computation at the head of this subsection gives 0.181 and 0.082 for the same two effects. Only the verse Kaiserchronik (0.064 and 0.103) approaches the text's order of size, and its length profile is the opposite of the text's. No manuscript prefers a paragraph-first symbol the way the text prefers p (45 percent against 1.3).

The line rules in every hand. The edge effects are not one scribe's habit. The five scribal hands are the writers that Lisa Fagin Davis told apart by their letterforms, and the effects appear in every one of them. In each hand taken alone, the first letters of line-first words diverge from those of interior words by 0.15 to 0.35 bits. The gallows glyphs, the tall glyphs, open the first line of a paragraph at 4 to 18 times their rate on other lines. Stripping the first letter of a line-initial word raises the share that is an attested word by 0.10 to 0.25. The differences between the hands at the line edges (0.04 to 0.06 bits) are no larger than the differences between the same hands in their interior vocabulary (0.02 to 0.12). So the rule is shared: the hands differ in what they write and agree in how they treat the edge. Two details do depend on the hand: hand 1 alone shows no shorter last word, and the adjacent-word information of 0.11 bits comes from hands 1 to 3 (0.06, 0.12 and 0.09 bits). Hands 4 and 5 give 0.003 ± 0.009 on their 440 and 779 word pairs, enough pairs to have shown an effect of hand 1's size. One half of the line-final finding also fails to survive the split by hand. The n form of a line-final m-word is attested no more often than the n form of an interior m-word in any hand (0.20 to 0.37 against 0.21 to 0.36). So the 85 percent above is a property of m-words and not of the line end. The register of the labels, which begin with o and avoid q, is shared by hands 1, 2 and 4, so it is not a habit of hand 4, whose pages are mostly labels. The rule is also partly spatial: where a drawing interrupts a line on the herbal pages, the word written after the pen crossed the drawing starts like a line-first word (below).

Herbal text lines cut by the plant stem, each resuming on its far side

Herbal text lines cut by the plant stem, each resuming on its far side
Four lines on f. 8r (herbal section, Currier A, hand 1) are cut by the stem of the plant. Each line stops at the drawing and resumes on its far side. In EVA, with the gap marked by a vertical bar, the four lines read as follows. First line: tchty sh,kcheals sho | okche do dchy dain al. Second line: chodar shy sy | chodaiin shokchy chor dy. Third line: qotor chor chor sheey | dchol shesed chof chy dam. Fourth line: okchey do r cheeey dy ky | scho chky ckooaiin chy,taiin. On the herbal pages of this hand, the word written after the pen crosses a drawing starts like the first word of a line. These four lines are one instance. Yale University Library, Beinecke MS 408, f. 8r (detail)

Earlier work. Currier (1976) observed that the line is a functional unit, and Vogt (2012) developed the observation. Timm (2014) added word length by position in the line, and Smith (2015) the stripping and adding of glyphs at line starts. All of it is reproduced here, and the derivation test gets a number, from comparing the attestation rate of stripped line-initial words with a matched baseline from the interior of the line. Feaster's (2022) rightward and downward gradients are reproduced for almost every frequent minimal pair, and they survive the removal of line-edge words. Stafford's (2022) per-scribe paragraph-initial shares are reproduced too, although the interaction proposed there between position inside the word and position in the paragraph is at its permutation null within each hand. The paired single-leg gallows known from Pelling's Cipher Mysteries blog as Neal keys are not there: gallows words stand left of centre on top lines, and their pairing and adjacency are at chance. The information a line break gives about the words on either side of it, measured here against the information inside the line, has no earlier application to this text. It is close to zero, and that is what makes the line a boundary. A voynich.ninja forum thread of 2019 examined the word after a drawing crossing, with the same reading, that the line rule re-applies at the crossing. No precedent was found for the per-hand results above. The line-position controls in earlier work are natural texts cut at an artificial width (Vogt at 62 characters, Gaskell and Bowern at 60), whereas the diplomatic manuscripts above are re-parsed at their own line breaks.

Two regimes, five hands

The transcribers split the pages into Currier "languages" A and B, the two writing styles that Prescott Currier identified. That split is real and binary, and it is a property of the glyphs, as a classifier shows. A logistic-regression classifier is a standard statistical rule that weighs several measurements to make a yes-or-no decision. Here it uses the frequencies of 18 letters over all the words of a page. Under leave-one-out testing, in which each page is classified by a rule fitted to all the other pages, it assigns 98.5% of the 197 labelled pages correctly. It scores 97.5% on the running text alone, 94 to 97% with other linear classifiers and 99% with word-final units, and only five pages are intermediate. The two regimes differ in glyph statistics far more than any two halves of a single reference text do. The divergence between their letter distributions is 0.017, against at most 0.005 for a text split in half, and the Old and New Testament halves of the Bible give 0.001. They also differ in letter-level predictability, by 0.15 to 0.19 bits depending on the symbol view, which is the choice of which glyph groups count as one symbol. Yet they share their vocabulary statistics: the Zipf slope, the growth of the vocabulary, the hapax rate and the word entropy cannot be told apart. The Zipf slope is how fast word frequencies fall from the commonest word down the ranks, and the hapax rate is the share of words that occur once. Three diagnostics do tell the regimes apart. Words ending dy are 6.6% of A and 23.8% of B, the letter e is 6.8% of A's letters and 12.0% of B's, and words ending ol are 14.2% of A and 7.1% of B.

Words characteristic of Currier A and of Currier B

Occurrences per thousand words of running text in each regime, and the ratio of the larger rate to the smaller

WordCurrier ACurrier BRatio
Characteristic of Currier A
otchol2.50.057.6×
dchy2.10.049.4×
dchor2.10.047.3×
cthor3.70.142.2×
cthol4.60.135.7×
cthy8.40.332.2×
ckhey2.30.126.8×
kchy2.70.215.4×
sho9.20.615.1×
kchol2.00.115.1×
Characteristic of Currier B
qokeedy0.013.3298.3×
chedy0.121.3238.6×
otedy0.06.1136.0×
qokedy0.111.9133.6×
okedy0.04.7105.9×
shedy0.218.3102.3×
qotedy0.03.884.5×
qoteedy0.03.373.9×
keedy0.02.760.2×
lchedy0.15.055.9×
The ten words most characteristic of each of Currier's two writing regimes. For each word the table gives its rate per thousand words in the running text of Currier A (11,194 words) and of Currier B (23,039 words), and the ratio of the two rates. Only words that occur at least 20 times in the regime they characterise are considered. A word absent from the other regime is given half an occurrence there, so that the ratio stays finite. The 547 words on pages without a regime code are left out. Prescott Currier identified the two regimes in 1976.

The split coincides with the scribal hands recorded in the transcription, the five writers that Fagin Davis told apart by their letterforms. Hand 1 wrote 112 of the 114 Currier A page sides and no B page, and the other two A sides are in hand 3. Hands 2, 3 and 5 wrote B, while hand 4 wrote only astronomical pages that are mostly labels. The regime therefore follows the writer more closely than the subject. Herbal pages exist in both regimes, 95 in A and 32 in B. Yet every A page side but two is in hand 1, and those two are the sides of one leaf, f58, in hand 3. Hand and section are partly confounded, because hand 1 wrote most of the herbal pages, so the text alone cannot fully separate their effects on the glyph-level regime. At the level of vocabulary, a stratified test, which compares pages only with pages of the same hand and regime, does find a section effect beyond hand and regime. Whether the regime difference reflects two conventions, two "dialects" of the system, or two stages of one system cannot be decided from the text.

The two writing regimes

Herbal page in Currier A by hand 110r · Herbal · Currier A · hand 1
Herbal page in Currier B by hand 239r · Herbal · Currier B · hand 2
Two herbal pages, one from each of Currier's two writing styles. On the left is f. 10r, in Currier A and written by hand 1. On the right is f. 39r, in Currier B and written by hand 2. Both pages show a plant. The style goes with the scribe. Yale University Library, Beinecke MS 408

The glyphs show two regimes more clearly than five scribes. Farrugia, Layfield and van der Plas (2022) attributed pages to the scribal hands with classifiers trained on sequences of characters, and that is reproduced here. Hand 1 is recovered with an F1 of 0.98, a score that combines how many of a hand's pages the classifier finds with how many of its calls are right. Five of the seven pages on which they disagreed with the palaeographic assignment are misclassified here too. But the hand signal in the text is almost entirely the signal of regime and section, as a permutation of the labels shows. Permuting the hand labels within regime and section gives each page another page's hand label while it keeps its own regime and section. With the labels permuted that way the classifier still scores 0.852, against 0.864 with the true labels. Inside Currier B alone, where hands 2, 3 and 5 all write, the balanced accuracy is 0.60 to 0.66 against a null of 0.55 (0.85 with a support-vector classifier on 81 pages). Balanced accuracy is the accuracy averaged over the hands, so a large hand cannot dominate the score. Burrows's Delta, the standard measure of stylistic distance between two texts, puts hands 2 and 3 only 0.31 apart. That is near the floor, the distance between two halves of one hand (0.18 to 0.21), while every distance to hand 1 is 0.69. Hands 4 and 5 are too small and too mixed to place. The positive controls show that two habits of one scribe's text would be separable at 96 to 97% from page-sized samples. So the glyph statistics support two regimes strongly and five hands weakly. The five hands rest on the letterforms, as Fagin Davis (2020) established them, and not on what the hands wrote.

Arutyunov and colleagues argued in a 2016 preprint that the text mixes two languages, on the evidence of how far the sorted glyph frequencies depart from a logarithmic law. That criterion has no power: of 40 single-language samples, 35 fail it, real Latin-Italian mixtures do not lower the fit, and the Voynich value (0.95 to 0.96) is inside the language range (0.75 to 0.99). A direct mixture test says something else. The glyph frequencies vary from page to page five to ten times as much as in any single-language text, and about 0.6 as much as in a text that alternates Latin and Italian. But the variation persists inside Currier A, inside Currier B and inside each hand, and a third component keeps improving the fit, which a genuine mixture of two languages would not show. Currier A and B are the first split of the glyph frequencies, with a purity of 0.88 in the fine alphabet, the transcription that keeps the most glyph distinctions. Purity is the share of pages that the split puts on the side of their own regime. Beneath that split the variation is graded page by page, and that is the pattern the drifting-state generator was built for.

The divergence between A and B is within a factor of two of the variation between two hands writing the same text. Two manuscripts of one text, transcribed at the diplomatic level, differ in their letter distributions by 0.011 to 0.012 bits, and their split halves by 0.002 to 0.006. Currier A and B differ by 0.020 by the same code, within a factor of two of scribal variation. At the word level, on the other hand, the divergence between A and B (0.495) exceeds both manuscript pairs (0.38 and 0.41). The difference in glyph predictability between A and B (0.19 bits) is matched by contiguous halves of B alone, so it is not quoted here as a mark of the regimes.

The regime travels with the sheet. The physical unit of the book is the bifolium, a sheet folded once to give two leaves, and the regime follows it. Surprisal is how unexpected each glyph of a page is to a model fitted on the other pages. Pages of one bifolium are alike in how predictable their text is, beyond what hand and section account for: the bifolium explains 56 percent of the variance of the held-out surprisal inside those strata, against 22 percent for pages shuffled among bifolia (p 0.005). Three generators run into the same pages reach at most 29 percent, and one of them, the local-copy generator, reaches p 0.03 on its own null at about half the manuscript's size. Four real texts poured into the same pages reach 29 percent for the Latin and 20 for the Italian, but 44 for the German and 44 for the English, both at p 0.005, so two ordinary books with none of the manuscript's writing in them show an effect of the same kind, and this test does not separate the sheet from the place in the book. The 22 percent is the null mean of the saved run. In another run of the same test it was 0.224 in place of 0.217, because the scripts were run without a fixed hash seed, and the saved run is the one reported here. Within one hand and section, the letter distribution steps more at a change of sheet than at the turn of a leaf (0.017 against 0.012 bits, p 0.013). None of seven control texts poured into the same skeleton of pages, the manuscript's own layout of pages and lines, shows that. Consecutive pages on one leaf also share more vocabulary (0.146) than consecutive pages on different sheets (0.131, a difference of 0.015 with a standard error of about 0.009). That vocabulary version of the step is weaker than the letter version. On the step test's own sample of 24 leaf and 16 sheet boundaries it is 0.148 against 0.130, at p 0.062, and that is a small sample. The same seven controls were scored on it, and none shows a leaf-over-sheet difference (sheet minus leaf -0.003 to +0.004, p 0.18 to 0.79). The drifting-state generator, whose state steps once per page, shares the same amount across either boundary (0.108 and 0.107). It also sits below the manuscript at every boundary type, 0.107 against 0.131 to 0.146, which the fit did not target, and a refit with its state stepping at the physical boundaries was not run, for want of compute. The quires that mix hands give a natural experiment, a quire being a gathering of bifolia bound together. A hand-2 bifolium bound inside a hand-1 quire is as far from its quire neighbours as from any other hand-1 bifolium. That holds on the Currier diagnostics, the glyph frequencies that tell A from B (p 0.005), and on vocabulary (p 0.011). So the regime travels with the sheet and its hand, and not with the quire the sheet was bound into. Five of the six shifts in ink darkness at page scale, counted over the 171 pages of the page-level series (below), fall at a sheet or quire boundary, against 2.8 of the six expected by chance (p 0.08). Four of those six shifts also coincide with a change of hand or of section, and the two that do not give p 0.58. The two shifts in ink hue fall at the turn of a leaf. The size of the darkness step is much the same at a sheet boundary as at a leaf boundary, 0.026 against 0.022, and the step at the turn from verso to recto is no larger than at the turn from recto to verso (p 0.24). So the placing of the change points is not established, and changes of hand and section account for most of it.

The hand labels rest on palaeography and on the letter statistics of the transcription. The letterforms on the aligned crops, the glyph images cut from the scans and matched to the transcription, give them weak independent support. A page classifier on gallows height, slant, descender length and pen lifts reaches a leave-one-page-out accuracy of 0.60 on 35 pages, against a permutation mean of 0.44 (p 0.04 to 0.05). For hand 1 against hand 2 inside the herbal section it reaches 0.82 against 0.50 (11 pages, p 0.03 to 0.05). On those features hand 2 leans right by about 10 degrees, writes a longer y tail and taller gallows, and lifts the pen more often. The support is weak because hands 2 and 3 are confounded with section, hand 5 has two aligned pages, and the heights of k and t separate no pair.

The verdicts inside each regime. The three measurements on which the verdict board's verdicts rest, the repeated-phrase counts, the glyph predictability and the spread of word lengths, were made on the whole running text. This examination therefore re-ran them, with the rest of the 23-statistic battery, on the Currier A pages alone (114 pages, 11,194 words) and on the Currier B pages alone (83 pages, 23,039 words). Each regime was run in its own page skeleton. The references were 697 samples cut to exactly those lengths from the language, medieval and list corpora, and meaningless imitations generated from each regime's own glyph statistics. No verdict reverses. In each regime the glyphs are more predictable than in any sample of the same length. The figures are 2.14 bits in A and 2.00 in B, against 2.68 for the nearest language and 2.48 for the nearest text of any kind. The spread of word lengths is 1.78 in both regimes, narrower than any alphabetic sample at either length (1.84 and 1.89), by a thin margin for A. The share of words used once is 0.69 in A and 0.68 in B. The one-edit density is the share of words within one glyph of a commoner word. It is 0.71 in A and 0.74 in B, against at most 0.60 and 0.65 for any text of the same length. A word comes back on the neighbouring line 3.1 times more often than chance in A and 3.2 times in B, against 4.6 for the whole text, and the imitations are flat. Page burstiness falls from 1.96 for the whole text to 1.44 in A and 1.60 in B, still above the imitations at 0.99 and 1.01. So a third to a half of the pooled clustering is the split of the vocabulary between the regimes, and the rest lies inside each. The word-order information is 0.066 bits in A and 0.128 in B, above the imitations at 0.008 and 0.051, and it comes mostly from B. One argument, the argument from repeats, needs the whole text: the repeated four-word count is 0 in A alone and 1 in B alone. At 23,039 words every language sample has at least 15 repeated four-word sequences, so the verdict holds inside B. At 11,194 words, however, the low-repeat prose samples have as few as 2, and a period list with its duplicated blocks removed can have none, so the A pages alone cannot support that argument. The statement that the text has one repeat where every sample has at least 67 therefore rests on the whole text. What separates the regimes is the glyph level. B is more predictable than A by 0.15 bits, and its words are longer on average (5.09 glyphs against 4.79). Its words also end in a narrower set of glyphs, and one word constrains the next more strongly. What the regimes share is the shape of the vocabulary, the spread of word lengths, the near-copying and the absence of long repeats.

Seven verdict statistics inside each Currier regime

Each row is one statistic on its own scale, once for the Currier A pages and once for the Currier B pages. The grey band is the range of the reference samples cut to that regime's length, the filled marker the regime's own value, the open circle a meaningless imitation of the same length, and the diamond the whole text. The repeated four-word count is on a logarithmic scale.

Glyph predictability (bits per glyph) below every sample at both lengths 2.68 A pages 2.14 2.67 B pages 2.00 2.0 2.5 3.0 3.5 Spread of word lengths (sd, glyphs) below every alphabetic sample at both lengths 1.84 A pages 1.78 1.89 B pages 1.78 2 3 4 Repeated four-word sequences within lines holds in B, but A alone is too short 2 A pages 0 15 B pages 1 0 1 10 100 1,000 Share of words used once inside the range at both lengths A pages 0.69 B pages 0.68 0.4 0.6 0.8 Page burstiness (1 = no clustering) inside the range, above the imitations A pages 1.44 B pages 1.60 2 4 Words one glyph from a commoner word (share) above every sample at both lengths 0.60 A pages 0.71 0.65 B pages 0.74 0.2 0.4 0.6 Word-order information (bits) inside the range, above the imitations A pages 0.07 B pages 0.13 0.0 0.5 1.0 Currier A pages Currier B pages whole text meaningless imitation at that length range of the reference samples cut to that length
Seven of the 23 statistics, measured inside each Currier regime. Each is set against samples cut to the same length from the reference corpora: 175 samples of 11,194 words for A and 150 of 23,039 for B. Each is also set against a meaningless imitation built from that regime's own glyph statistics. The glyph predictability stays below every sample at both lengths: 2.14 and 2.00 bits, against 2.68 and 2.67 for the nearest sample. So does the spread of word lengths, and the share of words one glyph from a commoner word stays above every sample. The share of words used once, the page burstiness and the word-order information sit inside the range at both lengths, the last two above the imitations. The repeated four-word count is 0 in A and 1 in B, against at least 2 and at least 15 in the samples. So that verdict holds inside B, and the A pages alone are too short to carry it. Bars through the markers are the page half-sample intervals.

Earlier work. The A and B regimes are Currier's (1976), and Zandbergen revisited them in page-by-page cluster analyses. That they coincide with the scribal hands, which Fagin Davis (2020) identified from the handwriting, is the standard account. Both are confirmed here by a leave-one-page-out classifier that places 98.5 percent of the labelled pages correctly, against the 89.2 percent reported for a held-out classifier in Parisel's 2026 preprint. Farrugia, Layfield and van der Plas (2022) classified pages into the five hands with character n-grams, which are short runs of characters. Their result is reproduced, including five of the seven pages on which their classifiers disagreed. It is qualified here, though, by the finding that most of what separates the hands is the A and B split. Arutyunov and colleagues (2016, a preprint) read the deviation of sorted letter frequencies from a logarithmic law as the sign of two sources, and that is not reproduced. With matched samples most single-language texts fail their criterion and real bilingual mixtures do not, so the log-law argument for a bilingual manuscript does not stand. The page-level heterogeneity behind it is real, however, and five to ten times that of any single-language text. The classification techniques here are established ones, and what they add is the stratified permutation that separates hand from section. That the regime follows the bifolium and not the quire has been argued before from three directions. Fagin Davis (2020) argued it from the palaeography, since she attributes the hands bifolium by bifolium and notes the mixed quires. Pelling (2022) argued it from glyph-pair densities mapped page by page across one quire, and proposed the bifolium as the unit. Parisel's 2026 preprint argued it from character-pair ratios across folio transitions, and finds the jump at a change of language within a quire larger than at a same-language transition. The boundary step test, the bifolium variance and the mixed-quire comparison above sort the boundaries by type within one hand and section, and add permutation nulls and poured controls to those observations. Hand attribution from the images has a precedent outside this manuscript in Popović, Dhali and Schomaker's (2021) analysis of the Great Isaiah Scroll, which located a change of writer at a sheet boundary. The letterform classifier above does the same on a small scale, on the aligned crops. Edwards (2025a) set the correlation of glyph frequencies between the A and B pages (84.8 percent) against medieval document pairs in one language (above 90) and in two (50 to 80). The calibration above uses two manuscripts of one text with a split-half floor instead.

The labels beside the drawings

The labels beside the drawings behave like names, and not like numbers or like ordinary text. The star labels are 83 distinct strings in 84 tokens, with no regularity between one label and the next. The 342 zodiac labels contain 278 distinct strings, of which only 37 recur on a second page. Labels begin with o 54% of the time, where the running text has 21%, and they almost never begin with q (1% against 15%). At the same time, 60% of the label words also occur somewhere in the running text.

The labels of page f101v

No.LabelElsewhere in the bookas running textas a label
1sairaly000
2otaldy742
3otol81654
4ytal14131
5dokor000
6orar770
7otarar440
8otoly843
9soraly000
10okol77646
11arom211
12oraram000
13oraeep000
14dytolg000
15olkor220
16dolary000
17odor770
18olaran000
The 18 labels written beside the drawings of page f101v (pharmaceutical section, Currier A, hand 1), in the order of the transcription. For each label, the table gives the number of times the same word occurs elsewhere in the book, as a word of the running text or as a label on another page. Where those two counts do not add up to the total, the remainder is text written around circular diagrams. 10 of the 18 occur elsewhere. Across the whole book, 623 of the 1,159 label words (54 percent) begin with o, the commonest first glyph of a label, against 21 percent of the words of the running text. René Zandbergen and, in 2017, Pelling noted that labels favour an initial o.

A jar and plant parts, each with a short label beside or above it

A jar and plant parts, each with a short label beside or above it
A jar and a row of plant parts on f. 99r (pharmaceutical section, Currier A, hand 1). Each has a short label written beside or above it. Over the whole book, labels are short. More than half of them begin with the glyph o, and three in five occur somewhere in the running text. Yale University Library, Beinecke MS 408, f. 99r (detail)

The labels as known plaintext. A crib is a piece of plaintext that is known or guessed, and a crib attack uses it to find the key. Suppose the plant labels name the plants drawn beside them and the star labels name the stars. Then the labels are the one place where a crib attack is possible in principle, because the candidate plaintexts, the words the labels might stand for, form a short list. This examination attacked the labels as a known-plaintext problem in four ways, and first showed that each way works on enciphered Latin names.

The first way used the substitution and homophonic solvers of the cipher section below. A substitution cipher replaces each letter by one fixed sign, and a homophonic cipher lets a letter be written as any of several signs. The tests ran the solvers on the label corpus alone: all 1,075 labels, and the 208 plant, 80 star and 340 zodiac labels separately, in two glyph views, against 25 language models. The crib corpus keeps only labels of at least two letters, which is why it counts 80 star and 340 zodiac labels where the census above counts 84 and 342. The solvers never score the labels better than a meaningless Markov imitation of the same labels beyond the solver's own scatter. Over 296 combinations of language and solver, the median advantage of the labels is 0.02 bits per glyph. In 4 cells the labels beat both imitation seeds by more than 0.1 bits, and in 27 cells the result goes the other way. The decoded labels stay above real text in every cell, by 0.4 bits at least and 2.0 in the median. The share of decoded labels within one edit of a vocabulary word is lower for the labels than for the imitation in 280 of 296 cells. The fault does not lie with the solvers, since the same solvers recover 99.9 percent of the characters of 970 enciphered Latin plant and star names under a simple key. Under a homophonic key of 36 symbols they recover 96 to 100 percent.

The second way was a direct dictionary search. It finds no simple substitution key that maps more than 5 of the 194 plant-label types onto the 876 plant-name forms in Pliny's Natural History, where the imitation and running-text controls reach 4, 4, 4 and 4. It finds none that maps more than 2 of the 79 star labels onto Latin star and calendar names (controls 4, 3, 4 and 3). Against the general vocabularies of Latin, Italian, French, German, Spanish, Greek, Hebrew and Arabic the labels never exceed their controls by more than the search's scatter. The search is sensitive enough to have found a key, since it maps 195 of 200 enciphered Latin plant names back onto the list.

The third way used statistics that no relettering can change, and they say why the searches fail. The labels have the length and the type-to-token ratio of a name list, where the type-to-token ratio is the number of distinct entries divided by the number of entries. But half or more of them begin with the same glyph: plant labels 50 percent, star labels 60, zodiac labels 67. Their first-glyph entropy, which measures how spread out the choice of first glyph is, comes to only 1.8 to 2.4 bits. No name list in any sampled language puts more than 27 percent on one initial (2.9 to 3.8 bits), and no one-to-one or homophonic relettering can remove that difference.

The fourth way was a crib test on the drawings, using a published compilation of plant identifications for 121 herbal pages (Sherwood, 111 with a Latin binomial). The test matched each page's label or first word against the binomial, its classical and vernacular synonyms and its English name, all under a single key. At most 4 pages are consistent with one key, against 3.1 ± 0.6 when the identifications are shuffled across the pages (200 shuffles). That last test has little power, because few pages offer a label and a name of compatible length and pattern, and it is reported for what it is. Taken together, the four ways show that the labels are not an enciphered name list of these languages under a simple or homophonic substitution. Whether they are names in some other notation is not decided by this test.

Nor are the labels running-text words with a grammatical layer removed. A proclitic is a short grammatical word that attaches to the front of the word after it. Stripping a proclitic layer (qo-, o-, y-, d-, s-) from the running text moves it away from the labels instead of towards them. The divergence of the initial glyphs is 0.155 to 0.441 bits, and the share of labels attested in the stripped text falls from 0.47 to 0.23. The same operation on the Hebrew Bible, by comparison, brings its text words to within 0.03 bits of its names.

The labels were also set against six real name lists of the period. The lists are the saints of the litanies, the names of the two martyrologies, the Durham Liber Vitae, the later York freemen roll, the plant names of the Sinonoma and the headwords of the Promptorium. The labels have the type-token ratio (0.72), the once-used share (0.83, the share of entries that occur only once) and the lengths of a name list. But 54 percent of them begin with one glyph (the zodiac labels 67 percent), where the six lists put 9 to 16 percent of their entries on their commonest initial. Windows cut from the lists average 14 percent or less, and none reaches 20. Of the label types, 37 percent lie one edit from a commoner type, against 7 to 14 percent in the lists. And 60 percent of the labels are attested in the running text, against 4 to 17 percent for real lists in prose (the saints in the Vulgate 51). So the labels have the size statistics of names and the shape statistics of the text's own words.

The twelve zodiac rings are not one list of names written under twelve alphabets, which is the arrangement a cryptanalyst would test first. The best single relettering maps 3.6 labels of one ring onto another, where the chance level is 3.4 to 4.0, whereas a planted list under twelve keys, with a fifth of its labels damaged, gives 19.2.

Earlier work. Zandbergen's writing-system pages describe how labels differ from running text. So does blog and forum work, above all Pelling's (2017) writing on labelese, the name given to the way the labels are written, and the forum discussion of the zodiac labels. Three observations are established there and reproduced here: labels favour an initial o, they almost never begin with q, and most of them also occur somewhere in the running text. What is added is the null model. This examination measures the cross-page recurrence of label types against shuffles of the label tokens, against draws from the running text and against Markov draws. It also measures the similarity of consecutive labels on a page against random pairs from the same page. That turns the informal reading of the labels as a numbering system into a test, and the reading fails it. No technique in this section is new in substance, and the contribution is that established observations are given controls, one of which, the consecutive-label similarity test, was not found in that form in earlier work. The Computational Attacks blog (JB 2011, 2016) compared labels with real name lists: with fifteenth-century female names on length and endings, and with the stone names of Alfonso X's Lapidario by structural matching. The six period lists above put the comparison on type-token, initial-letter and attestation statistics instead, with the same code as the language samples.