The data and how far to trust it
What are the sources, and how far can the findings be trusted? The sources are Yale University Library's photographs, three transcriptions, and samples of real writing. A transcription is the text typed out from the photographs, letter by letter. The three were made by others who have studied the manuscript. The samples of real writing number 243. They are passages cut from forty texts, and lists of the early fifteenth century. The texts are in thirty-one languages. One transcription is by René Zandbergen and Gabriel Landini together. Another is by Takahashi. Those two agree with each other on 84.1 per cent of the words. The third, Claston's, tells apart letter shapes the others merge. The findings hold whichever is used.
The primary sources are the Beinecke Library's scans of MS 408 and three complete transcriptions of the text, sign by sign, each made independently from those scans. The Zandbergen–Landini file, which has 5,385 text units, that is, lines and labels, is the most complete of the three and is the default here. The other two are Takahashi's file and Claston's "v101" file, and that last file uses a finer alphabet, which splits several shapes that the other two merge. One purpose-written reader parsed all three files, and it keeps every line's page, position and layout information. Glyphs are named throughout by the EVA convention, the "extensible Voynich alphabet", which gives each common shape a Latin letter. So a "word" such as qokeedy is a string of glyph names and not a reading.
Map of the page sides
Each cell is one page side, in manuscript order, read left to right along the upper strip and then the lower. The three rows give the section of the book, the Currier language and the scribal hand recorded in the page headers of the Zandbergen-Landini transcription. Vertical lines mark the boundaries between quires, the gatherings of folded sheets from which the book is sewn.
The sections of the book
A transcription is an interpretation of the page, so before measuring anything else this examination aligned the two independent readings that share an alphabet, Zandbergen–Landini and Takahashi, unit by unit. They agree on 84.1% of words overall and on 85.7% of words in running text, with 2.15 character-level differences per 100 characters. The disagreements are concentrated in a few shape pairs:
| Shape pair (EVA letters) | Share of all disagreements | Reliability for statistics |
|---|---|---|
| a vs o | 18% | unreliable |
| r vs s | 9% | unreliable |
| ch vs ee, and the count of e strokes | 9% | unreliable |
| number of i strokes (in / iin / iiin) | 5% | unreliable |
| k vs t | 4% | mostly safe |
| q, d, l, n, y; word spaces marked as certain | < 1% each | safe (certain spaces confirmed by the second reader 99.4% of the time) |
To see which reading was right where they differ, this examination checked twenty-seven lines and labels, spread over every section and all five scribal hands, against the scans at full resolution. Those lines held thirty-one disagreements between the two transcriptions, more than one on some lines, of which fourteen went to the default file and four to Takahashi. Thirteen stayed unresolved even at full resolution, among them all six r/s cases. Word gaps are visible on the page, but their width varies. In two of the sampled lines a word-sized gap falls inside a string that both readers wrote as one word. Agreement is best in the biological and star sections (hands 2 and 3) and worst on the zodiac and astronomical pages (hand 4), where the text is written around diagrams.
Because the readings differ in these ways, this examination checked every result below on both readings, on both conventions for uncertain spaces, and on the v101 alphabet. No verdict depends on an unreliable distinction, and where a number moves between readings, the report gives both values.
The comparison set is forty texts in thirty-one languages, with eleven list-like texts of the period beside them. Sixteen of the forty, in eleven languages and three scripts, enter every comparison. They are Latin (Caesar, Virgil, a 1570 natural history), Italian (Dante), German (Goethe, and a modern translation of the Nibelungenlied) and English (the King James Bible, Shakespeare). The others are French (Montaigne), Spanish (Cervantes), Hebrew (two rabbinic texts, a collection of Geonic responsa and Delacrut's sixteenth-century Tzel ha-Olam), Greek, Finnish (Kalevala), Hungarian and Czech. Nineteen more texts, in fifteen further languages, cover the abjads and the languages outside Europe. They are listed in the language survey, and this examination repeated every headline comparison below on them. Five more, in five further languages at the far ends of the range of language types (Hawaiian, Tongan, Maori, Inupiatun and Huasteca Nahuatl), complete the forty. Because not every text enters every comparison, the number of languages quoted beside a statistic, from 14 to 31, is the number for which that statistic was computed at matched size. The report compares faithfully transcribed medieval manuscripts and early prints separately.
Every comparison uses the same sample size as the Voynich running text, because vocabulary and entropy statistics (entropy being the bits of surprise per glyph) depend strongly on sample size. The sixteen core texts supply twenty random windows each, and since the windows are drawn independently, the windows of a short text overlap. The repetition and word-shape statistics of the wider language set use five windows per text at fixed offsets, which gives the 140 language samples cited in this report. The list texts use a sweep of ten windows. Seven of the sixteen core texts are shorter than the running text, so they enter the glyph-entropy comparison at a 10,000-word tier. A plug-in entropy, one computed straight from the counts, runs lower there, which makes the comparison conservative, because the bias moves those seven languages towards the text. At full size the nine longer texts span 3.03 to 3.36 bits, and at the smaller tier the seven shorter ones span 3.15 to 3.61.
Word spaces, measured on the page. The transcribers mark some word spaces as certain and others as uncertain, and the scans show the difference. This examination measured the gaps between words on the scans, for the pages whose glyphs were cut out of the images (described below). Spaces marked certain have a median width of 1.10 times the height of a small glyph. Spaces marked uncertain have a median width of 0.55, and the gaps between glyphs inside a word 0.04. Width alone predicts the transcriber's certainty flag with an area under the curve of 0.68. That score says how well one measure sorts two groups, where a half is chance and one is perfect. A permutation null, the same score with the flags shuffled, gives 0.53. Rozanova and Temerev reported a sharper separation, 0.905, from their own coordinates, and the automatic segmentation here is noisier than theirs, which explains the weaker figure. The uncertain spaces are also intermediate in the text's own statistics. Inside a word, each glyph tells something about the next. Across an uncertain space the glyphs keep a third to a half of that association (0.29 to 0.43 depending on the alphabet, where Rozanova and Temerev's preprint gives 0.49). Across a certain space they keep a tenth (0.09 to 0.14), and across a line break none (0.002). A segmenter, a program that guesses where the word boundaries fall, received the glyph stream with every space removed and nothing else. It places boundaries where the transcribers put certain spaces at the rate found for Latin, Italian and Hebrew. Its boundary F1 is 0.585, against 0.49 to 0.62 for the languages and 0.18 expected at chance. That score combines how many of the placed boundaries are right with how many of the real boundaries are found. At the uncertain spaces it places boundaries at 0.8 of that rate, far above its false-positive rate elsewhere, which is about 9%. This examination therefore treats an uncertain space as a weaker word boundary and not as noise. Splitting at uncertain spaces is the default convention, and every headline number is also given with them joined.
How much the sampling matters. Every headline value survives the removal of any one quire. The pages are bound in gatherings, called quires, whose contents differ, so this examination recomputed each headline value with one quire left out at a time. Glyph predictability stays between 2.10 and 2.15 bits, against 2.52 to 3.69 in the languages, the repeated four-word count at 0 or 1, and the repeated 30-glyph count at 0. The adjacent-word information, which is how much one word tells about the next, stays between 0.09 and 0.12 bits, where the languages have 0.11 to 1.23. The share of word types one edit from a commoner type stays between 0.74 and 0.76, where Latin has 0.27. Four statistics vary more between quires than page-level sampling noise allows, because the two writing regimes (Currier's A and B styles) and the sections differ. They are glyph predictability, the Zipf slope, the slope of word frequency against rank, page burstiness, how much a word's occurrences bunch on particular pages, and word length. For them the variation between quires is two to three and a half times the page-level noise, so an interval quoted from page halves understates their uncertainty by that factor. Where both are available, the quire-level figure is therefore the one to read. For the adjacent-word information that is a jackknife standard error of 0.03, from leaving out one quire at a time, where shuffles within lines give 0.016. The published form of this check is a bootstrap that resamples quires with replacement, drawing quires at random, with some drawn twice and others not at all. That is valid for per-token statistics but not for repetition counts, which duplicated pages inflate. The section-vocabulary effect reported below comes mostly from the biological gathering, which is one quire, and without it the statistic falls from 2.91 to 2.16. Without both it and the recipe gathering it falls to 1.54, which is still 5.6 standard deviations above the null that holds hand and regime fixed, recomputed on the same reduced set of pages (1.27). For comparison, the null for the full text is 2.05. So the effect is not an artefact of one section.
Method, sources and reproduction
The primary data are the page scans and three transcriptions. The scans come from the IIIF image service of the Beinecke Rare Book & Manuscript Library, which holds the manuscript as MS 408. This examination fetched all 226 page sides at 1600 px, with full-resolution regions for the 27 lines and labels checked against the transcriptions. Those are a different set from the 25 darker lines that the ink census leaves as retouching candidates. The transcriptions are ZL 3b (Zandbergen–Landini, May 2025), IT 2a (Takahashi) and GC 2a (Claston's v101 alphabet), all in the IVTFF 2.0 file format from voynich.nu. While the analyses ran, this examination consulted no prior decipherment claims, theories or secondary literature about the manuscript, and the analyses use only generic methods: information theory, corpus statistics, and cryptanalysis with controls (a control being a text of known origin run through the same test). The literature was read afterwards, for the comparison with published work above and to check each verdict against the alternatives that others have actually proposed.
The work ran as thirty-one independent analyses on the same parsed corpus, and each produced a report with its code. The first four dealt with transcription reliability, the information-theoretic profile (the predictability of the signs), word-internal structure, and layout and position. The next six dealt with controlled decipherment, generative models, transposition, codes and numerals, hidden channels, and content. Then came representation learning (computer models that learn to predict the text), the language survey, the drifting-state generator (a generator being a program that writes text by a set of rules), production order, glyph shapes and the information budget (the bits per word that a predictive model leaves unpredicted). Six more dealt with the Naibbe cipher, the medieval corpora, spelling and transcription sensitivity, unicity distances, artefacts of known origin, and the labels as known plaintext. Four covered the held-out generator competition (generators scored on pages they were not fitted to) and the page classifier (the program that learns to tell real pages from generated ones), the matched-inventory and typological comparisons, the list genres and composites, and isomorphs (glyph strings with the same pattern of repeated signs) and alphabet tests. The last three were the codicological measurements on the scans (measurements of the book as a physical object), the generator search with the model-class contest and the recurrent-model anatomy, and power, calibration and multiplicity. Two more followed: the verdict statistics inside each Currier regime, and the hand procedure of the period. Then came the search for a smaller procedure on four independent lines with a shared scoring rule, with the held-out and page-classifier checks of its leading rows and the label test. The derivation of every word from what the scribe had in sight followed. After all of them, a separate re-derivation with separately written code checked every headline number. Where two definitions gave different figures, for instance the strict and loose counts of one-edit neighbours, the report gives both. The archive alongside this page holds all scripts, intermediate tables and the analysis reports.
Power, multiplicity and a confirmatory run. Every null result has a minimum detectable effect, the smallest planted signal its test catches, and to find it this examination planted graded signals in the text and re-ran the test. At 80 percent power, meaning a four-in-five chance of catching the signal, the tests detect cross-line word dependence above 0.2 of the in-line level. They catch word dependence at distance two above 0.2 of the adjacent level, and at distance three only above 0.4. A nested glyph structure (classes within classes inside the word) shows up on 2 percent of the tokens, and a forward variant genealogy (later words derived from earlier ones, in page order) on 4 percent of the one-edit pairs. Labels attested on their own page show up above 5 percent of the label tokens, and consecutive labels that count (labels that run as a number sequence) above 5 to 10 percent of the label pairs. A paired-gallows key (a key carried by pairs of the tall gallows glyphs) shows up on 10 percent of the paragraph first lines. Near-repeat clustering shows up on 2 percent of the tokens, and a burst of function words (the small grammatical words) on 10 percent of their tokens. The tests also catch a single wrong entry in a transferred cipher key. Two statistics are weaker: the intermittency statistic registers a planted vocabulary shift only on 14 to 20 percent of the pages, and the fluctuation exponent registers none at all. Replacing a fifth of the lines by Latin lowers the recurrent model's gain from 0.13 to 0.08 or 0.10 bits. The analyses saved 445 p-values (a p-value is the chance of a result at least as large when nothing is there), of which 203 are below 0.05. A false-discovery correction within each family, which limits the share of false findings among those kept, leaves 185 of them (Benjamini and Hochberg 1995). A family-wise correction across all tests, which limits the chance of even one false finding, leaves 137 (Holm 1979). The effects that do not survive are the bifolium vocabulary test (a bifolium is a sheet folded once to give two leaves), three tests on label groups, one pair of sections and 16 positional and shape-channel tests. No verdict rests on any of them. The bifolium vocabulary test is the one in the section on what produced the text. There, the two leaves of one sheet share slightly more vocabulary than leaves the same distance apart on other sheets. Latin prose poured into the same page skeleton shares as much. The sheet finding in the section on what the text is like is a different test. It rests on held-out surprisal, the letter-distribution step at a change of sheet and the mixed-quire comparison, none of which the correction covers. A confirmatory run fixed the rules in advance and used Takahashi's transcription in place of the default. It finds 3 of the 23 battery statistics (the battery being the set of 23 statistics applied to every text) different between the two readings, all of them word-length statistics, and it rejects every generator jointly on both readings.
The minimum detectable effects follow the practice of the simulation literature, which plants effects into real data (Franklin, Schneeweiss, Polinski and Rassen 2014), at Cohen's (1988) conventional 80 percent power. On this text, Walsh's unreviewed 2026 preprint recovers planted page-level anchors above one effect size and not below it. Two other unreviewed 2026 repositories plant a signal of two or three sizes into one test of their own. No peer-reviewed Voynich study that the literature search found reports the detection rate of a planted signal, or a minimum detectable effect for a null result. The corrections are Holm's, Benjamini and Hochberg's, and Westfall and Young's, applied without change. Two unreviewed 2026 repositories apply a maximum-statistic permutation (a correction that compares each result with the largest that chance would give) to their own test families. One automated exploration (Aspect Research 2026) corrects its whole ledger of findings. None of the peer-reviewed Voynich studies that the search found corrects across its own tests. The confirmatory run joins two published practices. One is the finding that single statistics hold across transliterations, the different readings of the script (Bowern and Lindemann 2021; Zandbergen 2022). The other is fixing the decision rules before the run (Nosek, Ebersole, DeHaven and Mellor 2018). Ventura and Bosco's battery locks its panel of metrics across two transliterations, and states that it is not confirmatory without a dated deposit. No earlier work re-derives a verdict on a second transliteration under rules fixed in advance.
Smallest planted effect each null test detects
For each test that found nothing in the text, the smallest planted effect it detects with 80 percent power, that is, in four runs out of five. Signals of graded size were planted in the text and the test was run again with its own code. The unit of each effect is given under the test's name. The scale is logarithmic.
What was not done, and why
The set-off test needs full-resolution scans of every page side, and each pair of facing leaves registered (aligned image to image). At 1600 pixels the transfer of ink through a leaf is measurable, but the transfer between facing leaves is only borderline (p 0.12), and pricking and ruling are invisible at this resolution. Erasures, retouching and overwritten letters cannot be told from ordinary variation in the ink without a manual pass over full-resolution crops, so the census of corrections here is the transcribers' record. Classifying the allographs (the variant shapes of a letter) by hand is days of a palaeographer's work, and the pages of hand 4 have no full-resolution crops. A control set of digitised manuscripts, with line-level transcriptions and a known production mode, would have to come from several archives, under their licences. The copyist used to measure the line-end fit is therefore a rule and not a person. The Naibbe cipher and the generators were not re-fitted inside each Currier regime; only the meaningless imitation was, so the per-regime contrasts against the Naibbe cipher are those of the whole text. No hand procedure was carried out by hand; every candidate was simulated. The two obvious repairs of the first one, spelling rules that see two glyphs of context and a ruler that reaches three lines, were run in the search and closed no gap. The rows at the floor under the bound (a row being one candidate procedure of the search, and the floor the distance between the text's own two halves) were not made to write the labels. The step that makes the once-only words, stated from the page evidence, was built late into the recipe in the form that draws once, and into the tuned rows of the second line (one of the four independent lines of the search) in its full form. No row carries it and holds the once-used share on a half it was not tuned on. The simplest device (the smallest apparatus of tables and a die in the search) was not given the new-word step or a working sheet, and no row was carried out by hand with a clock.
The Codex Seraphinianus is in copyright and has never been transcribed, and the Rohonc Codex's transcription and code tables are in its decipherers' publications, in no machine-readable form. No transcribed corpus of glossolalia is public. Bulk texts in Linear B, Cherokee or Ethiopic were not found for download, and no Latin manuscript at the heaviest abbreviation density was available in diplomatic transcription, so simulations stand in for both. Name lists in Arabic and Hebrew, a period star catalogue and Samoan or Fijian scripture would have widened the label and language comparisons, and this examination obtained none of them. Three computations were not run, because each is days of computation or of construction. One is a refit of the drifting-state generator with its state stepping at the boundaries of the physical sheets instead of at page boundaries. Another is a multi-day adversarial search for a generator with a page-level state. The third is a word-level Naibbe cipher with page-persistent units. Examination of the ink and the parchment is impossible from scans. Each of these would sharpen a figure reported here, but none bears on the absence of repeated phrases, of a picture-to-word link, or of a recoverable channel, which the tests measure on the whole text.