Verdict
There are five verdicts, from the firmest to the most open. They are ruled out, very unlikely, unlikely, undecided, and the best guess, not confirmed. The list below says what each verdict means. Ruled out: a hidden message of 3,000 words or more, and random meaningless text. Very unlikely: a real language in a made-up alphabet. Very unlikely too: a simple cipher, if the text behind it repeated its phrases. A changing cipher is unlikely, since one kind of it was not tested. A list or catalogue is undecided. The best guess, not confirmed: a procedure carried out by hand with dice and tables of words.
This examination found no reading of the Voynich text. Its measurements make two explanations very unlikely: a natural language written in an invented alphabet, and a cipher of such a language under any fixed mapping. The measurements contradict every tested explanation of those kinds, and the family as a whole survives only on one condition. The original text would have to be unlike every one of the 140 language samples, 92 samples of transcribed medieval texts and eleven lists, glossaries and rolls of the period compared here. The text behaves instead like one produced by a procedure, and the tests found no content of the kinds they detect. The most economical account of that procedure, the one with the fewest parts, is consistent with every measurement in this report. In it, each word comes from a short template, a stock of word families supplies the words, and the part of that stock in use shifts slowly as the pages go by. The glyphs are bent to the position of the word on the line and the page, and each section makes a fresh start. That account is the best guess, not confirmed: it organises the observations, and it is not a recovered record of how the book was made. Generators built on it reproduce the text's properties one layer at a time, but none reproduces them all, and other procedures might leave the same traces. Whether the text has meaning is a question that these measurements do not decide.
Each explanation gets one of five verdicts. An explanation fails when text written by its rule comes out unlike the manuscript on the measurements. One fact runs through the verdicts. Every sample of real writing repeats some of its phrases, and the manuscript almost never does. A cipher disguises a message, and nobody can see the message behind a cipher. So nobody can see whether that message repeated its phrases. The verdicts assume it did, since every real text sampled does. From the most certain verdict to the least, they are:
- Ruled out: the tests exclude it, and no assumption is needed.
- Very unlikely: no way of doing it can pass, tested or untested, if the disguised text repeated its phrases as every real text does. This verdict rests on that assumption.
- Unlikely: every way we tested fails, but a way we did not test could pass. The assumption does not settle it.
- Undecided: nothing found supports it, and nothing rules it out.
- The best guess, not confirmed: closer to the manuscript than any other explanation, but not passing every test.
The verdict rests on four findings, listed from the most certain to the least.
| Finding | Decisive measurement | What it rules out | Confidence |
|---|---|---|---|
| The text is not a natural language written in an invented alphabet | Glyph predictability (how much the previous sign tells about the next) is 2.11 bits, against 2.52 to 3.69 in 40 texts of 31 languages. The text is 0.41 bits below Maori, whose alphabet is as small as the text's and which repeats thousands of four-word sequences. Once the other languages are merged to the same sign inventory, the text is about 0.3 bits below the nearest of them, and the order information inside the word, 1.14 bits against at most 0.95, accounts for the rest. A syllabic rendering of Sanskrit is 1.9 bits above the text. The text has one repeated four-word sequence, or at most 45 with every disputed glyph distinction merged, against 67 or more in every language sample, or 38 or more merged the same way. No faithfully transcribed medieval text comes near, and no simulated scribe varying at a plausible rate does. The word-length variance is 3.19, against 4.0 or more in the alphabetic samples. The abjads, scripts that write few or no vowels, have shorter written words and a variance of 1.6 to 3.7, and the other two properties separate them from the text. The narrowness of the word-length spread holds at every sign inventory, and its symmetry only at the letter level | Every language sampled, in the prose, verse and scripture genres sampled. Any other alphabetic or abjad language, or any other genre, on one assumption: that a text in it would fall inside the range set by the 232 samples, the five languages at the extremes of the world's language types and the eleven list-like texts, on the properties that relettering cannot change | high for the sampled languages; very unlikely, under that assumption, for the rest |
| No cipher, code or transposition (a rearrangement of the letters) of a natural-language text under a fixed mapping fits the text | With the spaces removed, the text has no repeated 30-glyph string. 131 of the 140 samples have 28 to 19,637; four samples of Dante's verse have none, and five prose samples have 4 to 26. Any fixed encoding preserves such repeats. The solvers, code-breaking programs, recover known keys from monoalphabetic ciphers (one sign per letter), homophonic ciphers (several signs to choose from for each letter), short verbose ciphers (a group of signs for each letter) and anagrammed ciphers (letters rearranged), in Hebrew, Arabic and Sanskrit as well as the European languages. On this text they reach only the level of the meaningless control, a text with no meaning run through the same programs. The text is 530 to 6,200 times longer than the unicity distance of each of those families, the length of ciphertext beyond which the language's statistics fix the key. So the solvers rule out keys of those kinds, and they tell nothing about keys outside their reach, as the Naibbe cipher shows (Greshko's published card cipher, which encodes Latin by drawing cards). Keys fitted on Currier A (one of the two writing styles Prescott Currier identified) lose 0.15 to 0.72 bits on B, where a real cipher loses none. Codebooks, number systems and syllabic codes fail on separate grounds | Simple substitution. Homophonic substitution with the homophone chosen at random. Fixed verbose codes. Transposition and anagram ciphers with a fixed substitution. Plain word codebooks. Positional numerals written one unit per digit, Roman-style numerals and ordered tables. Syllabic and letter-pair codes. All of these are ruled out for a plaintext that repeats phrases as every sampled text does. Neither the repetition argument nor the solvers reach encodings whose spelling varies by chance or with state. The published Naibbe card cipher reproduces the entropy (the bits of surprise per glyph), repetition and word-order statistics; the lexicon (the stock of distinct words) and the page statistics rule it out instead. The 21 schemes built here in which the substitution changes as the text goes on (polyalphabetic, autokey and position-keyed schemes) either keep the repeats or raise the entropy to 3 bits or more. They are ruled out as built, and the family is unlikely but not ruled out. One scheme is untested and so not ruled out: a word-level homophonic nomenclator (a codebook with several code words for each plaintext word) whose spelling persists across a page and depends on the neighbouring word, with separate tables for Currier A and B. It would hold about 9,000 words of Latin. One more is not ruled out by vocabulary shape or word order, though it fails the near-copy structure by five to ten times: a nomenclator with about two homophones per word and one null (a filler sign with no meaning) in eight | high for every form tested. For the fixed-mapping family as a whole, that confidence holds under the assumption that the original text repeats phrases as every sampled text does. High that the published Naibbe cipher is ruled out. The modified card cipher that would pass has not been built, so that family is unlikely and not ruled out |
| The text is not memoryless. The choice of words depends on something that persists across a page and a leaf and differs between sections. No generator (a program that writes text by a set of rules) without memory reproduces that | Word types come back above chance across the page (1.7×) and on the other side of the leaf (1.5×), with scribe and regime (the Currier A or B style) held fixed. That is the profile of Latin and German. Of four real texts poured into the same pages, the Latin and the German match the manuscript at the leaf, and the Italian sits near chance there. Every local generator has no memory beyond three lines. Each section has its own vocabulary (F = 2.91, a spread-between-groups statistic, against 2.05 for its chance level). Word families one edit apart cluster in context as strongly as in a language | random word salad, tables, grilles and copying generators with no memory beyond a few lines | high that it is not memoryless. A meaningless state that drifts slowly, fitted to the recurrence profile, reproduces the profile, but a text about something would leave the same profile, so the drifting state is one reading of the finding and not part of it |
| The tests found no content of the kinds they detect | Words do not follow the pictures they stand beside: the correlation r is −0.02, where a planted link gives +0.26. The commonest words are not a grammatical class. Zodiac labels come back across pages at the chance rate. The labels, attacked as known plaintext (with plant and star names tried as the words behind them), yield no key, where enciphered Latin names yield theirs. For n-gram models, which predict the next glyph from the few before it, the glyph stream is a fourth-order process: nothing beyond four glyphs of context helps, and neither does more data. A recurrent model, which keeps a running memory, gains 0.13 bits per glyph, about a tenth of a bit, from the current word and the few words before it, more than on any control, and nothing at longer range. Adjacent words share 0.11 bits of information, where prose in the sampled languages shares 0.26 to 1.23. Only Latin and Sanskrit epic verse come as low, and a codebook keeps its source's value. 29 derived channels (sequences drawn from the text: first glyphs, gallows, which are the tall glyphs, word lengths, stroke counts, single slots) show no structure beyond their carriers, where a planted Latin text is found from a few hundred symbols. The word-order information lies at the word boundary: the last glyph tells 0.19 bits about the next word's first glyph, and the next word's core adds 0.001 once its prefix is known. That is the mark of a glyph rule crossing the space, and not of words chosen for their meaning | A readable message in the words or their order, of the kinds and sizes the tests were shown to detect. A language-like message of 3,000 words or more in the 29 channels tested. Not ruled out: content in any other form, or below those sizes | moderate, for the forms tested; no bound on other forms. The best model leaves 11 bits per word unpredicted, which is room for a message but not evidence of one, since it measures what the model fails to predict and nothing about what the text means. A payload compressed, or keyed to random choices, is beyond any statistical test. The Naibbe cipher shows the limit in a concrete case: it is a genuine encoding of Latin, and several of these tests score it as meaningless |
How far the verdicts reach. The verdict rests on two kinds of argument, which have different weights, and it keeps them apart. The case against a natural language in an invented alphabet, and against any fixed mapping of a real text, rests on properties that an encoding cannot change. Relettering leaves glyph predictability, word-length shape and phrase repetition as they were, and a fixed mapping with fixed unit boundaries keeps a text's repeated passages. The measurements show that much. They cannot show the second premise, that the original text has those properties in the first place, which means that it repeats phrases and spells as predictably as the 140 language samples and 92 medieval transcriptions compared here. No sampled text contradicts that premise, and none of the 294 languages measured in the published literature does. It is still an assumption about language, genre and spelling, and those two claims reach their whole families only under it. Every other verdict rests on solvers and fitted models, so it reaches only the forms this examination tried. The models tested do not reproduce the text's full statistical profile, but that alone does not decide whether the text contains meaning, and it does not rule out every encoding of natural language. For encodings whose spelling varies by chance or with state the caution applies in full, because neither the repetition argument nor the solvers reach them. What rules out the published Naibbe cipher is the lexicon and the page structure, and a differently built scheme might match those. Because the text falls into two writing styles, the statistics behind these verdicts were also measured inside Currier A alone and Currier B alone, against samples of the same lengths, and none changes sides. The repetition argument is the one that needs the whole text, because the A pages alone are too short for it.
The procedure, as the best guess. The most economical description of the text is that people wrote it word by word, by a procedure with these parts. A template shapes the inside of each word, and a stock of word families supplies the words. That stock shows no trace of changing along the codex (the book, from its first page to its last), even though the test that looks for such change does catch an artificial text whose vocabulary evolves. At any moment only a shifting subset of the stock is in use, and the subset changes slowly within a page and faster at page and section boundaries. A memory keeps some two-word sequences, and rules govern the first and last word of a line, the first of those rules applying wherever the pen restarted after a drawing. The writers fitted each line to the margin, and on the herbal pages to the drawings, by choosing shorter final words and not by tightening the spacing. All of that is a description of regularities and the best guess about how they arose, and it is not a demonstration that anyone assembled words from a template, since other procedures could leave the same traces.
This examination tested one part of that procedure by prediction: the drifting state. It fitted a generator with nine rates to the recurrence profile (how often words come back) at six distances, and with no further fitting the generator then predicted three things. The first is that a word has no association with any word beyond its neighbours, and the second that frequent and mid-frequency words are equally bursty, equally prone to come in clumps. The third is that one-edit variants share contexts, at about 70% of the measured tendency, and the generator did all of that with no meaning in it. It missed the word-order information, the section vocabularies, the open vocabulary (new words keep appearing to the end) and the line-edge glyph effects. A separate one-parameter memory for two-word sequences reproduces the word-order information and predicts the measured absence of order at distances of two and three. Whatever it does at the line edges, it inherits from the template beneath it. No single generator has yet produced all of these together. The nearest is the hand procedure of the period, which comes within three tolerances on 21 of the 23 statistics and within tolerance on 9 (a tolerance being the range a statistic wanders over between random halves of the text's pages), and misses the spread of word lengths and the start-of-word information. The search for a smaller apparatus found a procedure of 2,870 cells (table entries) that reaches all 23 on the whole text, as one of 21,305 cells does. Neither holds on quires its tables were not built on, and both are told from the real pages. The derivation of every word from the page finds the rare words spelled fresh with their length settled first. The one candidate built with that step carries it at the price of other statistics. So the procedure remains the best guess, not confirmed, and that is why the confidence in this verdict stops short of high. The generators' shortfall can be measured on quires (the gatherings of folded sheets that make up the book) that they were not fitted to. Scored on random halves of those quires, the best of them misses the text by twice the floor, where the floor is the distance between the text's own two halves. The same three debts appear in every split (every division of the quires into a fitting half and a scoring half): the open growth of the lexicon, the tightness of the word-length distribution and the near-copy rate between adjacent words. A global search of the generator's parameters against the whole battery of statistics (the set of 23 statistics applied to every text), scored on held-out quires (quires the generator was not fitted to), does not close that gap. It reaches 14 to 15 of 23 statistics at best, against 17 for the text's own other half, and it misses the same statistics at every setting. A recurrent model trained on the generator's output loses 0.16 bits per glyph on the text, so the debt is not a matter of tuning.
What this examination cannot say. It cannot say why the text was made, because the measurements describe how the text behaves and not the intention behind it. An elaborate hoax, an exercise, a private notation with almost no content, or a text produced for its own sake would all leave the same traces. It cannot say whether the text has meaning, because the tests detect content of particular kinds, and finding none of those is not the same as finding that there is none. Nor can it name the language of a plaintext, when the tests give no reason to suppose that a plaintext exists. Three of the null results also come from weak instruments, and their limits go with them. Word dependence at distance three is detectable only above 0.4 of the adjacent level. The two fluctuation statistics, which track how the use of words rises and falls along the text, register a page-scale vocabulary shift only when it covers 14 to 20 percent of the pages, or not at all.
What the data in hand cannot settle. Some tests a practitioner would want are out of reach of scans at 1600 pixels and of public transcriptions, and several things about the pages cannot be read from the scans. The first is the order in which the scribes wrote the sheets. Set-off, the wet ink that transfers between facing leaves, could give that order, but a pilot test on these scans is suggestive and not decisive. The ruling and pricking of the leaves (the guide lines, and the holes pricked to lay them out) cannot be seen at this resolution. Erasures and retouching remain open, with 25 candidate lines awaiting a manual pass at full resolution. A palaeographer, a specialist in old handwriting, has not classified every letterform, and scans allow no examination of the ink or the parchment.
Four kinds of material that would calibrate the measurements are also missing. The first is a set of digitised manuscripts of known production, copies of known exemplars (the source a copyist worked from) and authors' working drafts, with line-level transcriptions and images. Against those, the examination could place the leftover space at the line end, the correction rate and the layout on a scale, instead of against a simulated copyist. The second is the Codex Seraphinianus, a modern book of invented script, which would calibrate the signatures of a text composed at the desk with no exemplar, but which is under copyright and has no transcription. The third is the Rohonc Codex, an invented code that has been read, but whose transcription and code tables are not public. The fourth is transcribed glossolalia, speech in no language, of which no corpus is public. Two kinds of comparison text are missing as well. Syllabary texts (one sign per syllable) in Linear B, Cherokee or Ethiopic were not available in bulk, so the Sanskrit syllable rendering stands in for them. No real Latin manuscript at the heaviest abbreviation density was available in a diplomatic transcription (one that records every sign as written), so a simulation stands in for it. Two constructions were not made either: a generator with a page-level state strong enough to defeat the page classifier (the program that learns to tell real pages from generated ones), and a card cipher of the Naibbe kind with word-sized, page-persistent units. Each of these would sharpen a figure, but none of them would supply what every test that could be run found missing.
What would overturn it. Any of six findings would overturn this verdict. The first is a derived channel, or a residual sequence (what is left after a model has predicted all it can), with structure beyond its carrier and its controls. The nearest thing found is the prefix residual, which is 0.02 bits per symbol above the best generator. The substitution solver does not turn it into Latin letters, which are the one payload the solver recovers in its control test. The second is a key that transfers from Currier A to Currier B with little loss, and the third a picture-to-word association above r = 0.1, with better botanical descriptors. The fourth is a set of labels that track a shared referent, the thing they would name, from page to page. The fifth is a card cipher of the Naibbe kind, rebuilt with word-sized units and page-persistent, neighbour-dependent spelling, that reproduces the lexicon and page statistics. The sixth is a decoding that reads better than chance on held-out pages. In the other direction, a generator whose pages a classifier cannot tell from the manuscript's on held-out quires would strengthen the production account, but a classifier tells every generator built here from the manuscript at 0.88 or better. None of these six appeared in any of the techniques applied under controls here.
Set against the published positions, this conclusion agrees with Bowern and Lindemann (2021) that the linguistic evidence does not support a natural language in an invented alphabet. They and Zandbergen state the entropy-invariance argument, that a substitution cipher cannot change the predictability of the signs, and this examination turns that argument into a measurement on ciphers with known keys. Rozanova and Temerev's 2026 preprint also rules out simple substitution, from a differently calibrated attack. The conclusion also agrees with the generative accounts of Timm and Schinner (2020) and Rugg (2004) that a procedure can produce most of what the text shows, and it separates the two. Word-by-word copying, the process in Timm and Schinner's account, reaches three lines, where the manuscript clusters across the page and the leaf. Rugg's table and grille (a table of word parts read through a card with holes) reproduces the qualitative claims of its papers, but it reproduces neither the word grammar, nor the lexicon, nor the word-order signal. With Gaskell and Bowern (2022) the conclusion agrees that the text belongs with human gibberish. It adds that human gibberish matches the text's vocabulary statistics and none of its glyph statistics, so a meaningless text is not the same thing as a text made the way volunteers make one. None of this answers the question any of those authors asked. What it records is which published measurements survive one pipeline (the same code and reference texts for every test) with controls, and how much room is left.