The findings in brief
A sign is one letter of the manuscript's own alphabet. The manuscript's next sign is more predictable than the next letter in thirty-one languages. Its phrases almost never repeat. So a real language in a made-up alphabet is very unlikely. A simple cipher always writes the same letter or word the same way. It is very unlikely too. A procedure is a fixed set of steps carried out by hand with dice and tables of words. It is the best guess, not confirmed. The page ‘Verdict’ says what each of the five verdicts means.
This report does not translate the manuscript, and the measurements say why a translation has been so hard to find. A natural language written in an invented alphabet is very unlikely. So is a cipher of one under any fixed mapping, that is, under any rule that always turns the same letter or word into the same glyphs, the signs of the script. The reason is repetition: an encoding that always turns the same passage into the same glyphs keeps the repeated phrases of the original. Every one of the 140 language samples, the 92 samples of transcribed medieval texts and the eleven lists, glossaries and rolls of the period compared here has dozens or thousands of repeated phrases, where this text has one. Its words repeat as often as the words of a language do, but its sentences never repeat. The ruling out of those explanations holds for every form tested, and for the family as a whole it rests on one assumption, which is that the original text repeats its phrases as every sampled text does. An encoding whose spelling varies by chance or with position escapes that argument, and the one published example, Greshko's Naibbe card cipher, went through every test here. It reproduces the statistics of the glyph sequence and of repetition, but it fails on the lexicon, the stock of words, and on the page structure. The verdict says what a variant would have to do to pass. The code-breaking solvers, programs that search for a cipher key, score that genuine cipher of Latin as they score a meaningless control text (a control being a text of known origin run through the same test). So when they find nothing in the manuscript, that is no evidence against ciphers of that kind.
What the text does have is a rigid structure driven by position, which no sampled language has. Each word is built from a short template of ordered slots, and the choice of glyph depends on where the word is on the line and on the page. Neighbouring words are near-copies of each other far more often than chance allows. Two writing regimes, two distinct styles of writing, coincide almost exactly with two groups of scribal hands. Because of that, the statistics behind the verdicts were re-measured inside each regime alone, against samples of the same length, and none of them changes sides. The repeated-phrase argument is the one that needs the whole text. The wider set of published statistics splits by range. The long-range statistics of the text are inside the language range. They include the shape of the network of words that occur together, the information held by page-sized blocks of text, the similarity of neighbouring pages, and the way the swings in word length grow with the length of text measured. The short-range statistics, which are the predictability of the glyphs, the word template and the near-copying, are outside it. The tests scored four other kinds of text on the same statistics. Two of them are human gibberish, meaningless text that volunteers wrote by hand, and Enochian, the "angelic" language of John Dee and Edward Kelley. The other two are a genuine eighteenth-century cipher and non-linguistic symbol systems, sets of signs that are not writing. Each of the four matches the text on one of those two levels and fails on the other.
For each remaining family of explanation this examination built a generator, a program that writes text by that family's rule, fitted it to the text, and scored it against 23 measured properties. None of them reproduces the text. The closest is a hand procedure of the period, built from tables, lots, a ruler and a renewed working sheet, which comes within three tolerances on 21 of the 23 properties (a tolerance being the range a property wanders over between random halves of the text's pages) and misses two. A search over some forty families of smaller procedure, the kinds of smaller procedure only, finds a procedure of 2,870 table cells (a cell being one entry written out in a table) that matches all 23 on the whole text, as one of 21,305 cells does. Neither holds on pages its tables were not built on, where every candidate misses the spread of word lengths and the count of words used once. The derivation of every word from what the scribe had in sight finds no ruler and less copying than any imitation. It finds the rare words spelled fresh with their length settled first, a step that only the last two candidates carry, and both fail on other counts. What survives, then, is a list of properties and not a mechanism. The list is a rigid word template, an open vocabulary in which new words keep appearing, and similar words clustered by line and by page. It also includes rules for the first and last word of a line, and a trace of word order at the boundary between words, where a rule about glyphs crosses the space. The procedure described below is the best guess, not confirmed. It organises those measurements, and which process produced the text is not established. The sections below test each family of explanation one after another, and the closing verdict says how far each verdict reaches and how sure it is.
Against earlier work. The findings above rest on measurements that other people made first. In 1976 Bennett measured the entropy of the glyphs, the amount of surprise each glyph holds once the one before it is known, and found it unusually low. Lindemann and Bowern set that value below several hundred languages in 2021, in a preprint not yet peer reviewed. Tiltman in 1967, Currier in 1976 and Stolfi in 2000 described the rigid word template, while Zandbergen recorded the absence of repeated phrases and Timm counted it in 2014. What this examination adds comes to three things. Every test runs in one pipeline, a single chain of code, and has its own control. The repetition result is recast in a form that does not depend on where the word breaks fall, so that no fixed encoding can remove it. And each explanation is built as a generator and scored against the whole set of measurements together. The examination also finds that some published results do not survive a wider comparison set. Among them are the claim of Amancio and colleagues in 2013 about how the frequent words are spaced through the text, and the reading-direction asymmetry that Parisel reported in a 2025 preprint. The two-language reading of Arutyunov and colleagues, a 2016 preprint, is another. One published cipher, Greshko's Naibbe cipher of 2025, reproduces the entropy and repetition statistics, and this examination rules it out on the lexicon and the page structure instead. Each section below says who measured what before, whether it held here, and what is new, and the works referred to are listed at the end.
What is new in the method. This examination used 211 techniques. A literature check made after the analysis finds 32 of them with no earlier application to the Voynich text, 102 that adapt published methods with new controls or validation, and 77 that are established. A second count sets aside four unreviewed preprints from 2025 and 2026 and the forum and student projects, and against peer-reviewed and long-established work only the split is 55 new, 103 adapted and 53 established. The new techniques fall into a few groups. Several are tests of what a cipher could and could not hide. A repeated-string test strips out every space and counts the repeated runs of glyphs, which any fixed cipher would keep, and it finds that the text has no long one where nearly every real text has many. An order-invariant repetition test counts repeated runs of words by the letters they contain, so that an anagram cannot hide a repeat, and it finds that anagrammed Latin keeps its repeats while the text has almost none. A first-glyph test measures how much the first glyph of a word tells about the later glyphs. In a nested category code, where the first sign names a broad class and later signs narrow it, that would be a great deal, and in the text it is almost nothing. A homophone test looks for pairs of glyphs that stand for the same letter, recovers planted homophones, and finds none in the text. A grid of simulated codebook ciphers, built from five real texts, shows that no codebook matches the text on word order, vocabulary and near-copying together. Others look for content or a hidden message. A picture-to-word test asks whether the words on a herbal page follow what the picture shows, and it found a link planted as a control and nothing in the manuscript. A residual-channel test looks for a message in the part of each word that the best predictive model cannot predict, and it too found a planted message and nothing in the text. Others measure the text against new kinds of comparison. A measurement on the scans compares the space each line leaves at the right margin with a simulated copyist filling the same lines. A measurement of the information across line breaks shows that the end of a line tells almost nothing about the start of the next. Languages keep a third to two thirds of what they have inside a line. A genuine historical ciphertext and several non-linguistic symbol systems were scored beside the text as reference classes, known kinds of text to compare against. A word-pattern fingerprint records which glyphs in a word are the same glyph, a pattern that no substitution can change, and it places the text nearest to scripts written without vowels. A word-space recovery test gives a program the glyphs with every space removed and asks it to find the word breaks, which it finds better than it finds the breaks in Latin. A recurrent neural model, scored against models that predict each glyph from the few before it, gains a little, and the gain comes from inside the word and the few words before it. A calibrated score of closeness to known languages ranks no candidate language. Once the preprints are set aside, three more count as new: the verbose-cipher solver, the key-transfer test between the two writing regimes, Currier A and B (the two writing styles that Prescott Currier identified in 1976), and the controlled re-runs of the preprints' own measurements. The chart below gives every technique its status under both views, and a table lists what goes beyond the preprints.
The sections below give the measurements behind each sentence of this summary, and for each one they say what it rules out, what remains undecided, and what evidence would be needed to go further. Every number in this report was produced by code run on the primary data during this examination, and separate code derived each number a second time before it was used here.
Verdicts on the main explanations
This examination turned each family of explanation into measurable predictions and tested those predictions against the whole running text, so each verdict below reflects only what the measurements support. Some of the verdicts rest on a property that no encoding can change, and a verdict of that kind has two premises. The first is that the encoding keeps the property, which can be shown. The second is that the original text had the property in the first place, and that is an assumption about language, genre and spelling. The comparison texts support that second premise, but they cannot settle it, which is why each row says which premise its verdict rests on.