Beinecke MS 408 · independent examination
Voynich Text Examination

What the text of the Voynich manuscript is like, measured on the page scans and on full transcriptions, and compared with real writing in thirty-one languages.

Compared with published work

What is new here, and what was known before? A technique is one method of measuring or testing the text, such as one of the 23 measurements. We used 211 techniques in all. They fall in three groups. The first group is 77 techniques that had been applied to this text before. The next is 102 published methods used here for the first time. The last is 32 that are new. Every technique was also run on a text of known origin. About thirty of the findings had already been published by others. The page also says which published claims do not hold up on these tests.

This examination finished its measurements before it checked the literature, so no earlier theory could steer them. It then listed every technique it had used, 211 in all, and searched for an earlier application of each one to the Voynich text. The search covered journals, the 2022 Voynich conference proceedings, arXiv preprints up to August 2026, René Zandbergen's site, and the main blogs and forums. Each technique was then sorted under deliberately conservative rules, which lean towards calling a technique established. A technique counts as established if someone had applied the same measurement to this text before, whatever controls were used here. It counts as adapted if this examination changed a published method in a named way. A named change is one of three things: a new null model, a new invariant or a validation step. A null model shows what the measurement gives by chance. An invariant is a quantity that some kind of encoding cannot change. A validation step runs the method on a text with a known answer. A technique counts as new only if the search found no earlier application at all.

The count depends on what is allowed to count as prior work, so the report gives it two ways. The first view counts every source the search found, including four unreviewed preprints from 2025 and 2026, student projects, code repositories and forum threads. Under that view, 32 techniques are new, 102 adapted and 77 established. The second view counts only peer-reviewed journals and the long-established sources, which are Currier, Tiltman, D'Imperio, Bennett, Stolfi and Zandbergen's site. Under that view, 55 are new, 103 adapted and 53 established. The two views disagree on 36 techniques, and each of those has a preprint, a student project or a forum thread as its only precedent. "New" means that the search found no earlier application. It does not mean that nobody has tried the technique.

The techniques counted by area and by status, under each of the two views of prior work. Hover over a segment to see the names of the techniques in it.
NewAdaptedEstablished

Counting every source, preprints and unpublished work included

Counting peer-reviewed and long-established sources only

The full table of techniques, with the nearest prior work under each view.

In the table, h1, h2 and h3 are the entropy of a glyph, which is the number of bits of surprise in it, given none, one or two preceding glyphs. MI is mutual information, which is how many bits one quantity tells about another, and PMI is pointwise mutual information, the same measure for one particular pair of values. MDL is minimum description length, BPE is byte-pair encoding and kNN is k nearest neighbours. CI is confidence interval, LM is language model and KN is Kneser–Ney smoothing. PPMI-SVD is a word-embedding method, and LDA, LSA and NMF are topic models. IIIF is the image service of the Beinecke scans and IVTFF is the transcription file format. F/T/P means fitted, trained or predicted. ZL, IT and GC are the Zandbergen–Landini, Takahashi and Claston transcriptions.

#TechniqueStatus, all sourcesStatus, reviewed sources onlyNearest prior workWhat differs
1Conditional character entropies h1/h2/h3 of the glyph stream vs matched-size language samplesEstablishedsameBennett 1976, Scientific and Engineering Problem-Solving with the Computer (Prentice-Hall)The measure is Bennett's. Added here are exact size matching, with 20 windows per corpus, and three conventions for the space. Neither changes the known result of 2.0-2.2 bits against 3.0-3.6 for the languages.
2Sensitivity of entropy to transliteration and alphabet conventionEstablishedsameLindemann & Bowern 2021, arXiv:2010.14697, (partition by transcription system)No material difference
3Within-word letter-shuffle control attributing the h2 deficit to glyph order inside wordsAdaptedsameLindemann & Bowern 2021, arXiv:2010.14697…Lindemann and Bowern reached the same conclusion by studying where each glyph can stand inside a word. Here the letters are shuffled inside each word and the entropy is measured again. Rozanova and Temerev shuffle whole words within the line, which is a different control.
4Character mutual information at distance d=1..20 with a letter-shuffle floor (the value the shuffled letters give) and a word-shuffle decompositionAdaptedsameLandini 2001, Cryptologia 25(4), (spectral analysis of the text without spacesLandini, Schinner, Amancio and colleagues, and Montemurro and Zanette measured long-range correlation in other ways (spectral analysis, random walks, word-level information). The curve here is character-level mutual information at each distance, read against a letter-shuffle floor, and it is split into a word-template part and a word-order part.
5Sukhotin vowel/consonant separation with alternation-ratio diagnosticEstablishedsameGuy 1991, Cryptologia 15(3) 207-218 and 258-262 (two foliosGuy, and later Reddy and Knight, ran the same algorithm. The alternation ratio and the comparison with languages are small additions.
6Word-length distribution shapeEstablishedsameStolfi 1997-2000, Voynich pages (binomial word-length distribution)No material difference
7Lexical statistics at matched sizeEstablishedsameLandini 2001, Cryptologia 25(4), (Zipf)No material difference beyond exact size matching
8Adjacent-word mutual information in excess of a word-shuffled baselineEstablishedAdaptedRozanova & Temerev 2026, arXiv:2608.17096, (adjacent-token MI minus the mean of 100 within-line shufflesRozanova and Temerev, and Parisel, had the same idea and found the same qualitative result. Their shuffle moves words within each line. The shuffle here moves words across the whole window.
9Immediate repeats and edit-distance-1 near-repeats of adjacent words vs shuffled baseline and vs languagesEstablishedsameTiltman 1967 / Currier 1976 as summarized at voynich.nu 'Sentences etc', (doubled and tripled words)Tiltman, Currier and Timm counted the doubled words. Two things are added here: the ratio to a shuffled baseline, and the finding that prose languages fall below their own chance level. The technique is not new.
10Currier A against B (the two writing styles Prescott Currier identified), measured against half-splits of reference corpora and against halves of B aloneAdaptedsameCurrier 1976, reproduced and revisited at voynich.nuThe difference between A and B has been known since Currier. New here is the yardstick: disjoint halves of 16 reference corpora show how large a difference one text can produce on its own.
11Effective alphabet sizeEstablishedsameZandbergen, voynich.nu character statisticsNo material difference
12Entropy-invariance argumentEstablishedsameZandbergen, voynich.nu 'What we learn from entropy', (substitution leaves entropy unchangedNo material difference. The numbers here reproduce the known behaviour.
13Positional statistics of glyphsEstablishedsameTiltman 1967, NSA Technical Journal (reprinted), no stable URLNo material difference
14Adjacency and co-occurrence constraintsEstablishedsameTiltman 1967 (as above)The constraints are classic observations, from Tiltman and Stolfi onward. The shuffle nulls, which say how often each constraint would be broken by chance, are an addition.
15Unsupervised discovery of multi-glyph units by PMI merging and MDLEstablishedAdaptedRozanova & Temerev 2026, arXiv:2608.17096…Rozanova and Temerev learn the units by byte-pair encoding. Here they are learned by merging on pointwise mutual information and by minimum-description-length segmentation. The purpose is the same and the same units come out.
16Branching entropy inside wordsEstablishedsameZandbergen, voynich.nu 'From bigram entropy to word entropy'…Zandbergen measured the profile from the start of the word. The profile from the end of the word, read right to left, and the successor-variety counts are additions.
17Ordered-slot grammar learned by greedy set coverAdaptedsameStolfi 2000, (crust-mantle-core grammar)Stolfi and Zattera built their slot grammars by hand. The grammar here is learned automatically. It is scored by a curve of rule budget against coverage, by the productivity gap and by the shuffled-word selectivity test, and the same scoring is applied to Latin, Italian and Hebrew.
18Edit-distance-1 word familiesEstablishedsameTimm 2014, arXiv:1407.6639Timm and Schinner described the families. The language baselines at matched size (Latin, Italian, Hebrew) are an addition.
19Hapax legomena as concatenations of two attested shorter wordsEstablishedsameTimm & Schinner 2020, Cryptologia 44(1)…Timm and Schinner made the observation. Here the share is quantified and set against Latin and Italian at matched size.
20Sequential vs hierarchical within-word prediction testNewsamenone foundNo earlier test asks whether the first unit of a word constrains the later units beyond what the adjacent unit tells. That constraint is the signature of a nested category code. A left-to-right slot process does not have it.
21Frequency-band feature analysisEstablishedsameCurrier 1976No material difference
22Cross-transcription and convention sensitivity of word-structure measuresEstablishedsameLindemann & Bowern 2021, arXiv:2010.14697 (by transcription system, hand, Currier language)No material difference
23Line-initial and line-final letter and word distributions vs interior wordsEstablishedsameCurrier 1976 ('line as a functional entity')No material difference
24Reference texts poured into the Voynich page/paragraph/line skeleton as the null for layout effectsEstablishedAdaptedRozanova & Temerev 2026, arXiv:2608.17096…Rozanova and Temerev wrap prose controls to the line template. The verse and wrapped-line references, which have line units of their own, are a small addition.
25Paragraph-initial and paragraph-final effectsEstablishedsameCurrier 1976 and D'Imperio 1978 (split gallows on first lines), summarized atNo material difference
26Word length by position in the lineEstablishedsameTimm 2014, arXiv:1407.6639, (first word longer, second shorter, decline along the line)Timm measured word length by position. The curve of how the edge effect decays along the line is a small addition.
27Line-edge derivation testAdaptedsameSmith 2015, 'Linestart words', (initial o stripped at line startSmith proposed the stripping and adding transformation from glyph frequencies. Here it is tested by counting how often the stripped forms are attested words, against a matched interior baseline. The test of replacing final -m is added.
28Neighbour similarityAdaptedsameTimm 2014, arXiv:1407.6639, (similarity to words on the same line and the line above)Timm measured similarity to the words on the same line and on the line above. Added here are position-matched controls, which remove the confound of the line-initial word, and the flat profile over lags.
29Adjacent-repeat and near-repeat rates against a within-page permutation null, and runs of >=3 near-identical wordsEstablishedsameTimm 2014, arXiv:1407.6639Timm and Schinner counted the chains of similar words. The permutation null is an addition.
30Page-vocabulary Jaccard overlap by pair categoryAdaptedsameZandbergen, 'The Currier languages revisited', (page-by-page distance matrix from bigram/word distributionsZandbergen, Montemurro and Zanette, and Sterneck, Polish and Bowern clustered pages by their text statistics. Here the overlap is measured as an excess over a word shuffle, by pair category and by physical distance, and a narrative poured into the same skeleton is the reference.
31Word burstinessAdaptedsameMontemurro & Zanette 2013, PLoS ONE 8(6)…Montemurro and Zanette measured the same property with a different statistic. Here it is document frequency and recurrence against an exact expectation, by frequency band. The conclusion agrees.
32Currier A/B page classificationEstablishedAdaptedParisel 2026, arXiv:2604.25979, (supervised classifier, a program that learns to sort pages, predicting A/B of held-out folios, the folios kept out of its training, 89.2%)Parisel classified held-out folios from character statistics, so the technique exists. The feature set and the cross-validation by hand against section differ in detail.
33Label statisticsEstablishedsameZandbergen, voynich.nu (labels in the writing-system and analysis pages), andNo material difference
34Numbering-system test for labelsAdaptedNewvoynich.ninja thread 'Zodiac labels'…Forum posts on voynich.ninja have compared adjacent zodiac labels, but they are unreviewed. The test here compares consecutive labels with random pairs from the same page, as a check for a numbering system, on the star and the zodiac labels. It was not found in that form.
35Repeated word n-grams (runs of n words)EstablishedsameZandbergen, ('curiously lacks repeated phrases of 2 or more words')Zandbergen and Timm noted the missing phrases. The poured references at matched size and the shuffled controls are additions.
36ZL vs IT agreement at locus, word and character level by sequence alignment, with per-symbol dispute rates, confusion pairs, and breakdown by section/hand/languageAdaptedsameZandbergen 2022, CEUR Vol-3313 keynote 'Transliteration of the Voynich MS text'…Landini and Stolfi aligned the files of several transcribers long ago. New here are the dispute rate for every symbol and the ranking of confusion pairs between the ZL and IT transcriptions, by section and by hand. The agreement percentages in the earlier sources were not verified.
37The v101 alphabet as a third symbol viewAdaptedsameZandbergen, voynich.nu transliteration pages and IVTFF tools…Zandbergen's conversion tables were written by hand. The table here is estimated from the paired words, and it measures purity, which is how deterministic the correspondence is in practice.
38Word-space reliability between readersEstablishedsameZandbergen, IVTFF format (uncertain space convention)Rozanova and Temerev's 2026 preprint validates the uncertain spaces from glyph coordinates. That is more thorough than the scan sample used here.
39Targeted inspection of disputed glyphs on full-resolution IIIF cropsEstablishedsamestandard practice of every transcriber (Takahashi, Zandbergen, Claston)No material difference
4024 generative models from six families scored against a fixed 23-statistic battery, with fitted/trained/predicted bookkeeping per statistic, page-half-sample CIs and z-scoresAdaptedsameTimm & Schinner 2020, Cryptologia 44(1), (generator, a program that writes text by a rule, vs a set of statistics)Timm and Schinner, Gaskell and Bowern, Parisel, and Rozanova and Temerev all compared a generator with a set of statistics. Here 24 models are scored on one fixed battery, and a ledger records for each statistic whether the model was fitted to it, trained on it or predicts it. Only genuine predictions count.
41Autocopy generatorsEstablishedsameTimm & Schinner 2020, Cryptologia 44(1)Timm and Schinner fix the copying kernels. Here they are fitted to the text. The family of generators is theirs.
42Slot-template generatorsEstablishedsameRugg 2004, Cryptologia 28(1) (table-and-grille generation)No material difference in kind
43Word-level Markov chainsEstablishedAdaptedParisel 2026, arXiv:2604.19762…No material difference
44Hierarchical category-code generatorNewsamenone foundNo earlier test was found that builds a Wilkins-style nested category code as a text generator and scores it against the text's statistics. Bowern and Lindemann discuss Friedman's philosophical-language idea, but no generator had been built from it.
45Codebook/nomenclator generatorsNewsamenone foundNo earlier work scored a rank-matched codebook text against the Voynich statistics. Kuykendall's annealer has a nomenclator mode, but as a solver, and Greshko's Naibbe cipher is a substitution cipher. Neither is a codebook generator.
46Bias-corrected adjacent-word MI with an interior-only shuffleAdaptedsameRozanova & Temerev 2026, arXiv:2608.17096…Rozanova and Temerev's null shuffles every word in the line. The null here keeps the first and last word of each line in place, because edge words are different kinds of words (once-only words are twice as common among them). With a whole-line shuffle the correction flips sign (-0.12 bits) on this data. The edge-fixed version is what separates the real text from every copy and template generator. Whether the whole-line version does the same is not reported in the preprint.
47Simulated-annealing monoalphabetic solverEstablishedsameHauer & Kondrak 2016, TACL 4, (LM-based decipherment of the VMS as substitution/anagram over 380 languages)Hauer and Kondrak, Kuykendall, and Rozanova and Temerev use the same approach. The coverage here (21 language models, 3 symbol views, both boundary modes) and the cost charged for surplus symbols are details.
48Many-to-one (homophonic) solverAdaptedsameRozanova & Temerev 2026, arXiv:2608.17096…Rozanova and Temerev, and Kuykendall, fix the number of classes in advance. Here the annealer chooses how many symbols share a letter, under a description-length penalty, and the collapse of the unpenalised objective to 'iiii...' is documented.
49Verbose/null-cipher solverAdaptedNewRozanova & Temerev 2026, arXiv:2608.17096, (BPE-learned multi-glyph units mapped to lettersRozanova and Temerev fix the segmentation first, by byte-pair encoding, and substitute afterwards. The solver here searches the segmentation, the nulls and the letters together, with a stated cost for each null, and it reports word-level recovery on verbose ciphers with a known key.
50Meaningless controls for every solver runEstablishedAdaptedRozanova & Temerev 2026, arXiv:2608.17096…Rozanova and Temerev use the same idea with an order-3 imitation. The imitation here is order 2, and the within-word shuffle is added as a second floor.
51Held-out key-transfer testAdaptedNewParisel 2026, arXiv:2604.25979…Parisel tested the same idea, that one key cannot serve both A and B, from the switching of character ratios. Fitting a real decipherment key on A, scoring it on B, and comparing the loss with a genuine Latin cipher (0.000) and a Markov control was not found.
52Dictionary hit rate of decoded tokens vs the decoded Markov controlEstablishedsameHauer & Kondrak 2016, TACL 4, (dictionary-based word accuracy of decodings)Hauer and Kondrak measured the dictionary hits of their decodings. The comparison with a decoded control is an addition.
53Segmentation-free repeated-substring testNewsamenone foundZandbergen and Timm noted that the text repeats no long phrases. Counting repeated long strings with the spaces removed, so that the count cannot be changed by the cipher type or by errors in word division, was not found in earlier work.
54Cipher transformations of LatinAdaptedsameBowern & Lindemann 2021, Annual Review of Linguistics 7 (bigraph/verbose encodings of Latin lower h2Bowern and Lindemann, and Rozanova and Temerev, enciphered known text and measured h2. The joint criterion, low h2 together with no long repeats, and the position-keyed variant are additions here.
55Variant-collapse recount of repeated word 4-grams after merging unreliable or variable distinctionsAdaptedsameTimm 2014, arXiv:1407.6639, (the few repeated sequences differ by spelling variants and word order)Timm and Zandbergen raised the idea that spelling variation hides repeated phrases. The test here merges the variable distinctions, recounts the repeats and compares them with a shuffled control. It was not found in earlier work.
56Label-free unigram profile matchingEstablishedsameZandbergen, voynich.nu character statisticsComparing frequency profiles is standard, from Guy to Lindemann and Bowern. Reading the profile as a statistic that no transposition can change is a new framing. The measurement is not new.
57Within-word bag-of-letters statisticsAdaptedsameHauer & Kondrak 2016, TACL 4…Hauer and Kondrak proposed that the words are alphabetically sorted anagrams. The test here counts precedence violations, with a lower bound that no ordering of the alphabet can beat, and measures types per bag and multi-order bags. The controls are Latin texts sorted or anagrammed with a known key.
58Anagram-invariant substitution solverEstablishedsameHauer & Kondrak 2016, TACL 4…Hauer and Kondrak solved the same problem with the same kind of objective. Added here are the Markov and shuffled controls, positive controls in five languages scored against held-out lexicons, and both mapping modes.
59Route / columnar transposition testNewsamenone foundD'Imperio listed transposition among the possibilities without testing it, and Rozanova and Temerev put it outside their scope. No earlier work re-reads the glyph stream under route geometries and scores the readings on entropy and repeats with known-key controls.
60Order-invariant phrase-repetition testNewsamenone foundTimm and Zandbergen recorded the missing repeated phrases. The statistic here counts repeated runs of letter bags, which anagramming within words and relabelling cannot change. It was not found in earlier work.
61Entropy constraint under transpositionAdaptedsamePelling 2018, Cipher Mysteries, (qualitative: adjacent-letter structure rules out arbitrary anagramming)The argument is standard cryptanalysis, and Pelling made it in words for the Voynich text. The numbers here come from Latin of matched size transposed with known keys. They show the one exception, sorted anagrams, which keep h2 low.
62Codebook / nomenclator simulation gridNewsamenone foundThe rank-matched codebook generator (row E6) is extended into a grid of homophones, nulls and word splitting, scored on the word-order information as a budget and on the 1-edit rates. Pelling argued in words that the text has too few shapes for a nomenclator. No earlier quantitative simulation of a nomenclator against the Voynich statistics was found.
63Words-as-numerals test via slot dependenceAdaptedNewFeng & Hu 2016 (supervisors Abbott & Ng), Adelaide project 'Cracking the Voynich manuscript code'…Feng and Hu pursued the reading of the words as numbers by pattern matching. The test here measures how much one end of a word tells about the other, with generated numeral tables as positive controls.
64Labelling-free leading-unit distribution testNewsamenone foundNo earlier application of Benford's law to the text was found. The test turned out to have no diagnostic power, so nothing rests on it.
65Ordered-table testsNewsamenone foundThe ordered-table idea has been discussed on voynich.ninja. No earlier quantitative test of it with generated tables as controls was found.
66Syllabic and letter-pair codes of real languagesAdaptedsameStolfi 2002, 'Chinese theory redux', (syllable-inventory argument for an East Asian syllabic reading)Stolfi, and Rozanova and Temerev, compared a real syllabic text, Chinese in pinyin, with the Voynich lexicon. Here European texts are written out in syllables and in letter pairs. The entropy and inventory comparison and the argument that a word-level code cannot change the repeated 4-gram count are added.
67Derived-channel batteryAdaptedsameMatlach, Janeckova & Dostal 2022, PLoS ONE 17(1) e0260948…Matlach, Janeckova and Dostal proposed one steganographic scheme and one diagnostic, the autocorrelation of symbol reuse. Here 29 derived channels are each tested for any learnable sequential structure, with held-out models and shuffle nulls.
68Planted-message positive controlsAdaptedsameMatlach, Janeckova & Dostal 2022, PLoS ONE 17(1) e0260948…Matlach, Janeckova and Dostal built a Voynich-like text from a message. Here a message is planted into one derived channel of the real word sequence. The power of the channel test to detect it is then measured as a function of length.
69Chain-rule decomposition of the bias-corrected adjacent-word MI by slotAdaptedsameRozanova & Temerev 2026, arXiv:2608.17096…Rozanova and Temerev, and Smith, measured the coupling between the glyphs at word edges. Here the whole word-to-word information is split by the chain rule into the parts that the prefix, the core and the suffix of the next word account for.
70Word recurrence by physical distance classAdaptedsameMontemurro & Zanette 2013, PLoS ONE 8(6)…Montemurro and Zanette, Schinner and Timm established that words recur over long ranges. Here the distances are binned by the physical make-up of the book (line, page, leaf, quire, the gathering of leaves), with an exact expectation and a null stratified by hand and Currier language. The result is compared with fitted copying and page-subset generators.
71Held-out within-page predictive gainAdaptedsameSterneck, Polish & Bowern 2021, arXiv:2107.02858…Sterneck, Polish and Bowern ran topic models on the pages. Added here are a held-out predictive gain with a buffer between the halves and stratification by hand and Currier language. The tests also show that LDA, the topic model, has no power at this size even on natural-language controls.
72PERMANOVA pseudo-F on page-by-page JSDAdaptedsameMontemurro & Zanette 2013, PLoS ONE 8(6), (section vocabularies and links between sections)Montemurro and Zanette, Zandbergen and Currier established that sections have their own vocabulary. Here the test is a PERMANOVA, a permutation test of group differences on a distance matrix. Its nulls preserve contiguity or are stratified by hand and Currier language, and the pages are subsampled first to correct the entropy bias.
73Image-text association on the herbal pagesNewsamenone foundMontemurro and Zanette, and Sterneck, Polish and Bowern, compared vocabulary with illustration themes at the section level. No earlier page-level statistical test links measured image features to page vocabulary with a planted positive control.
74Function-word testAdaptedsameMontemurro & Zanette 2013, PLoS ONE 8(6)…Montemurro and Zanette applied the burstiness criterion for function words to the text. Added here are the contrast of the top-30 words with the mid-band under stratification by hand and Currier language, and the positional and neighbour-entropy features. A page-subset generator reproduces the Voynich pattern.
75Label testsAdaptedsamePelling 2017, 'Voynich labelese'Pelling and Zandbergen discuss how labels differ from running text. The shuffle and draw nulls for recurrence across pages, and the derangement null for similarity to the same page, are additions here.
76Held-out Kneser-Ney cross-entropy against n-gram order 1-8AdaptedsameLindemann & Bowern 2021, arXiv:2010.14697, (conditional entropy h2 across 316 texts)Lindemann and Bowern, Parisel and Zandbergen stop at low orders and plug-in estimates. Here the gain beyond order 3 is cross-validated, and the learning curve against training size is compared with imitations of the text at matched order.
77Word-level held-out bigram gainAdaptedsameRozanova & Temerev 2026, arXiv:2608.17096, (shuffle-corrected adjacent-token MI)Rozanova and Temerev, and Parisel, estimate the same quantity by plug-in counts corrected with a shuffle. Here it is a held-out predictive gain.
78PPMI-SVD word-embedding geometryEstablishedAdaptedPerone 2016, 'Voynich Manuscript: word vectors and t-SNE', (word2vec + t-SNE on Voynich words)Perone, Baum and colleagues, and an EPFL course paper built word embeddings of the text. Added here are matched meaningless controls, and the hubness result comes from them.
791-edit word pairs as distributional neighboursAdaptedsameEPFL MLO course paper 2022, 'Word embeddings for the morphosyntactic analysis of the Voynich manuscript'…The EPFL course paper and Perone used embeddings to study word variants. The test here, 1-edit pairs against rank-matched pairs in nearest-neighbour lists, with split-half and context-restricted controls and generator baselines, was not found.
80Unsupervised cross-lingual embedding alignmentEstablishedAdaptedBaum et al. (n.d.), voynich2vec (seminar project under C. Bowern)…Baum and colleagues tried the alignment in a seminar project. The result here is negative. The known-answer and supervised controls show that the method has no power at this corpus size.
81Extended language surveyAdaptedsameLindemann & Bowern 2021, arXiv:2010.14697…Lindemann and Bowern, Hauer and Kondrak, and Matlach and colleagues surveyed many languages on one statistic each (entropy, language-model likelihood, autocorrelation). Here the profile has several statistics that a fixed cipher cannot change, with a combined distance ranking, and it includes abjad and Indic corpora.
82Offset sensitivity of the repetition countsEstablishedAdaptedRozanova & Temerev 2026, arXiv:2608.17096, (quire bootstrap and disjoint-half resampling)A check of sampling variability. No new technique.
83Known-key cipher positive controls in Hebrew, Arabic and Sanskrit and the solver gridEstablishedsameHauer & Kondrak 2016, TACL 4, (HebrewThe substitution solver (row F1), which maps each symbol to 1 letter, and its meaningless controls (row F4) are run on more languages. The technique is the same.
84Drifting-state generatorAdaptedsameTimm & Schinner 2020, Cryptologia 44(1)…Timm and Schinner's generator clusters words by copying nearby ones, and other generators use fixed page or line vocabularies. The generator here makes the clustering a hidden state with fitted timescales. It is fitted only to the clustering statistics, and every other statistic is reported as an out-of-sample prediction. No latent-state generator for this text was found.
85Generator fitting protocolAdaptedsameTimm & Schinner 2020, Cryptologia 44(1)…Timm and Schinner, Parisel, Rozanova and Temerev, and Rugg report which statistics their generators match. The protocol here (the rules the fitting follows) states the objective, its tolerances (the range a statistic wanders over between halves of the text) and the seed variance (the spread between runs with different random seeds), and it separates the statistics that were fitted from those that are predicted.
86One-parameter verbatim pair memoryAdaptedsameParisel 2026, arXiv:2604.19762, (slot-based parametric generator with 12 ablationsParisel's word-Markov simulations fit the whole transition table. Here one memory rate inside the drift generator is fitted to one statistic, and the longer-range information and the held-out gain are the checks.
87Section-base variant of the drift generatorAdaptedsameMontemurro & Zanette 2013, PLoS ONE 8(6), (section vocabularies and links between sections)Montemurro and Zanette, and Sterneck, Polish and Bowern, measured the section vocabularies. Whether a fitted drift generator reproduces them, and at what cost to the other statistics, had not been tested.
88Re-implementation self-checkAdaptedsameRozanova & Temerev 2026, arXiv:2608.17096…A verification step. It is not a new measurement, and it is listed because the generator's scores depend on it.
89Variant genealogy along codex orderAdaptedsameTimm & Schinner 2020, Cryptologia 44(1), (network of similar words spanning the manuscriptTimm and Schinner infer the process from the network of similar words. The test here asks whether the codex shows the ordering that a vocabulary evolving during writing would leave. It uses page-order permutation nulls and positive controls for a global and a local genealogy. It finds no ordering.
90Fit-to-space at the line end measured on the scansNewsamenone foundVogt, Feaster and Smith measure word length and glyph forms by position in the transliteration, and Zandbergen describes the right margin in words. Measuring the physical space left at the margin, against a simulated copyist filling the measured width, was not found.
91Conjugate-leaf vocabulary sharingAdaptedsameFagin Davis 2020, Manuscript Studies 5(1) (five scribes identified palaeographicallyFagin Davis, Zandbergen and Pelling raised the idea that the bifolium (a sheet folded once to give two leaves) was the unit of production, from palaeography and codicology. The test here uses the vocabulary, with controls matched by distance, hand and language, and with language texts re-flowed into the same pages.
92Ink change points against vocabulary driftNewsamenone foundThe ink has been described by eye on voynich.ninja and analysed chemically by McCrone Associates. A page-by-page ink series compared with a vocabulary-drift series was not found.
93Transcription-free glyph segmentation of full-resolution scansAdaptedNewPonzi (n.d.), 'Neural-network handwritten text recognition for the Voynich manuscript', Medium/ViridisGreen…Ponzi's pipeline is a supervised recogniser trained on transliterations. The pipeline here cuts units without any training and aligns them to the transliteration by dynamic programming. Every crop gets an EVA label, and the quality of the alignment is stated for each line.
94Global shape clustering of aligned glyph cropsNewsamenone foundNo unsupervised clustering of glyph images against the transliteration alphabet was found for this manuscript. Confidence is low, because forum or blog experiments of this kind may exist.
95Allograph-sequence channel testNewsamenone foundTranscribers have recorded the form variants, and Painter and Bowern and Fagin Davis used them to separate hands. Testing the sequence of variants for message-like structure, with shuffle nulls, the scribe's habits as covariates and a substituted-language control, was not found.
96Whole-shape stream statisticsAdaptedsameZandbergen, (entropy under different transcription alphabetsZandbergen, and Rozanova and Temerev, measured entropy under alternative transliterations. Applying it to a symbol stream derived from the images, and reading the excess as clustering noise, is the extension.
97Held-out information budgetAdaptedsameZandbergen, (word entropy 9.9 bits at 8,000 words vs Dante 9.1, Pliny 10.6Zandbergen's estimate is a word entropy, and Bennett's, Lindemann and Bowern's and Parisel's are character entropies. The model here is a held-out structured model. Each stage prices one known regularity, and the residual is compared with generators of the same shape.
98Residual-channel testsNewsamenone foundMatlach and colleagues, and Rozanova and Temerev, ran sequential tests on symbol streams. Testing the residual of a fitted predictive model as a channel, stream by stream, against generators of the same shape was not found.
99Planted low-rate channel control and residual decipherment stepAdaptedsameMatlach, Janeckova & Dostal 2022, PLoS ONE 17(1) e0260948…Matlach and colleagues planted a message in a text, and Hauer and Kondrak and Rozanova and Temerev validated solvers on known keys. Applying those controls to the 2 residual streams of a fitted model, prefix and suffix, with 1 planted letter per word, is the extension.
100Naibbe cipher re-implemented from the published code with fixed seedsEstablishedsameGreshko 2025, Cryptologia, (Naibbe cipherGreshko compared the Naibbe ciphertext with the text on fewer statistics, and Rozanova and Temerev on a few more. Here the full battery, the content and embedding diagnostics and the order-2 Markov imitation are added. The extra statistics do not change the class.
101Substitution solvers on Naibbe ciphertextAdaptedsameRozanova & Temerev 2026, arXiv:2608.17096…Rozanova and Temerev score the Naibbe output with 1 language-match differential. Here the three substitution solvers and the key-transfer test are run on it, scored against the known plaintext, and the failure is stated as a limit of the solvers.
102Homophone-exchangeability testAdaptedsameRozanova & Temerev 2026, arXiv:2608.17096, (Brown clustering of 88 glyph units into 23 classes by context)Rozanova and Temerev, Matlach and colleagues, and Guy group glyphs or units by their contexts. The test here asks whether frequent word types are exchangeable in context, with a permutation null and a positive control of known homophone classes.
103Spelling-variation sensitivityAdaptedsameBowern & Lindemann 2021, Annual Review of Linguistics 7…Bowern and Lindemann, and Timm and Schinner, compare with real historical texts. Here synthetic variation is added in measured doses, and the repetition, entropy and family statistics are read together.
104Transcription-collapse sensitivity on the Voynich sideAdaptedsameZandbergen, (entropy under different transcription alphabetsZandbergen, and Rozanova and Temerev, report entropy and repeat counts under merged alphabets. The progressive collapse here, with shuffled chance levels from 3 seeds, the 1-edit family collapse and the matched collapses of Latin and Italian, extends the variant-collapse recount.
105Verbose-expansion matchingAdaptedsameRozanova & Temerev 2026, arXiv:2608.17096, (variable 1-3-glyph groups in the calibrated attackTiltman raised the verbose-cipher idea long ago, and Rozanova and Temerev calibrated the group sizes. Reading the repetition threshold in plaintext-equivalent lengths under both size conventions was not found.
106Eighteen deterministic state-dependent encodings of Latin in the skeletonAdaptedsameBennett 1976, Scientific and Engineering Problem-Solving with the Computer (character entropiesThe entropy argument against polyalphabetic substitution is standard, from Bennett onward. Here it is applied together with the repeat, family and information-placement statistics to eighteen fully specified schemes, among them autokey, state machines and counter-chosen nulls.
107Solver-detection check on state-dependent ciphersAdaptedsameRozanova & Temerev 2026, arXiv:2608.17096, (calibrated attack with Markov surrogates)Hauer and Kondrak, and Rozanova and Temerev, validated their solvers on known keys. Running the detection protocol on ciphers outside the set of ciphers the solver can find, to state what it covers, was the added step.
108Diplomatic medieval reference setAdaptedsameLindemann & Bowern 2021, arXiv:2010.14697, (18 transcribed historical texts in 8 languagesLindemann and Bowern, Gheuens and Boxer compared a few historical texts, mostly on entropy. Here the transcriptions are diplomatic, with the abbreviation marks kept as symbols, and the same text appears at several levels of faithfulness. The full repetition, family and adjacency profile is measured with class ranges over windows.
109Transcription-policy effect table on the reference sideAdaptedsameLindemann & Bowern 2021, arXiv:2010.14697, (abbreviation and scribal-convention effects on entropy examined)Zandbergen, and Rozanova and Temerev, examined the effect of transcription policy on the Voynich side (alphabets, separators), and Lindemann and Bowern examined it for entropy on the reference side. The step-by-step table over the whole profile of statistics is the extension.
110Simulated scribal abbreviation and spelling variation as a rule modelAdaptedsameLindemann & Bowern 2021, arXiv:2010.14697, and Bowern & Lindemann 2021, Annual Review of Linguistics 7…Synthetic variation on reference texts is shared with the spelling-variation test (row R1). The calibration against the scribal inconsistency measured in real manuscripts is the added step.
111Word co-occurrence network topologyEstablishedsameAmancio, Altmann, Rybski, Oliveira and Costa 2013, PLoS ONE 8(7) e67310, (EVA transcription)Amancio and colleagues made the same measurement. Added here are a second null, the line-interior shuffle, which removes the assortativity excess, and 14 language chunks of matched size. Also added are 9 generators poured into the same page and line skeleton and a triad census verified against the reference implementation.
112Intermittency of word recurrenceEstablishedsameAmancio et al. 2013Amancio and colleagues compared the text with one book in 15 translations. Here the reference set is 28 corpora of matched size, some of one genre and some of mixed topics. Each Currier language, each scribal hand and nine generators are scored separately.
113Information over scales curveEstablishedsameMontemurro and Zanette 2013, PLoS ONE 8(6) e66344Montemurro and Zanette made the same measurement. Added here are stratified, within-page and within-paragraph shuffles, which take the curve apart, and a moving-block bootstrap on the position of the peak (512 in 156 of 200 resamples). Nine generators are scored too, the fitted drift generator among them.
114Keyword ranking by per-word information or block entropy with bootstrap significance, and the share of significant keywordsEstablishedsameMontemurro and Zanette 2013Montemurro and Zanette made the same measurement. Here the identical criterion is applied to 28 languages of matched size and nine generators. The counts are repeated per hand and per Currier language, and the resulting keyword list is tested for concentration in one section.
115Paragraph-similarity network modularity against paragraph-shuffled textEstablishedNewde Arruda, Marinho, Costa and Amancio 2018, Paragraph-based complex networks: application to document…De Arruda, Marinho, Costa and Amancio compared the text with a paragraph shuffle. Here the reference set is 28 languages of matched size and nine generators poured into the same paragraph skeleton. A null stratified by hand and Currier language is added beside the paragraph shuffle.
116Adjacent-page similarityEstablishedsameReddy and Knight 2011, What we know about the Voynich manuscript, LaTeCH at ACL…Reddy and Knight made the same measurement. Added here are a free permutation null and one stratified by hand and Currier language, a tf-idf variant, two transliterations, 28 languages, nine generators and a planted positive control.
117Two-state letter HMM and the a*b word grammarEstablishedsameReddy and Knight 2011Reddy and Knight, and Acedo, made the same measurement. Added here are three symbol views in place of one and random restarts beside the a*b initialisation. The within-word shuffle and the Markov-2 imitation are the nulls, and Latin without vowels, Hebrew and Arabic are the positive controls.
118Word-class inductionEstablishedsameReddy and Knight 2011Reddy and Knight made the same measurement. Added here are the interior-shuffled text as a floor for the gain, a word-Markov generator as a ceiling and five folds. The induced classes are compared with the function words of the Latin control.
119Reading-direction cross-entropy asymmetry Delta_n = X_LTR - X_RTLEstablishedNewParisel 2025, arXiv:2509.10573, (Kaggle code linked from the paper)Parisel made the same measurement. Added here are three transliterations, 15 languages, the within-word and character shuffles (both at zero) and five generators. Two smoothing families and three orders are run, and a worked reconstruction shows what an asymmetric edge convention would produce.
120Boundary signaturesEstablishedNewParisel 2026, arXiv:2604.19762Parisel made the same measurement. Added here are three transliterations, the within-word shuffle as a negative control, and the Markov-2 imitation and five generators scored on the same four signatures. The boundary information is bias-corrected by the line-interior shuffle.
121Internality index of uncertain word separators and line breaksEstablishedNewRozanova and Temerev 2026, arXiv:2608.17096Rozanova and Temerev made the same measurement, but their description of the index admits two readings, and both are computed here. Added are four symbol views and two transliterations, 17 poured languages and a planted internal-boundary positive control. Three groups of Middle High German manuscripts whose scribal and editorial boundaries are known are added as well.
122Physical gap width on the scans as a predictor of transcriber space certaintyEstablishedNewRozanova and Temerev 2026Rozanova and Temerev made the same measurement from published word boxes. Here the gaps are measured from the scans. Added are blind predictions written down before the values were read, a within-line permutation null and a breakdown by hand. A positive control compares certain spaces with within-word gaps (AUC 0.913).
123Dependence gap across BPE merge levels D_k = HEstablishedNewRozanova and Temerev 2026Rozanova and Temerev made the same measurement without controls. Added here are three space conventions, a page-block bootstrap, the Miller-Madow correction, nine languages, the within-word shuffle, the Markov-2 imitation and five generators.
124Quire bootstrap and leave-one-quire-out confidence intervals for the headline statisticsEstablishedNewRozanova and Temerev 2026Rozanova and Temerev used the same resampling design. Added here are a page bootstrap run beside the quire bootstrap for comparison, and matched PERMANOVA nulls without the two most influential quires. The tests also show that resampling with replacement corrupts every repetition and information statistic.
125Table-and-grille generatorEstablishedsameRugg 2004, Cryptologia 28(1):31-46, (abstract only)The generator family is Rugg's. Added here are the full 23-statistic battery with page-half confidence intervals, extra repetition and boundary-information statistics, and 168 configurations over five seeds. A null device without locality isolates the effect of renewing the table.
126Wheel (volvelle) grille generatorAdaptedsameZandbergen 2021, The Cardan grille approach to the Voynich MS taken to the next level, arXiv:2104.12548Zandbergen's proposal is a sketch of a mechanism. Added here is a direct product-set test, which asks how much of the real vocabulary a bounded wheel product can cover. The full 23-statistic battery is run as well, with extra statistics for lexicon growth, the recurrence profile and boundary information, over six devices and five seeds.
127Human-produced gibberish reference class and the 42-metric random-forest classifierEstablishedsameGaskell and Bowern 2022, Gibberish after all? Voynichese is statistically similar to human-produced samples…Gaskell and Bowern's classifier and metrics are used unchanged. Added here are the two shuffles and the Markov-2 imitation passed through the same forest. The full battery is run on all 38 gibberish and 71 meaningful texts at their own 200-word scale, and a pooled profile is ranked against 29 languages.
128Detrended fluctuation analysis, Hurst exponent and spectral slope of word-level seriesEstablishedsameSchinner 2007, The Voynich Manuscript: evidence of the hoax hypothesis, Cryptologia 31(2):95-107…Schinner, and Montemurro and Pury, report an exponent or an autocorrelation without taking it apart. Here three targeted shuffles (line-interior, line-order, page-order) locate the correlation between the line and the page, and 30 languages, six generators and the residual-surprisal series are added.
129Spectral analysis of the spaceless glyph streamEstablishedsameLandini 2001, Evidence of linguistic structure in the Voynich manuscript using spectral analysis, Cryptologia…Landini and Arutyunov made the same spectral measurement. Added here are ten shuffle seeds, a white-noise shuffle, the within-word shuffle, the Markov-2 imitation and four languages. A second symbol view with merged glyphs is used, and the line-band component that the published abstract only asserts is tested directly.
130Distribution of the distance to the next similar word and its test against the geometric lawEstablishedsameSchinner 2007 (abstract only)Schinner and Timm tested the same distance law. Added here are a hazard-ratio decomposition by distance band, the text's own shuffle as the memoryless null, 30 languages and seven generators. A copying positive control and a page bootstrap on the short-distance excess are added as well.
131Word and glyph mutual information across line and paragraph boundariesNewsameCurrier 1976 (line as a functional entityNo earlier measurement of mutual information across the line break was found. Currier called the line a functional entity, and Landini and Vogt took the idea up. The design here has three independent nulls: a same-page permutation, a generator with a line-start state as the negative control, and poured languages as the positive control (z 4-43). Two transliterations and a glyph-level counterpart are added.
132EnochianEstablishedsameBoxer 2022, Fingerprinting gibberish: a quantitative comparison of the Voynich and Sloane MS 3188, CEUR-WS…Boxer's two statistics are used unchanged. Added here are the word-order shuffle of each text as a drift baseline, page-block resampling without replacement and the Markov-2 imitation. The comparison set has 29 languages, and the full battery is run at the Enochian sample size.
133Polygraphia III style letter-to-invented-word cipher as a reference textEstablishedsameHermes 2022, Polygraphia III: the cipher that pretends to be an artificial language, CEUR-WS Vol-3313 paper 7Hermes compared the cipher with the text descriptively. Here the cipher class is generated at matched size over 14 table settings and three seeds. It is scored on the same 23-statistic battery as every other generator, with extra repetition and boundary-information statistics.
134Non-linguistic symbol corporaNewsameSproat 2014, A statistical comparison of written language and nonlinguistic symbol systems, Language…Sproat's and Nair's scorecards had never been applied to this text. Here they are, under two stated mappings (glyph in word, word in line), with permutation nulls for every scorecard statistic, the within-word and word-order shuffles and the Markov-2 imitation. The published calibration ranges were re-derived as a check.
135Genuine historical ciphertextsNewsameKnight, Megyesi and Schaefer 2011, The Copiale cipherNo genuine ciphertext had been compared with this text. The comparison here is made at matched size on symbol streams without spaces. A homophonic cipher of the same plaintext was built as a positive control, a symbol shuffle is the null, and 29 languages and three synthetic cipher families are measured alongside.
136Index of coincidence, periodic IoC and Kasiski spectra at lags up to 200, and the cipher-type feature vector of Nuhn and KnightAdaptedsameNuhn and Knight 2014, Cipher type detection, EMNLPThe index of coincidence is a classical statistic, going back to Friedman. Here it is run at every lag up to 200 in three symbol views, with permutation nulls. A full set of synthetic positive controls is added (Vigenere of three periods, autokey, word-keyed, state-dependent, verbose, homophonic, transposition), and the index is measured page by page beside matched Latin.
137Word-pattern (isomorph) fingerprintNewsameHauer and Kondrak 2016, Decoding anagrammed texts written in an unknown language and script, TACL 4:75-86Hauer and Kondrak use the word patterns as a decipherment tool for anagrammed text. Here the distribution of patterns is a language fingerprint, measured at matched size with a sampling floor, three symbol views, a monoalphabetic positive control, and generator and cipher controls.
138Glyph-level homophone detection by hierarchical clustering of symbol contextsNewsameLehofer 2022, Applying hierarchical clustering to homophonic substitution ciphers using historical corpora…Lehofer demonstrated the method on historical ciphers. Here it is applied across three symbol views, with two synthetic homophonic positive controls and natural-language negative controls. The within-word shuffle is a sensitivity check, and a token alignment identifies which fine distinctions are exchangeable.
139Kober triplets and paradigm consistency across stemsAdaptedsameKober 1948, The Minoan scripts: fact and theory, American Journal of Archaeology 52:82-103…Kober's heuristic is qualitative. Here it becomes three quantitative statistics with a permutation null. They are applied to nine languages, three symbol views and five generators, with a monoalphabetic positive control and Latin written in syllables as a second positive control.
140Scribe attribution by character n-gram classifiers and Burrows' Delta against the Fagin Davis handsEstablishedsameFarrugia, Layfield and van der Plas 2022, CEUR-WS Vol-3313 paper 5, (ZL transcription)Farrugia, Layfield and van der Plas used the same classifiers. Added here are a label permutation null within strata of Currier language and section, sub-tests within one language and within one section, a poured-Latin negative control and the drift generator. Two positive controls are added as well: one text in two simulated scribal habits (0.96-0.97) and two manuscripts by different scribes (1.00).
141Rightwardness and downwardness scores of graphemic minimal word pairsEstablishedsameFeaster 2022, Rightward and downward grapheme distributions in the Voynich manuscript, CEUR-WS Vol-3313 paper…Feaster's minimal-pair test is used unchanged. Added here are a within-line permutation null, an interior-only variant that removes the known line-edge effect, and a control family of every other single-glyph substitution. Both transliterations are used, poured languages are scored on their own pairs, and word-final m and g are the positive control.
142Per-scribe paragraph-initial habits and the interaction of within-word position with position in the paragraphEstablishedsameStafford 2022, Seven habits of highly eccentric paragraphs, CEUR-WS Vol-3313 paper 13Stafford described the same habits. Added here are exact counts for all 740 paragraphs with a page bootstrap, and two permutation nulls for the distance between hands (within section and across all pages). Poured Latin and Italian, the within-word shuffle and the Markov-2 imitation are the controls for the position interaction.
143Segment information up to the uniqueness point as a function of word frequencyEstablishedsameLayfield, van der Plas, Rosner and Abela 2020, Word probability findings in the Voynich manuscript, LT4HALA…Layfield and colleagues made the same measurement. Added here are the within-word shuffle, the Markov-2 imitation and six generators run through the identical trie pipeline, and 30 languages at matched size. A partial correlation on word length and a breakdown by frequency band are added as well.
144Word length vs frequencyAdaptedsamePiantadosi, Tily and Gibson 2011, PNAS 108(9):3526-3529Piantadosi, Tily and Gibson used very large corpora and a surprisal model trained on them. Here both are estimated on matched samples of 34,780 tokens with a cross-validated surprisal, and the within-word shuffle, the Markov-2 imitation and six generators are scored alongside 30 languages.
145Menzerath-Altmann law between word length in units and mean unit lengthNewsameMenzerath-Altmann law (Altmann 1980, Glottometrika 2)No value of this law had been published for the text. Here it is fitted through one segmentation pipeline applied identically to every text, with the within-word shuffle as the null. Three alternative segmentations, six languages, five generators and a bootstrap over types are added.
146Unsupervised word segmentation of the spaceless glyph stream and its agreement with the transcribed spacesNewsameJin and Tanaka-Ishii 2006, COLING/ACLNeither of the two segmentation methods, Jin and Tanaka-Ishii's and Goldwater, Griffiths and Johnson's, had been applied to this text. Both are run here at a rate-matched operating point, with the within-word shuffle as a floor and the Markov-2 imitation as a ceiling. Three languages of matched size are scored too, and recall is measured separately at the uncertain positions.
147Long-context neural language model gain over Kneser-Ney n-gramsNewsameStandard methodNo peer-reviewed neural comparison exists for this text. The design here fits the neural model and the n-gram baseline on exactly the same held-out folds. It adds the Markov-2 imitation as a zero-gain reference and a copying generator as a positive one, and it measures the gain as a function of context length.
148Paired single-leg gallows on paragraph top linesAdaptedNewPelling, Cipher Mysteries blog posts of 2020-08-29 and 2025-07-12, (grey literaturePelling's claim is grey literature with no statistics. Here it is given a binomial placement model, a within-line permutation null and a breakdown by hand. A Markov-2 imitation control and two placebo families in languages, chosen by the same rule, are added.
149Stroke-level symbol viewAdaptedNewCham 2014, Curve-line system, (grey literature)Cham's proposal is grey literature with no numbers. Here a documented stroke table is written out and applied to the text, and to Latin and Italian with a comparable table. Two table permutation controls separate the effect of the recoding from the text.
150Fractal dimension, sliding-window network centrality maps and visibility-graph statistics of text seriesEstablishedsameZelinka, Lara, Windsor and Lozi 2023, Softcomputing in identification of the origin of Voynich manuscript by…Zelinka and colleagues made the same measurements. Added here are the text's own word shuffle and within-word shuffle as nulls, the drift generator and the Markov-2 imitation, and six languages. A language-to-language distance is the yardstick for differences between maps.
151Two-language mixture test from the deviation of sorted letter frequencies from a logarithmic law, and spectral portraits of two-letter distributionsEstablishedNewArutyunov, Borisov, Fedorov, Ivchenko, Kirina-Lilinskaya, Orlov, Osipov, Pyrkov, Shilin and Zeniuk 2016…Arutyunov and colleagues argue from a fit statistic on sorted frequencies. Here the statistic is given single-language references of matched size and constructed bilingual references in the same page skeleton. A held-out two-component mixture model with a third-component check, restrictions to one hand and to one Currier language, and a spectral portrait of the bigram matrix are added.
152Phonetic-prior decipherment into candidate known languages with a language-closeness scoreNewsameLuo, Hartmann, Santus, Barzilay and Cao 2021, Deciphering undersegmented ancient scripts using phonetic…Luo and colleagues' released model needs a GPU and phonetic transcriptions of every candidate vocabulary. The substitute here scores closeness as the gain of a validated homophonic solver on the text over the same solver on the text's own Markov-2 imitation. A monoalphabetic cipher is the positive control and a language isolate the negative control.
153Autoencoder visual similarity of glyph shapes within the Voynich alphabet and to the letterforms of other scriptsEstablishedsameZelinka, Lara, Windsor and Lozi 2023, open copyZelinka and colleagues made the same autoencoder comparison from font renderings. Here the glyphs are aligned crops from the scans. Added are a null script of random strokes and the spread within a class as the scale of 'same shape'. The positive control is a set of known alphabets rendered in other fonts and distorted, which must find their own script.
154Unicity-distance table per cipher familyAdaptedsameShannon 1949, Communication theory of secrecy systems, Bell System Technical Journal 28(4):656-715, (theoryVon zur Gathen computed the unicity distance for one historical cipher, the Zodiac-340, with one key count and one corpus entropy. The published solver calibrations report accuracy against ciphertext length for English and do not refer to the unicity distance. The one computation in a Voynich context is a rough forum figure for one proposed cipher. Here the distance is tabulated for every cipher family against the manuscript's own ciphertext lengths, in three symbol readings, per section and for the labels. The redundancy is the most conservative of four estimators over nine candidate languages. The Naibbe key is counted three ways, with its expansion and randomisation folded into an effective redundancy. The solvers are calibrated against the distance on Latin ciphers with a known key. Each solver failure is then sorted into one of two kinds. It is informative when the text is far above the distance and the key space was searched. It is uninformative when the family lies outside the searched key space, or when the text is at or below the distance.
155Book of Soyga tables regenerated from Reeds's recurrenceAdaptedsameDaruka 2021, On the Voynich manuscript, Cryptologia 45(1):44-80 (online February 2020)…Daruka made the one published comparison of the manuscript with a magical table text, Dee's Liber Loagaeth. How that text was generated is unknown, and the comparison reports statistical-linguistic matches at the whole-text level. The Voynich literature describes the Book of Soyga as algorithmically generated, but nobody had measured it as text. Here the tables are rebuilt from Reeds's published recurrence, so the generator is known exactly. They are scored on the full battery at matched size, in both reading directions and in a word view poured into the manuscript's page skeleton. The within-word shuffle, the Markov-2 imitation, three languages and a positive-control language are scored on the same statistics.
156Hildegard of Bingen's Lingua Ignota scored as a type listAdaptedNewveriarch.com 2025, The Abbess's Code: testing Hildegard's lingua ignota, unsigned web page dated 21 August…The one earlier measurement found is on an unsigned web page: a normalised character entropy of the glossary against Latin, Middle High German and a random string. The published treatments of the Lingua Ignota, Higley's edition and Skowronska's classification, are qualitative. Boxer's quantitative comparison of an invented vocabulary with the manuscript concerns Enochian running text, on once-only words and repetition. Here the glossary is scored as a type list on the word-level statistics used throughout, beside size-matched type lists of the manuscript, three languages, Enochian and the Trithemius conjurations. The controls are the letters-within-word shuffle and a Markov-2 word chain fitted to the list. Two independent sources of the glossary are checked against each other, and a permutation test is run on its semantic ordering.
157The conjurations of Trithemius's SteganographiaAdaptedsameReeds 1998, Solved: the ciphers in Book III of Trithemius's Steganographia, Cryptologia 22(4):291-317…Reeds, Ernst and Ponzi solved and described the Steganographia's ciphers. Boxer (2016) built tools that apply its techniques to any text and rank the outputs by letter-frequency divergence. Matlach and colleagues proposed a Trithemius-style scheme for the manuscript and tested it by symbol autocorrelation. Hermes profiled the output of Polygraphia III against the manuscript. Nobody had measured the conjurations as text, or planted the alternate-letter concealment described for them into a carrier of the manuscript's length. Here both are done, with the battery and the channel score. A Markov-2 imitation fixes the channel's baseline. An invented-word cover is generated from the conjurations, a wrong-phase control is added, and a detection curve is drawn from 500 to 5,000 channel symbols.
158Known-plaintext and dictionary attack on the label setsAdaptedsameZandbergen, Some special properties of labels in the Voynich MS, voynich.nu, updated 11 September 2025…The label literature reports frequency properties. Examples are the share of labels that occur once, ok and ot as the initials of 52% of zodiac labels, and labels beginning with o at probability 0.48. Bax, Sherwood, Tucker and Talbert offer crib readings of single labels without a test. The one dictionary attack found, a genetic-algorithm n-gram mapping of 111 herbal first words with 31 hits, has no null. Hauer and Kondrak's decipherment attack treats the whole text. Here the solvers are run on the label corpora alone, against their own Markov imitations, with a planted positive control at the same size and budget. The dictionary search reports the largest number of labels that one key maps onto a name list, beside imitation and running-text controls and a planted key. The first-glyph concentration is compared with real name lists, as a statistic that no one-to-one or homophonic relettering can change. The published identifications are scored under a null that shuffles the identifications.
159Held-out model competition on random quire splits with the scoring set fixed in advanceAdaptedsameTimm and Schinner 2020, A possible generating algorithm of the Voynich manuscript, Cryptologia 44(1):1-19…Timm and Schinner (relative errors), Rugg and Taylor (frequency and length features) and Parisel (four signatures) each compared one generator's output with the whole text on a list of statistics. In each case the generator was tuned on the same text. Cross-validated model selection is standard elsewhere but had not been applied to Voynich generators. Here every generator is refitted on half the quires and scored on the other half. The unit of distance is the between-quire jackknife error, and the floor is the text's own half-to-half distance. The scoring set was fixed before fitting, five splits give a stability measure, and the shuffles and the glyph imitation are placed on the same scale as the generators.
160Held-out discriminatorAdaptedsameGaskell and Bowern 2022, Gibberish after all? Voynichese is statistically similar to human-produced samples…The published Voynich classifiers separate meaningful text from human gibberish, Currier A from B, or scribe from scribe on held-out folios. Gaskell and Bowern's random forest on 42 metrics, with the manuscript as a test case, is of the first kind. In other fields a classifier is used as a two-sample test to score generative models. Here the two classes are the manuscript's own pages and a generator's pages in the same skeleton. The split is by quire, so that no page's neighbours leak into the training set. The null is a label permutation refitted 100 times. A nested feature-removal analysis measures how much of the separation survives without the leading statistics. A Latin-versus-imitation positive control and an arbitrary-label negative control bracket the result.
161Akshara rendering profile with the Kober positive controlAdaptedsameLindemann and Bowern 2021, Character entropy in modern and historical texts: comparison metrics for an…Lindemann and Bowern compared the text with texts as natively written in abugidas or syllabaries, by character-set size against h2. Other work induces units from the text's own glyphs. Nobody renders one language at several sign granularities, so that the same words are measured as letters, graphemes, consonant-vowel signs and aksharas. Here one abugida-written language is rendered by rule into four sign streams and profiled per sign on the battery. The battery covers h1 to h4, word-length shape in signs, one-edit share, repetition, and shuffle and Markov nulls. The Kober grid is given its first positive control from a real syllabary. The closeness score is calibrated with a renamed-sign control. That control shows the score to have no power for the syllabary question, while the excess over the real language after the best key does have power.
162Matched-inventory PMI mergingAdaptedsameRozanova and Temerev 2026, A glyph is not a letter, a token is not a word, a space is not a space: what the…Rozanova and Temerev, and thevoynichproject, published merge trajectories that index the text and a few controls by the number of merges and report the dependence gap. Lindemann and Bowern, and the survey, vary the text's own character divisions and study the effect of set size. Bowern and Gaskell's fixed bigraph encodings go the other way, expanding letters into digraphs. Here 28 languages are brought to the text's own sign counts (24 and 38 signs for 99% coverage) by the text's own PMI unit-induction procedure. The entropy margin is then read at equal inventory, and the text's letters merged by the same code give a procedure-matched comparison. Finding 1 rests on the within-word shuffle gap at matched inventory, and not on the raw h2. Word length in signs is measured at the same targets.
163Syllable canon and harmony test from the Sukhotin classesEstablishedAdaptedGuy 1991, Statistical properties of two folios of the Voynich manuscript, Cryptologia 15(3):207-218…Guy, Reddy and Knight, and Lindemann and Bowern published the Sukhotin classes. In the grey literature, thevoynichproject ran the syllable-shape shares and a harmony test with Turkish and Finnish as positive controls, using a bipartition pure-word statistic against vowel-resampled surrogates. Ponzi and Smith scored syllable-based word grammars for coverage. Here the canon is counted as pattern-coverage numbers and parsability shares against 28 languages. The harmony statistic is a mutual information across the consonant run. The null is the text's own Markov-2 imitation, which uses adjacency only, and it reproduces every canon statistic and all but 0.005 to 0.045 bits of the vowel association. The statistic and the null differ from the earlier work. The measurements and the conclusions are the same: a phonotactically unremarkable canon, and no harmony.
164Clitic-stripping name testNewsameNearest, none making the measurement: Bowern and Lindemann 2021, The linguistics of the Voynich manuscript…The lack of qo- on labels and the preference for o-, ok- and ot- are recorded in Bowern and Lindemann's peer-peer-reviewed survey and on voynich.nu. Sazonov counted prefixes with and without a root in the running text. Label attestation in the running text has been counted here and in the grey literature. No source strips the putative proclitic layer and re-measures the relation between labels and text, and none calibrates the operation on a language with real proclitics. The logic of the test is simple. If the prefixes are clitics, stripping them should bring text words closer to the labels, as it brings Hebrew text words closer to the names. Neither that logic nor the result, that stripping moves the text away from the labels and halves attestation, is in the searched literature.
165Agglutinative slot, family and sequence controls with successor-variety signaturesEstablishedsameReddy and Knight 2011, What we know about the Voynich manuscript, ACL LaTeCH workshop…Reddy and Knight published unsupervised affix-signature extraction on the Voynich text, with Linguistica. Lindemann (2022) published the comparison with agglutinative languages on word statistics, and thevoynichproject made it in the grey literature on induced morphemes per word. The first part applies the three measurements of the earlier rows to five more languages, which adds new controls to the same measurement. The second part counts signatures at matched vocabulary with a successor-variety cut instead of Linguistica's minimum-description-length cut, and adds the Markov-2 imitation and the shuffle as nulls. The measurement, affix signatures found by unsupervised segmentation of Voynich words, is the published one.
166Dense-abbreviation simulation at Cappelli densityAdaptedsameBowern and Lindemann 2021, The linguistics of the Voynich manuscript, Annual Review of Linguistics 7:285-308…Bowern and Lindemann measured real abbreviated texts: the entropy and character-set size of diplomatic against normalised transcriptions, and the abbreviated against the plain Secreta. In the grey literature, Edwards abbreviated Latin by rule and reported mean word length, and thevoynichproject reported entropy, mean length and a merge signature under a subsequent verbose encoding. Here the abbreviation is pushed to a stated density, 177 and 229 marks per 1,000 letters, which is the Cappelli-density target of the audit. It is applied consistently and inconsistently. The statistics at issue are the word-length variance and skew and the repetition counts (rep4 and rep30), which the mean cannot show. Consistent dense abbreviation matches the mean but moves the word-length shape and h2 the wrong way. Inconsistent abbreviation lowers repetition only by raising h2 further.
167Mixture and coined-noun plaintextsNewsameNearest, none making the measurement: Arutyunov, Borisov, Fedorov, Ivchenko, Kirina-Lilinskaya, Orlov…Mixed-language readings of the text exist as proposals: Arutyunov and colleagues fitted a two-language mixture from letter frequencies, and macaronic and bilingual readings appear on the forums. The text's repeat rates have been compared with single-language corpora. No source constructs sentence-level, word-level, macaronic or coined-vocabulary mixtures and profiles them against the text on entropy, phrase repetition, adjacent-word information and one-edit density. The result is not in the searched literature. Every mixture raises h2. Only an unnatural word-by-word alternation removes the repeated phrases, and it does so at the cost of the adjacent-word information that the text keeps.
168Facsimile line-break re-parse of diplomatic manuscriptsAdaptedsameVogt 2012, The line as a functional unit in the Voynich manuscript, self-published PDF…Currier, Feaster, Stafford, and Rozanova and Temerev measured line position in the text alone. Vogt, Gnuchev, and Gaskell and Bowern compared it with natural texts cut at artificial widths. Vogt used 62 characters, Gnuchev wrapped prose, and Gaskell and Bowern a 60-character wrap where original breaks were unavailable. The diplomatic corpora with real manuscript line breaks have been used for entropy but not for line effects. Here manuscripts transcribed at the facsimile level, with allographs and abbreviation signs as symbols, are cut at their actual line and paragraph breaks. They are run through the text's own line-edge and paragraph-first statistics with a shuffle floor. That is the test of the scribal-convention reading (allographs at line edges, elongated first lines) which the artificial-wrap controls cannot make. Verse is included as the one genre with line-edge effects of the text's order.
169Two-manuscript scribal calibration of the Currier A and B divergenceAdaptedNewEdwards 2025, Voynich reconsidered: scribes and languages, Medium, 7 February 2025…Currier, Zandbergen, Lindemann and Bowern, and Parisel established the difference between A and B. The peer-peer-reviewed survey and voynich.nu say that the calibration against known-language variation has not been done. One grey source, Edwards, has since calibrated a glyph-frequency correlation against same-language and different-language document pairs. Here the calibration is against the specific case that the two-regimes reading (regime meaning one of the two writing styles) needs: two hands writing one text, and two collections of one genre. It is made at the diplomatic level, with the normalised form as the edited control, using the R1-A10 divergence and h2 statistics with a split-half null. It finds the letter divergence between A and B within a factor of two of that between two hands, and the h2 difference reproduced inside B alone.
170List-genre repetition ladderAdaptedsameZandbergen, Voynich MS, Sentences etc., voynich.nu…Zandbergen raised the list-genre explanation of the missing repeated phrases on voynich.nu. The grey literature tested it with other statistics, adjacent-repeat rates and long-range class information, on one recipe compendium and one Vulgate register. Peer-reviewed work compares the text's word-repeat rate and the 42-metric profiles of herbal texts, without a list genre or a phrase-repeat count. Here the genre is sampled as eleven entry-per-line texts of the period. They are measured at the text's own size in its own page skeleton, with the four-word phrase count that the verdict rows rest on and a word-shuffle chance level. The frequent-word head is the second reading, and positive controls and a window sweep are added. The finding is new: every bare list repeats at least 24 four-word sequences, and the two glossaries that reach the text's zero lack its frequent-word head.
171Typological-extreme language profileEstablishedsameBennett 1976, Scientific and Engineering Problem-Solving with the Computer, Prentice-Hall, chapter 4…Bennett (1976) compared the text's conditional entropy with Polynesian languages, Hawaiian first of all. The peer-peer-reviewed survey, the Lindemann and Bowern preprint (with Maori, Tongan, Samoan, Greenlandic, Inupiak and Nahuatl) and voynich.nu repeated the comparison, with Stallings's caveat about spelling. Here the same languages are read from nineteenth-century scripture in period spelling, which lowers the floor further (Maori 2.52). The instruments of the earlier rows are added: repetition counts, the 14-statistic distance, and the closeness score with known-key controls. The measurement is the published one on new corpora, and only the added legs are new here, so the technique is not new.
172Labels against period name listsAdaptedNewJB 2011, Naked ladies, Computational Attacks on the Voynich Manuscript, 31 May 2011…The grey literature has compared labels with real name lists. One blog compared them with fifteenth-century female names on length and endings, and another source matched the Lapidario's stone names by structure. The initial-glyph and hapax profiles of the labels are published on voynich.nu and in the survey's remark on nouns. Here six real lists of the period are sampled at the label sets' sizes, with window resampling and whole-list initial concentration. The labels are placed on type, length, initial-letter, attestation, one-edit and numbering statistics together, against both the lists and the running texts of their languages. The finding is not in the searched sources. The labels have a list's type-token and singleton profile, but an initial concentration (54%) and a one-edit density (37%) that no real list shows.
173Non-professional writer profile with the Paston lettersEstablishedAdaptedLindemann and Bowern 2021, Character entropy in modern and historical texts, arXiv:2010.14697…Lindemann and Bowern compared a private, non-professional early-modern hand transcribed as written, Napier's casebooks, with the text on conditional entropy. Early-spelling English is among the peer-reviewed comparison corpora on word-level metrics. Here a fifteenth-century family correspondence in original spelling is added. It is profiled on the instruments of the earlier rows: the one-edit family share and the phrase-repeat counts as well as h2. The corpus is new and the added statistics come from the earlier rows, so the measurement is the published one on another text.
174Adversarial composite batteryAdaptedsameGreshko 2025, The Naibbe cipher: a substitution cipher that encrypts Latin and Italian as Voynich…Encoded natural-language plaintexts have been scored against the text before, one encoding family at a time. Greshko scored the Naibbe cipher on seven prose plaintexts against 42 word-level metrics, and Bowern and Gaskell scored 22 manipulations by distance. Rozanova and Temerev measured unit statistics on Naibbe streams, and the earlier rows here scored fixed codes on Caesar. Here the plaintext side is chosen adversarially, from the list-like, low-repetition and low-entropy sources that the earlier items found closest, and crossed with the fixed and Naibbe encodings. The composites are poured into the text's page and line skeleton, so that the line-edge and page statistics can be scored. They are judged by a fixed 23-statistic pass count with Markov and shuffle nulls and a positive control. An oracle union bounds what any composite of these parts could match. The result, 6 of 23 at best and twelve statistics matched by none, is a bound on the whole family, which the earlier single-encoding comparisons could not give.
175Segment isomorph scanAdaptedsameFriedman and Callimahos, Military Cryptanalytics, Part II…The classical isomorph attack of Friedman and Callimahos is a manual search for pairs of isomorphic sequences in a few messages, used to reconstruct alphabets once a pair is found. Here it is made a census over all 174,000 windows at every length from 12 to 30. The pattern counts and the far-pair counts are compared with Markov and shuffle nulls and with thirteen synthetic Latin ciphers at the text's size. The statistic therefore reports the size of effect each polyalphabetic scheme would leave (70 to 405 pairs) against the text's 0. The excess at 15 to 24 glyphs is attributed by word content to the known word series. Bowern and Lindemann's published argument against polyalphabetic ciphers rests on word recurrence and does not search for disguised repeats.
176Mismatch-tolerant near repeatsAdaptedsameBaeza-Yates and Perleberg 1992, Fast and practical approximate string matching, Combinatorial Pattern…The published near-repeat work on the Voynich text is at word level: edit distance between word types, their co-occurrence within three lines, similar-word networks and the self-citation model. The gapped n-grams of the Zodiac wiki are 2 to 4 symbols long. Here whole windows of 20 and 30 glyphs, across word boundaries, are compared exhaustively at Hamming distance at most k, by the k+1-block pigeonhole of Baeza-Yates and Perleberg. They are counted against Markov and shuffle nulls in three views. The count is calibrated with sparse-homophone (p = 0.05 and 0.10) and verbose-homophone Latin ciphers at the text's size. The target is a plaintext repeat disguised by a few homophones or nulls, which is a different thing from the word families that the word-level measures describe.
177Zodiac ring depth testNewsameZandbergen, Some special properties of labels in the Voynich MS, voynich.nu, updated 11 September 2025…Every source found compares zodiac labels across rings under the identity mapping, by duplicate counts, initial-glyph shares or edit distances, or reads them as a numbering without a test. None asks whether a single glyph permutation maps one ring's labels onto another's. Here the comparison is made invariant under relettering, by pattern overlap and pattern kappa at the best cyclic shift. A maximum bijection-consistent pairing is searched for all 66 ring pairs and the star pages, against three nulls (re-dealt, Markov-imitated and borrowed labels). A planted one-list-twelve-keys control at 20% damage (19.2 recovered against 3.65) and a twelve-lists-twelve-keys negative control state the power of the null result.
178Same-alphabet chi and kappa tests between partitions of the textAdaptedsameKullback 1938, Statistical Methods in Cryptanalysis, revised edition, Register No. 193…The published comparisons between Currier languages, sections, pages and hands measure frequency differences, correlations, Euclidean bigram distances, n-gram classifiers or a mixture model on substitution ratios. They ask whether the groups differ. The cross-coincidence chi, normalised by both indices of coincidence, asks the cryptanalytic question instead. Are two groups written in one alphabet, which gives a value near 1, or are they reletterings of each other, which gives 0.4 to 0.55? Here it is run on seven partitions, including three defined by markers that no source groups by: the paragraph opener, the line opener and the stretches between gallows. It has half-split baselines, permutation nulls of the minimum, Markov seeds, six keyed Latin controls at the text's size, and a per-page-pair distribution against two-key and five-key ciphers. A direct search for a glyph permutation between A and B is added, which no source found attempts.
179Unit-level homophone exchangeabilityAdaptedsameLehofer 2022, Applying hierarchical clustering to homophonic substitution ciphers using historical corpora…The forum measures whether word pairs that differ in one unit both exist, and how their frequencies compare. That shows alternation but not exchangeability of context. Greshko's Naibbe paper builds homophonic prefix and suffix units without testing the manuscript's own. Here Lehofer's context statistic, validated at glyph level in R4-Q28, is moved to the slot grammar's prefix, core and suffix units in a position-classed spaceless stream. It has a pooled-and-resplit null, two context definitions, and Markov and drift negative controls. The Naibbe ciphertexts are the positive controls, both with the key's own units (AUC 0.93 to 1.00) and with the same inferred segmentation as the text (AUC 0.45 to 0.66). That states the power limit of the null result.
180Core-only repeat countsNewsameTimm 2014, How the Voynich manuscript was created, arXiv:1407.6639…No source counts repeated word or glyph sequences after the affixes are removed. The published and grey work describes affix variation between similar words, decomposes words into slots, or measures entropy after deleting glyph types. Here two grammars and six position reductions are applied to the text and to its shuffle and Markov imitations, each parsed with its own grammar. They are also applied to Latin with planted nulls at one or both word ends. The full repeat family (rep3, rep4, rep12 to rep30 and the longest repeat) is recounted on each. A fixed cipher of the cores with nulls or state-dependent affixes would then show as the Latin control does. There rep30 goes from 0 to 132 once the ends are dropped. The text's 0 is read at that power.
181Key-table tests for the f57v ring, the f49v and f76r glyph columns and the f66r word columnAdaptedsameD'Imperio 1978, The Voynich Manuscript: An Elegant Enigma, NSA Center for Cryptologic History, section 4.3…D'Imperio and later writers catalogue the four loci and note the f49v cycle and the near-exact f57v repetition qualitatively. Some read the ring as an instrument scale, Brumbaugh used the sequences inside a decipherment without a test, and the forums count f66r words elsewhere informally. Here each locus is put to a stated test with a null. The ring order is tested against 10,000 random orders on three correlates. The ring as an Alberti disk is scored over all shifts against 200 random rings and the Markov imitation. The columns' repeats are tested against 2,000 permutations, and the columns as running keys against 200 key permutations with two statistics. Agreement with line-initial glyphs is counted. The f66r attestation rate is tested against three draws, which turns the forum's counts into a comparison with the label classes.
182Solver size and inventory ladderAdaptedsameRavi and Knight 2008, Attacking decipherment problems optimally with low-order n-gram models, EMNLP…The published curves, Ravi and Knight's among them, report a solver's accuracy against cipher length (2 to 10,000 letters) for English or a few languages. For homophonic ciphers they report it against alphabet size (27 to 100 symbols). The Voynich decipherment papers validate at one benchmark size. Here the three solvers of the earlier rows are calibrated on Latin at the manuscript's token counts, 500 to 35,000 tokens and up to 247,760 letters. They are also calibrated on a wrong-genre plaintext and on a 144-symbol inventory whose frequencies copy the manuscript's finest transliteration at its own glyph count. The true key's objective is scored beside the found key's, so that each miss is attributed to the search or to the objective. That turns the question of size and inventory into a stated operating range for the verdict rows.
183Drawing-interruption edge testEstablishedNewMarcoP 2019, Lines interrupted by drawings, voynich.ninja thread 2945, started 25 September 2019…MarcoP made the same two measurements on the voynich.ninja forum: the edge-glyph distributions and the glyph association at image breaks, set against line edges and interior boundaries. Here the design differs. A positional null of random interior boundaries in the same lines, and a line-edge reference, place each statistic on a 0-1 edge index. Jensen-Shannon divergence replaces raw shares. A bias-corrected mutual information with a within-page permutation null replaces a summed deviation from independence. There are size-matched in-line and line-break references, per-hand strata (the effect belongs to hand 1, the first of the five scribal hands, with an index of 0.76 against 0.15 for hand 2), and a page bootstrap. Poured Latin and Markov-2 controls in three seeds show that the information test has power. The forum's result, that association at image breaks is as weak as at line breaks, becomes a partial retention of 15-20% against none at line breaks.
184Leftover to the drawing outlineNewsamenone foundThe literature, Zandbergen among others, states the fit qualitatively, or measures the lengths of the words beside drawings in the transcription. Here the space is measured on the scans in glyph units for 1,442 segments and compared with a simulated copyist who fills the same width with the same words. The design is that of the line-end test in R3-P2 (row 90), moved from the paragraph margin to the drawing outline, and so to hand 1. A page bootstrap and a per-bifolium table are added.
185Word-gap profile along the line from full-resolution word boxesNewsamenone foundThe published gap measurement on the scans compares classes of separator and does not look at position along the line. Whether the line is fitted by spacing or by the choice of word had been argued from transcribed word lengths. Here the physical gaps and word widths of 544 lines are profiled by position, with a within-line permutation null, per hand. They are tied to the measured leftover at the margin, which separates the two mechanisms directly.
186Physical-boundary step testAdaptedsameParisel 2026, A quantitative confirmation of the Currier language distinction, arXiv:2604.25979…Parisel types folio transitions by language and quire and compares the jumps in character-pair ratios. The precedents in manuscript studies locate a change of hand or spelling and note that it falls at a sheet or quire boundary. Here the unit is the page, so the turn of a leaf, the centre of a quire, the change of sheet and the change of quire are separated. A folio-level design cannot see the turn of a leaf. The comparison is restricted to boundaries inside one hand and one section. The statistics are the letter-distribution divergence and the vocabulary overlap of matched samples, with a label-permutation null. Seven poured and generated controls are written out into the same pages. The ink change points and the drift generator's step unit are tested against the same typing.
187Mixed-quire natural experimentAdaptedsameCurrier 1976 (as G30Currier observed that the herbal B bifolia keep their own statistics inside herbal A quires. That observation is the basis of the scribal attribution by bifolium. Parisel's published test compares the jumps between consecutive folios by transition type. Here the twelve bifolia of the mixed quires are compared pairwise, in four classes of hand by quire, on seven diagnostics and on vocabulary overlap. The contrast of hand against quire is given a permutation null that shuffles the hand labels within each quire. So the attribution is tested from the text alone, and the quire is shown not to be the unit.
188Bifolium variance of held-out surprisalAdaptedNewPelling 2022 (as G33Pelling's bifolium-level consistency is a visual map of three glyph-pair densities in one quire. The published per-page statistics (bigram vectors, character-ratio switches) are not decomposed by bifolium. Here a held-out per-token predictability is averaged per page and its variance is partitioned by hand, section and bifolium. The bifolium effect is isolated inside hand-by-section strata with a permutation of pages to bifolia. Three generators and four poured real texts in the same skeleton are scored the same way. That turns the observation into a test with a null and a control set.
189Physically-above word testAdaptedsameTimm 2014, How the Voynich manuscript was created, arXiv:1407.6639 (v3, December 2015)…Timm counted vertical similarity by line index, as the same position in a previous line or any word one to three lines up. Otherwise it has been seen by eye on a page image. Here the word physically above is found from measured word boxes and set against the same-index and same-fraction candidates on the same records. The test has a page-shuffle null, a fourth candidate from another line pair and a sign test on the discordant subset. A planted 10% physical-copy control fixes the detectable copy rate. The immediate precedent is the index- and fraction-matched test of R1-C6 (rows 28 and 41).
190Line-resolution ink runsNewsamevoynich.ninja thread 4377, Ink/pen dynamics and the rhythms of writing, September-October 2024…No source measures the ink darkness of a manuscript line by line as a series, or tests its persistence against a line-order null. None gates change points by permutation with a false-discovery rate from permuted pages, or asks whether the shifts fall on paragraph starts. The Voynich sources, such as the voynich.ninja thread on ink dynamics, describe the ink variation by eye and explain it by re-inking. The ink models of the adjacent field concern single strokes or ink identity. The nearest earlier measurement is the page-level change-point row R3-P4, with six shifts over 171 pages. The line series refines it to 130 shifts, with a breakdown by hand and a paragraph-start test.
191Gregory's-rule side checkAdaptedsameGregory 1885, Les cahiers des manuscrits grecs, Comptes rendus des seances de l'Academie des Inscriptions et…In codicology Gregory's rule is applied by reading the sides by eye. On the Voynich manuscript the one remark found is that the sides can barely be told apart. No side sequence has been recorded and no image statistic has been tried. Here the rule is turned into a test on the scans. It uses the recto-verso background difference per leaf and its sign alternation between consecutive leaves of a quire, against the 50% chance level. The ink darkness index is checked for the same asymmetry. That sorts a recto-lighter imaging bias from a parchment-side signal and clears the ink index of it.
192Layout regularityAdaptedsameDe Stefano, Fontanella, Maniaci and Scotto di Freca 2011, A method for scribe distinction in medieval…The absence of ruling in the manuscript is a qualitative observation in the grey literature. De Stefano and colleagues use interlinear spacing as one of several layout features for telling scribes apart on a ruled book. Here the spacing and the baselines are measured as regularity statistics: the pitch coefficient of variation at two resolutions, the within-line baseline slope, page skew and the pitch trend. They are compared with ruled and freehand benchmarks to ask whether the book was ruled. A permutation null within sections tests the spread between hands, and the quire effect is tested within hand-by-section strata. The measurement's own noise floor on a ruled page remains uncalibrated, because no ruled scan was on disk.
193Letter-form hand metrics from the aligned glyph cropsAdaptedsameFagin Davis 2020, How many glyphs and how many scribes? Digital paleography and the Voynich manuscript…Image-based writer identification with statistical validation is published for other manuscripts. The Voynich hands were assigned by eye, by Fagin Davis, and tested only with text n-grams. The one forum project to put numbers on the letterforms published none. Here palaeographic quantities are measured from the crops of the automatic alignment and averaged per page. They are the height of the gallows and of sh, slant, descender length, stroke width and pen lifts. They are classified by leave-one-page-out against a null that permutes the hand labels within a section. That separates the shape signal from the section signal, which dominates the text classifiers. Hands 2 and 3 remain confounded with section, and hand 5 cannot be tested.
194Census of the transcribers' comments and uncertainty codesAdaptedsameD'Imperio 1978, The Voynich Manuscript: An Elegant Enigma, NSA, section 4.2…The Voynich sources make the observation by eye: Nill, reported through D'Imperio, and Zandbergen. Pelling estimated an uncorrected-error rate on one paragraph, and the forum disputes the claim. The survey in the adjacent field counts corrections in manuscripts of a known language. Here every correction, insertion and uncertainty mark recorded by transcribers who examined each glyph is counted as a rate per 1,000 letters, by hand, section and locus type. The spread between hands is tested by permuting pages to hands within a section. The marked words are tested for line-edge position, and darkness outliers of the line series are listed for a manual pass. The result is one suspected correction per 9,000 letters in every hand, 20 to 50 times below a copyist's rate, with the marks clustered at the line edges.
195Set-off contrastAdaptedsamePelling, Voynich codicology, Cipher Mysteries (after The Curse of the Voynich, 2006)…Pelling read contact transfers in the manuscript by eye, to argue about the original order of the bifolia and the timing of the painting. One forum poster overlaid a flipped page to test a single mark. Document image analysis registers the verso's mirror image in order to remove it. Here the mirror-image transfer is measured as a contrast statistic for 106 pairs, scored against a null of random same-section masks. It is calibrated by show-through as the positive control, with a detection rate of about 40% at 1600 px. It is applied to the three pair classes that the folding and the binding predict to differ: conjugate inner faces, conjugate outer faces and the facing pages of the binding.
196Battery-wide generator searchAdaptedsameTimm and Schinner 2020, A possible generating algorithm of the Voynich manuscript, Cryptologia 44(1):1-19…Published Voynich generators, such as Timm and Schinner's, are tuned by hand or swept over a parameter grid and scored on the whole text. The one evolutionary search on the manuscript optimises a decipherment fitness. Fitting a simulator by matching summary statistics, and checking it by recovering known parameters, are methods of the adjacent field that had not been applied to Voynich generators. Here a global optimiser is run over every free parameter of four families against all 23 battery statistics together. The tables are built from training quires and the optimum is re-scored on held-out quires. A recovery control and a sensitivity scan say which parameters the battery can identify and which it cannot. It can identify the inclusion probabilities, but not the rates of change, the recency copy or the section base. The failure of every family to reach the held-out floor is therefore charged to the families and not to local fitting. The held-out competition of row G6 refitted generators by coordinate descent. This row asks whether a global search changes that answer, and it does not.
197Shared-escape prequential class contestAdaptedsameDawid 1984, Present position and potential developments: some personal views: statistical theory: the…Prequential codelength comparison, two-part codes, class n-grams and cache models are standard in the adjacent field. On the Voynich text, the published work induces word classes and measures bigram predictability, and the grey neural experiments score n-gram reproduction or a single model's loss. Here the three accounts of the text (a language with word classes, a self-citing procedure, and a word codebook into a known language) are each given their strongest likelihood-bearing member. All run under one shared spelling model on the same page folds, so that only the sequence prediction differs. An oracle-table codebook control fixes the cost of the codebook class even when the table is known. A class-detection control is run on each account's own kind of text. It shows that the class instrument fails its own positive control at this text size, and that the codebook loss says nothing about hidden content.
198Gain anatomy of the character LSTMAdaptedsameKhandelwal, He, Qi and Jurafsky 2018, Sharp nearby, fuzzy far away: how neural language models use context…The interpretation tools are published for natural-language models by Khandelwal and colleagues, Hupkes and colleagues and others. They are context ablation by shuffling, retraining on permuted text, and diagnostic classifiers on hidden states. The only Voynich neural interpretation found is a grey saliency map on a small GPT. Here the quantity taken apart is the recurrent model's advantage over the best n-gram. It is split by position in the word and by word length, against four controls on their own models. The real-trained model is cross-scored on three generator outputs and a generator-trained model on the real text, with the per-line difference as a discriminator. Retraining on a word-order shuffle measures how much of the gain needs the order of the words. A word-identity model tests whether that order is held in word identities. Probes with majority and shuffled-label baselines complete the set. Together they locate the gain in the interior glyphs of the current word. It is drawn from the glyphs of the preceding words and from the page's regime, and not from word identities or long words.
199Planted-signal power ladderEstablishedAdaptedCohen 1988, Statistical Power Analysis for the Behavioral Sciences, 2nd ed., Lawrence Erlbaum, reissued…The grey 2026 repositories plant a signal at two or three sizes into one of their own tests (page-level anchors, planted numerals, attribute-label bindings) and report whether it is recovered. Here the same design is run as a ladder of six sizes with ten replicates, over fourteen null results already published in the earlier rows. Each has a planting recipe that copies the text's own dependence law, so that the planted size eps is measured in units of the in-line or adjacent level. The tests keep their own permutation nulls, with matched nulls for the mutual-information and glyph-model tests. An interpolated minimum detectable effect at 80 percent power is given for every null. That turns each null verdict into a bound. Cross-line dependence is ruled out above 0.2 of the in-line level, hierarchical glyph structure above 2 percent of tokens, and a wrong key entry at 1 of 23. It also identifies three nulls with no useful power (MIx3, Q2 and Q18). Against the textbook form of Cohen, power is estimated by planting into the real text with test-specific recipes, and not by an effect-size convention or a parametric simulation.
200Half-sample calibration against seed spreadAdaptedsamePolitis and Romano 1994, Large sample confidence regions based on subsamples under minimal assumptions…Politis and Romano validated subsampling intervals by theory, and the literature otherwise validates them by Monte Carlo coverage on a known process. Here each generator is treated as the known process. The text-side half-sample estimator is applied to single realisations and divided by the generator's seed-to-seed standard deviation. That gives a calibration factor per statistic and per generator: 1.05 for a stateless page process, and 0.44 for a process with state that spans pages. A quire block bootstrap is added as an upper scale. The battery's within-tolerance counts are recomputed on five scales, so that the published rule can be labelled conservative or liberal for each generator. The grey Voynich battery of 2026 reports a folio bootstrap and generator seeds side by side without comparing them, and the 2026 preprint reports seed ranges only.
201Leave-one-seed-out Mahalanobis joint testEstablishedAdaptedMahalanobis 1936, On the generalised distance in statistics, Proceedings of the National Institute of…Three grey 2026 repositories place the real text in a Mahalanobis region of generator seeds with a regularised covariance, which is the same measurement. Here the covariance is shrunk by the Ledoit-Wolf estimator over 23 statistics and 50 seeds. The reference distribution is the seeds' own leave-one-out distances, in place of a chi-square or an empirical 95 percent region. The calibration is checked with held-out seeds of the same generator (one rejection in ten) and seeds of another generator (five of five). The test is run for five generators and repeated on a second transliteration. The design matches the leave-one-out goodness-of-fit statistic of the approximate Bayesian computation literature more closely than the textbook Mahalanobis test.
202Energy-distance separation of generator cloudsNewsameSzekely and Rizzo 2004, Testing for equal distributions in high dimension, InterStat, November (5), cited…The energy test is applied in its textbook form, as a two-sample E-statistic with a permutation reference, without change. What is specific is the object: clouds of 23-statistic battery vectors produced by generator seeds and by quire resampling of the text. There are three uses: seed-cloud homogeneity, pairwise generator separation, and separation of the text from each generator. No earlier Voynich use of any multivariate two-sample test between generator output clouds was found. The classification is therefore NEW as a standard method applied to this text for the first time, and not as a new statistic.
203Max-T over half-sample vectorsEstablishedNewWestfall and Young 1993, Resampling-Based Multiple Testing: Examples and Methods for p-Value Adjustment…Two grey 2026 repositories apply a Westfall-Young max-statistic permutation to their own families of tests on the Voynich text. Under the conservative rule that makes the method's use on this text established. Here the family is the 23-statistic generator battery. The resampling unit is the page half-sample of the real text, in place of a permutation of labels. The output is a family-wise critical |z| of 3.16 against 1.96, applied to every generator's misfit count on two variance scales. The count is re-derived on a second transliteration. The design is the textbook max-T with the battery's own resampling scheme substituted for the permutation scheme.
204Multiplicity ledger of harvested p-valuesEstablishedNewHolm 1979, A simple sequentially rejective multiple test procedure, Scandinavian Journal of Statistics…Several grey 2026 projects apply Bonferroni or Benjamini-Hochberg over their own search space. One of them, the automated exploration of April 2026, corrects over its whole ledger of findings, which is the same measurement in kind. Here the ledger is built mechanically from every saved p-value of a 153-technique examination, and not from nominated findings. The correction is run three ways: Holm within family, Holm across all, and Benjamini-Hochberg. Every demoted effect is named against its threshold. The floor set by permutation resolution, which makes Holm across all unattainable, is stated. The techniques are counted per verdict row, so that each row's evidence base is sized. Against the textbook procedures nothing is changed.
205Held-out transliteration protocolAdaptedsameNosek, Ebersole, DeHaven and Mellor 2018, The preregistration revolution, PNAS 115(11):2600-2606…That single statistics hold across transliterations is established: entropy across transcription systems in the peer-reviewed literature, and a locked ten-metric equivalence panel across editions in the grey 2026 battery. Preregistration, as Nosek and colleagues describe it, is a published practice for new data. Here the two are combined, so that a second transliteration of the same manuscript is the confirmatory sample for a generator-comparison battery. The decision rules (per-statistic tolerance, family-wise misfit count, joint leave-one-out p, and the distinction between fitted and predicted statistics) and the sampling scales are written down before the run. The verdict is then re-derived, and not only the statistics compared. The object confirmed is the ledger of generator verdicts, and not the values of the statistics.
206Bias-corrected entropy and complexity profile with quire jackknifeAdaptedsameBennett 1976, Scientific and Engineering Problem-Solving with the Computer, Prentice-Hall (record atThe conditional glyph entropies are established, by Bennett (1976), Stallings (1998), Lindemann and Bowern (2021) and voynich.nu. Gaskell and Bowern (2022) have a compression metric on excerpts, and grey sources report a Miller-Madow-corrected h1 and a bare LZ complexity. The published Voynich figures use plug-in counts at one length without bias treatment, and the sample-size question raised in the forum literature is answered nowhere. Here the same quantities are placed in one profile, all on the same 30 language samples and generator seeds. The profile has four estimators and their disagreement as a function of block length, an excess-entropy curve, a normalised LZ76 ratio and a match-length rate. A delete-one-quire jackknife gives a bias estimate as well as a spread, and a page half-sample interval is added. An estimator self-test at the text's length fixes the usable block length (L at most 4 or 5) and the match-length bias. The finding is that the text has the lowest E(4) of all the languages, and that it is more compressible and more predictable than its own drift generator.
207Skeleton modelAdaptedsameTimm and Schinner 2020, A possible generating algorithm of the Voynich manuscript, Cryptologia 44(1):1-19…Published and grey generators handle line lengths in one of three ways. Timm and Schinner fix them by a character limit. Rozanova and Temerev, and the borrowed skeleton of the earlier rows, copy the manuscript's own line template. The 2026 automated exploration draws words per line independently from the empirical histogram. The model here is hierarchical (section, page width class, paragraph, line class, and the previous line's length) and fitted. It is checked against the real distributions with KS distances and the lag-1 correlation of consecutive line lengths. It is used for a controlled comparison of the same generator seeds in the borrowed and the fitted skeleton. That measures how much of a generator's battery result the borrowed layout is responsible for: two statistics of twenty-three.
208Line-length regression with a within-page permutation nullAdaptedNewVogt 2012, The line as a functional unit in the Voynich manuscript, self-published PDF, 27 November 2012…The coupling between words per line and word length is a grey measurement. Vogt (2012) inspected the distributions, and the 2026 automated exploration used Pearson correlations per position class against a simulation. Word-length selection near drawings and line edges is measured in a preprint and a workshop paper. Under the conservative rule the fit-to-space component alone would be established. Here the statistic is a held-out R squared decomposition over three nested term groups: mean word length, paragraph position class, and a vocabulary term that cannot encode the count. It is run on the real text against two controls poured into the same skeleton, each line given the real line's number of words. Such a pour leaves a control with no dependence by construction, so the two controls set the floor of the statistic and cannot fail. The vocabulary gain, which asks whether word identities beyond their lengths predict the line's width, is tested against a within-page permutation null in two seeds. No precedent was found for the vocabulary test or for the poured-control design.
209Paragraph topic gain against the drift generator with a codebook controlAdaptedsameSterneck, Polish and Bowern 2021, Topic modeling in the Voynich manuscript, arXiv:2107.02858, and the Yale…Sterneck and colleagues (2021) ran LDA on the Voynich text, and Reddy and Knight (2011) measured page topicality. Montemurro and Zanette (2013) measured section keyword structure, and the grey 2026 multiscale audit measured held-out cache gains at several scopes. An earlier row here measured within-page cache and cross-page LDA gains. Here the document-completion evaluation of the topic-model literature is moved to the paragraph half. The statistic is the topic gain beyond what the paragraph's own words explain (LDA plus cache, minus cache). The null is a stateful generator without topics, run in the same strata with five seeds, in place of a shuffle. A codebook over English text is planted as the positive control, and it is shown to be detectable only through far recurrence of page-dominant topics. The result is a stated power limit: topics of English strength in the Currier B strata, and nothing weaker.
210Verdict statistics re-measured inside each Currier regime against samples of the same lengthAdaptedsameCurrier 1976, Some important new statistical findings, seminar paper, consulted through the summaries…Earlier work measured the two regimes to show that they differ. Here each regime is scored on the full verdict battery against reference samples cut to its own length and against imitations of its own glyph statistics, to test whether any verdict drawn from the whole text depends on pooling A and B. None reverses, and the repeated-phrase argument is found to need the whole text because the A pages alone are too short for it.
211A hand procedure of the period, simulated and scored on the full batteryAdaptedsameRugg 2004, An elegant hoax? A possible solution to the Voynich manuscript, Cryptologia 28(1)…Rugg's table and grille and Timm and Schinner's self-citation are hand methods, and each was shown to reproduce some properties of the text. Here a procedure restricted to devices in use before 1400 (tables, lots or dice, a ruler and the page, with no Cardan grille and no cipher disk) is simulated with every random act counted, its rates fitted, and its output scored on the 23-statistic battery, the extras and the recurrence profile beside the published generators on the same footing. Ablations tie each device to the statistics it buys, two table sets (one set of tables for each writing style) test the Currier regimes, and the cost per word is stated. The cost accounting, the per-device ablation on a common battery and the regime test have no precedent found for this text. The search that followed has none either: some forty families of smaller procedure, the kinds of smaller procedure only, on four independent lines under a bound on the apparatus, the count of cells (table entries) and sides beside every score, held-out quires with the tuned splits named, a page classifier, and the derivation of every word from what was in sight with description length as the judge. Greshko's Naibbe cipher (2025) is the one published key of the period's kind whose sign variants are chosen by chance, and the dice-read tables here are its relatives without a plaintext.

The techniques that are new under both views.

The techniques that become new once the unreviewed work is set aside. These twenty-three have a preprint, a forum thread or a student project as their nearest precedent. The peer-reviewed literature has no precedent for them.

The adapted techniques that matter most. Two of them, the verbose solver and the key-transfer test, count as new once the preprints are set aside.

What goes beyond the 2025 and 2026 preprints

Four preprints posted between September 2025 and August 2026 cover part of the same ground. Parisel wrote three of them (arXiv 2509.10573, 2604.19762 and 2604.25979) and Rozanova and Temerev wrote the fourth (arXiv 2608.17096). None has been peer reviewed. This examination therefore treats them as unreviewed claims. It does not treat them as a benchmark. The first count above includes them as prior work and the second sets them aside. Everything said about the preprints here comes from their own texts, and this examination did not run anything taken from them. The table lists, area by area, what they report doing and what this examination does that they did not.

AreaWhat the preprints did, by their own accountWhat this examination adds
Word-order informationRozanova and Temerev measured the information between adjacent words as the excess over the mean of 100 shuffles, each of which moves every word of a line, with the vocabulary capped at 2,000 types. They found 0.066 bits. Parisel measured the glyph-level information across the word boundary.This examination corrects the estimate for bias and compares it with a shuffle that keeps the first and last word of every line in place, so that line-edge effects do not count as word order. That gives 0.113 bits. Splitting the same information by word slot places it at the boundary between words. This examination also measures a held-out bigram gain with an open vocabulary and the character information at distances of 1 to 20 glyphs. It re-measures the preprints' own boundary signatures and merge-level dependence gap with a Markov imitation and a within-word shuffle as controls, which the preprints lack. It measures the information across line and paragraph breaks as well.
Substitution-cipher attacksRozanova and Temerev ran one calibrated attack. It merged 88 units into 23 classes, used a near one-to-one mapping and seven trigram language models, and scored held-out lines. Parisel ran none.This examination ran three kinds of solver: one-to-one, many-to-one with a description-length cost, and verbose with nulls, which the Rozanova and Temerev paper puts out of scope. Each solver ran over 21 language models of order 4 and order 5, which predict each letter from the last few letters, and over three symbol views. Each was validated on ciphers with a known key. Every run had a meaningless control. This examination also ran the key-transfer test between Currier A and B, measured dictionary hit rates against the decoded control, and added 19 further corpora with known-key checks in Hebrew, Arabic and Sanskrit.
TranspositionRozanova and Temerev put transposition out of scope. Parisel tested the reading direction only.This examination ran six tests. They include a solver that anagramming cannot defeat, checked on known keys in five languages, nine re-readings of the glyph stream by route and by column, and a count of repeated letter-bag sequences, which anagramming cannot hide. It also re-ran the reading-direction test with a convention that treats both directions alike, and under that convention the published asymmetry vanishes.
Codes and numeralsRozanova and Temerev used five glyph-level homophonic set-ups, and took the Naibbe cipher, a Linnaeus catalogue and Chinese in pinyin as comparison texts. Parisel did nothing in this area.This examination built a word-level codebook grid of 620 texts. It ran the words-as-numerals slot test, a leading-unit test against Benford's law and the ordered-table statistics. It also tested syllabic and letter-pair codes of real languages.
Hidden channelsNeither preprint.This examination tested 29 derived channels for a message, with planted messages as positive controls.
Text generatorsRozanova and Temerev ran the Timm and Schinner copying generator and an order-3 glyph generator, on a few statistics. Parisel built a slot-based generator with 12 ablations and a Cardan-grille generator.This examination scored 24 generators from six families on a fixed battery of 23 statistics, with a ledger of what each generator was fitted to. It added a table-and-grille generator over 168 configurations and a wheel generator. It counts the repeated strings without word segmentation, which no generator is fitted to. It runs the first-glyph test, which separates one kind of nested category code from a left-to-right template. It scores a recurrent model against the n-grams on the same folds.
Line and page layoutRozanova and Temerev measured the divergence of line starts with a relabelling null, stratified by Currier language, and measured the enrichment of gallows. Parisel measured positional polarisation and a boundary anomaly.This examination pours reference texts into the actual skeleton of pages, paragraphs and lines and uses that as the null. It measures the line-edge derivation rates, and the recurrence of words by physical distance with a null stratified by hand and Currier language. It runs a permutation test of the section vocabularies. It tests the claim of paired gallows on top lines against a placement model, and it measures the rightward and downward drift of minimal pairs with nulls.
Transcription reliabilityRozanova and Temerev used a second reader, measured agreement on separators, validated the uncertain spaces from glyph coordinates and ran a blind ink audit. Parisel ran cross-transcription checks.This examination aligns the Zandbergen–Landini and Takahashi files at the level of text unit, word and character, with a dispute rate for every symbol. It recomputes every headline statistic under three symbol views and two word-boundary conventions. It crops the disputed glyphs from the scans. It measures the widths of the uncertain spaces on the scans against blind predictions, and it recovers the spaces from the space-free glyph stream with an unsupervised segmenter.
Currier A and BParisel built a classifier that predicts held-out folios at 89 percent and found near-categorical switching of four letter pairs. Rozanova and Temerev only used A and B to stratify their measurements.This examination runs the key-transfer test, which asks whether one key explains both dialects. It classifies pages with one page left out at a time. It measures the differences between A and B against half-splits of real texts. It attributes scribes by character n-grams and by Burrows' Delta and compares the result with the palaeographic hands. It fits a two-source mixture model with single-language and bilingual references in the same page skeleton.
Meaning, pictures, labels, embeddingsNeither preprint.This examination runs a picture-to-word test on 119 herbal pages with a planted control, a function-word test, and label recurrence and numbering tests. It fits word embeddings and aligns them across languages, with a power check on known language pairs. It builds the paragraph-similarity network, the keyword ranking and the information-over-scales curves, with 28 languages and nine generators poured into the same skeleton. It adds gibberish, glossolalia, ciphertext and symbol systems as reference classes.
VerificationRozanova and Temerev resampled by quire, left one quire out at a time and compared disjoint halves.Separately written code re-derives every headline number. The repetition counts are re-run over 140 contiguous samples, and the repeated 4-grams are recounted after the transcription variants are merged. Every headline statistic gets a leave-one-quire-out interval. A demonstration shows that a quire bootstrap with replacement corrupts the repetition and information statistics.

The preprints also did some things that this examination did not, so this section is not a superset of them. Rozanova and Temerev ran a blind ink audit and used a catalogue and a pinyin text as comparison texts. Parisel measured the switching of four letter pairs between Currier A and B and built a Cardan-grille generator in Parisel's own form. A Cardan grille is a card with holes in it, laid over a table so that the holes pick out the parts of a word. The table-and-grille generator tested here is Rugg's. Four of the preprints' measurements were re-run here with controls: the validation of uncertain spaces from glyph coordinates, the reading-direction test, the byte-pair dependence gap and the resampling by quire. The section on the data and the sections each measurement bears on say what came of it.

The findings, old and new. About thirty of the results in this report reproduce published ones. Among them is the low glyph predictability, which is the low entropy of a glyph given the one before it. So is its cause, the fixed order of glyphs inside a word. The Zipf-like vocabulary, with a few common words and a long tail of rare ones, was known, and so were the narrow spread of word lengths and Sukhotin's vowel classes. The line was known to be a functional unit, with words behaving differently at its two ends. The absence of repeated phrases was known, and so were the networks of near-copy words. Currier A and B were known to follow the scribal hand. Earlier work had ruled out simple substitution, in which each sign stands for one letter, and had found that verbose encodings, in which one letter becomes several signs, survive. It was known that meaningless generators reproduce most of the single statistics. Rare words were known to come in bursts within a page, and each section was known to have its own vocabulary. The modularity of the paragraph-similarity network, the two-state letter grammar and the long-range correlations between words had all been reported. So had the low glyph-level rigidity on the symbol-system scorecard and the width of the uncertain spaces on the scans.

The search did not find the following results anywhere else. With the spaces removed, the text has no repeated string of 30 glyphs. Once the adjacent glyph is known, the first glyph of a word tells nothing about its later glyphs, which counts against a nested category code. On word-order information measured against the edge-fixed null, a shuffle that leaves the first and last word of each line in place, the real text separates from every meaningless generator. A cipher key fitted on Currier A loses information when it is scored on Currier B, where a Latin cipher loses nothing. Rank-matched codebook texts fail on the statistics of similarity between neighbouring words. The word-order information lies at the boundary between words. The line-edge derivation rates were quantified.

The search did not find these results elsewhere either. One slowly drifting state reproduces the whole multi-scale clustering of the vocabulary. The one-edit variants of words show no genealogy along the codex, that is, no line of descent as the pages go on. The line ends are fitted to the margin. The held-out predictive model leaves 11 bits per word unpredicted and gains 0.02 bits from word order. The Naibbe cipher separates from the text on vocabulary and page structure. Its agreement with the text on entropy and repetition is in the preprint by Rozanova and Temerev, whose Naibbe numbers agree with the ones found here. The argument from missing repeats holds when spelling variation is allowed for. No information crosses a line break. The claim of paired gallows on the top lines of paragraphs fails. The recurrent model gains over the n-gram models, and the gain lies within the word and the few words before it. The calibrated closeness score picks out no candidate language.

The later comparisons add more results that the search did not find elsewhere. The unicity distance of each cipher family, which is the length of ciphertext needed to pin down its key, was set against the length of the text. The Book of Soyga, the Lingua Ignota and the Steganographia were scored beside the manuscript as texts of known status. The first-glyph test does not separate a constructed vocabulary from the text. The generators competed on held-out quire halves, and a page classifier tells every generator from the text. The entropy margin below the nearest languages was measured at a matched sign inventory. The list genres of the period were compared at matched size. There are no isomorphic repeats (two strings with the same pattern of repeated signs) at 25 and 30 glyphs under any relettering. The lines of the first hand fit the drawings by word choice, and the line rule was tested again where the pen restarts after a drawing. The bifolium, a sheet folded once to give four pages, is a unit of the writing on two tests. Within one hand and section the letter distribution steps more at a change of sheet than at the turn of a leaf, and none of seven poured and generated controls in the same pages shows that step. And a bifolium of one hand bound in a quire of another keeps its own hand's statistics, not its quire's. Every generator is rejected jointly, with a calibrated tolerance. The minimum detectable effect of each null test was measured.

Five points differ from published readings. The size of the word-order signal depends on the null model. Rozanova and Temerev find 0.066 bits with a whole-line shuffle, and this examination finds 0.113 bits with the edge words held fixed. The signature of copying is diffuse here, where Timm and Schinner's generator copies locally. Landini and Montemurro read the long-range glyph correlations as linguistic organisation. This examination attributes them to the word template plus repetition. Parisel's reading-direction asymmetry does not survive a convention that treats the two directions alike. Arutyunov and colleagues inferred a mixture of two languages from the letter frequencies. This examination reproduces that gain, but the gain persists inside each scribal hand. So it reads as variation from page to page. It does not read as two languages.