Compared with published work
What is new here, and what was known before? A technique is one method of measuring or testing the text, such as one of the 23 measurements. We used 211 techniques in all. They fall in three groups. The first group is 77 techniques that had been applied to this text before. The next is 102 published methods used here for the first time. The last is 32 that are new. Every technique was also run on a text of known origin. About thirty of the findings had already been published by others. The page also says which published claims do not hold up on these tests.
This examination finished its measurements before it checked the literature, so no earlier theory could steer them. It then listed every technique it had used, 211 in all, and searched for an earlier application of each one to the Voynich text. The search covered journals, the 2022 Voynich conference proceedings, arXiv preprints up to August 2026, René Zandbergen's site, and the main blogs and forums. Each technique was then sorted under deliberately conservative rules, which lean towards calling a technique established. A technique counts as established if someone had applied the same measurement to this text before, whatever controls were used here. It counts as adapted if this examination changed a published method in a named way. A named change is one of three things: a new null model, a new invariant or a validation step. A null model shows what the measurement gives by chance. An invariant is a quantity that some kind of encoding cannot change. A validation step runs the method on a text with a known answer. A technique counts as new only if the search found no earlier application at all.
The count depends on what is allowed to count as prior work, so the report gives it two ways. The first view counts every source the search found, including four unreviewed preprints from 2025 and 2026, student projects, code repositories and forum threads. Under that view, 32 techniques are new, 102 adapted and 77 established. The second view counts only peer-reviewed journals and the long-established sources, which are Currier, Tiltman, D'Imperio, Bennett, Stolfi and Zandbergen's site. Under that view, 55 are new, 103 adapted and 53 established. The two views disagree on 36 techniques, and each of those has a preprint, a student project or a forum thread as its only precedent. "New" means that the search found no earlier application. It does not mean that nobody has tried the technique.
Counting every source, preprints and unpublished work included
Counting peer-reviewed and long-established sources only
The full table of techniques, with the nearest prior work under each view.
In the table, h1, h2 and h3 are the entropy of a glyph, which is the number of bits of surprise in it, given none, one or two preceding glyphs. MI is mutual information, which is how many bits one quantity tells about another, and PMI is pointwise mutual information, the same measure for one particular pair of values. MDL is minimum description length, BPE is byte-pair encoding and kNN is k nearest neighbours. CI is confidence interval, LM is language model and KN is Kneser–Ney smoothing. PPMI-SVD is a word-embedding method, and LDA, LSA and NMF are topic models. IIIF is the image service of the Beinecke scans and IVTFF is the transcription file format. F/T/P means fitted, trained or predicted. ZL, IT and GC are the Zandbergen–Landini, Takahashi and Claston transcriptions.
| # | Technique | Status, all sources | Status, reviewed sources only | Nearest prior work | What differs |
|---|---|---|---|---|---|
| 1 | Conditional character entropies h1/h2/h3 of the glyph stream vs matched-size language samples | Established | same | Bennett 1976, Scientific and Engineering Problem-Solving with the Computer (Prentice-Hall) | The measure is Bennett's. Added here are exact size matching, with 20 windows per corpus, and three conventions for the space. Neither changes the known result of 2.0-2.2 bits against 3.0-3.6 for the languages. |
| 2 | Sensitivity of entropy to transliteration and alphabet convention | Established | same | Lindemann & Bowern 2021, arXiv:2010.14697, (partition by transcription system) | No material difference |
| 3 | Within-word letter-shuffle control attributing the h2 deficit to glyph order inside words | Adapted | same | Lindemann & Bowern 2021, arXiv:2010.14697… | Lindemann and Bowern reached the same conclusion by studying where each glyph can stand inside a word. Here the letters are shuffled inside each word and the entropy is measured again. Rozanova and Temerev shuffle whole words within the line, which is a different control. |
| 4 | Character mutual information at distance d=1..20 with a letter-shuffle floor (the value the shuffled letters give) and a word-shuffle decomposition | Adapted | same | Landini 2001, Cryptologia 25(4), (spectral analysis of the text without spaces | Landini, Schinner, Amancio and colleagues, and Montemurro and Zanette measured long-range correlation in other ways (spectral analysis, random walks, word-level information). The curve here is character-level mutual information at each distance, read against a letter-shuffle floor, and it is split into a word-template part and a word-order part. |
| 5 | Sukhotin vowel/consonant separation with alternation-ratio diagnostic | Established | same | Guy 1991, Cryptologia 15(3) 207-218 and 258-262 (two folios | Guy, and later Reddy and Knight, ran the same algorithm. The alternation ratio and the comparison with languages are small additions. |
| 6 | Word-length distribution shape | Established | same | Stolfi 1997-2000, Voynich pages (binomial word-length distribution) | No material difference |
| 7 | Lexical statistics at matched size | Established | same | Landini 2001, Cryptologia 25(4), (Zipf) | No material difference beyond exact size matching |
| 8 | Adjacent-word mutual information in excess of a word-shuffled baseline | Established | Adapted | Rozanova & Temerev 2026, arXiv:2608.17096, (adjacent-token MI minus the mean of 100 within-line shuffles | Rozanova and Temerev, and Parisel, had the same idea and found the same qualitative result. Their shuffle moves words within each line. The shuffle here moves words across the whole window. |
| 9 | Immediate repeats and edit-distance-1 near-repeats of adjacent words vs shuffled baseline and vs languages | Established | same | Tiltman 1967 / Currier 1976 as summarized at voynich.nu 'Sentences etc', (doubled and tripled words) | Tiltman, Currier and Timm counted the doubled words. Two things are added here: the ratio to a shuffled baseline, and the finding that prose languages fall below their own chance level. The technique is not new. |
| 10 | Currier A against B (the two writing styles Prescott Currier identified), measured against half-splits of reference corpora and against halves of B alone | Adapted | same | Currier 1976, reproduced and revisited at voynich.nu | The difference between A and B has been known since Currier. New here is the yardstick: disjoint halves of 16 reference corpora show how large a difference one text can produce on its own. |
| 11 | Effective alphabet size | Established | same | Zandbergen, voynich.nu character statistics | No material difference |
| 12 | Entropy-invariance argument | Established | same | Zandbergen, voynich.nu 'What we learn from entropy', (substitution leaves entropy unchanged | No material difference. The numbers here reproduce the known behaviour. |
| 13 | Positional statistics of glyphs | Established | same | Tiltman 1967, NSA Technical Journal (reprinted), no stable URL | No material difference |
| 14 | Adjacency and co-occurrence constraints | Established | same | Tiltman 1967 (as above) | The constraints are classic observations, from Tiltman and Stolfi onward. The shuffle nulls, which say how often each constraint would be broken by chance, are an addition. |
| 15 | Unsupervised discovery of multi-glyph units by PMI merging and MDL | Established | Adapted | Rozanova & Temerev 2026, arXiv:2608.17096… | Rozanova and Temerev learn the units by byte-pair encoding. Here they are learned by merging on pointwise mutual information and by minimum-description-length segmentation. The purpose is the same and the same units come out. |
| 16 | Branching entropy inside words | Established | same | Zandbergen, voynich.nu 'From bigram entropy to word entropy'… | Zandbergen measured the profile from the start of the word. The profile from the end of the word, read right to left, and the successor-variety counts are additions. |
| 17 | Ordered-slot grammar learned by greedy set cover | Adapted | same | Stolfi 2000, (crust-mantle-core grammar) | Stolfi and Zattera built their slot grammars by hand. The grammar here is learned automatically. It is scored by a curve of rule budget against coverage, by the productivity gap and by the shuffled-word selectivity test, and the same scoring is applied to Latin, Italian and Hebrew. |
| 18 | Edit-distance-1 word families | Established | same | Timm 2014, arXiv:1407.6639 | Timm and Schinner described the families. The language baselines at matched size (Latin, Italian, Hebrew) are an addition. |
| 19 | Hapax legomena as concatenations of two attested shorter words | Established | same | Timm & Schinner 2020, Cryptologia 44(1)… | Timm and Schinner made the observation. Here the share is quantified and set against Latin and Italian at matched size. |
| 20 | Sequential vs hierarchical within-word prediction test | New | same | none found | No earlier test asks whether the first unit of a word constrains the later units beyond what the adjacent unit tells. That constraint is the signature of a nested category code. A left-to-right slot process does not have it. |
| 21 | Frequency-band feature analysis | Established | same | Currier 1976 | No material difference |
| 22 | Cross-transcription and convention sensitivity of word-structure measures | Established | same | Lindemann & Bowern 2021, arXiv:2010.14697 (by transcription system, hand, Currier language) | No material difference |
| 23 | Line-initial and line-final letter and word distributions vs interior words | Established | same | Currier 1976 ('line as a functional entity') | No material difference |
| 24 | Reference texts poured into the Voynich page/paragraph/line skeleton as the null for layout effects | Established | Adapted | Rozanova & Temerev 2026, arXiv:2608.17096… | Rozanova and Temerev wrap prose controls to the line template. The verse and wrapped-line references, which have line units of their own, are a small addition. |
| 25 | Paragraph-initial and paragraph-final effects | Established | same | Currier 1976 and D'Imperio 1978 (split gallows on first lines), summarized at | No material difference |
| 26 | Word length by position in the line | Established | same | Timm 2014, arXiv:1407.6639, (first word longer, second shorter, decline along the line) | Timm measured word length by position. The curve of how the edge effect decays along the line is a small addition. |
| 27 | Line-edge derivation test | Adapted | same | Smith 2015, 'Linestart words', (initial o stripped at line start | Smith proposed the stripping and adding transformation from glyph frequencies. Here it is tested by counting how often the stripped forms are attested words, against a matched interior baseline. The test of replacing final -m is added. |
| 28 | Neighbour similarity | Adapted | same | Timm 2014, arXiv:1407.6639, (similarity to words on the same line and the line above) | Timm measured similarity to the words on the same line and on the line above. Added here are position-matched controls, which remove the confound of the line-initial word, and the flat profile over lags. |
| 29 | Adjacent-repeat and near-repeat rates against a within-page permutation null, and runs of >=3 near-identical words | Established | same | Timm 2014, arXiv:1407.6639 | Timm and Schinner counted the chains of similar words. The permutation null is an addition. |
| 30 | Page-vocabulary Jaccard overlap by pair category | Adapted | same | Zandbergen, 'The Currier languages revisited', (page-by-page distance matrix from bigram/word distributions | Zandbergen, Montemurro and Zanette, and Sterneck, Polish and Bowern clustered pages by their text statistics. Here the overlap is measured as an excess over a word shuffle, by pair category and by physical distance, and a narrative poured into the same skeleton is the reference. |
| 31 | Word burstiness | Adapted | same | Montemurro & Zanette 2013, PLoS ONE 8(6)… | Montemurro and Zanette measured the same property with a different statistic. Here it is document frequency and recurrence against an exact expectation, by frequency band. The conclusion agrees. |
| 32 | Currier A/B page classification | Established | Adapted | Parisel 2026, arXiv:2604.25979, (supervised classifier, a program that learns to sort pages, predicting A/B of held-out folios, the folios kept out of its training, 89.2%) | Parisel classified held-out folios from character statistics, so the technique exists. The feature set and the cross-validation by hand against section differ in detail. |
| 33 | Label statistics | Established | same | Zandbergen, voynich.nu (labels in the writing-system and analysis pages), and | No material difference |
| 34 | Numbering-system test for labels | Adapted | New | voynich.ninja thread 'Zodiac labels'… | Forum posts on voynich.ninja have compared adjacent zodiac labels, but they are unreviewed. The test here compares consecutive labels with random pairs from the same page, as a check for a numbering system, on the star and the zodiac labels. It was not found in that form. |
| 35 | Repeated word n-grams (runs of n words) | Established | same | Zandbergen, ('curiously lacks repeated phrases of 2 or more words') | Zandbergen and Timm noted the missing phrases. The poured references at matched size and the shuffled controls are additions. |
| 36 | ZL vs IT agreement at locus, word and character level by sequence alignment, with per-symbol dispute rates, confusion pairs, and breakdown by section/hand/language | Adapted | same | Zandbergen 2022, CEUR Vol-3313 keynote 'Transliteration of the Voynich MS text'… | Landini and Stolfi aligned the files of several transcribers long ago. New here are the dispute rate for every symbol and the ranking of confusion pairs between the ZL and IT transcriptions, by section and by hand. The agreement percentages in the earlier sources were not verified. |
| 37 | The v101 alphabet as a third symbol view | Adapted | same | Zandbergen, voynich.nu transliteration pages and IVTFF tools… | Zandbergen's conversion tables were written by hand. The table here is estimated from the paired words, and it measures purity, which is how deterministic the correspondence is in practice. |
| 38 | Word-space reliability between readers | Established | same | Zandbergen, IVTFF format (uncertain space convention) | Rozanova and Temerev's 2026 preprint validates the uncertain spaces from glyph coordinates. That is more thorough than the scan sample used here. |
| 39 | Targeted inspection of disputed glyphs on full-resolution IIIF crops | Established | same | standard practice of every transcriber (Takahashi, Zandbergen, Claston) | No material difference |
| 40 | 24 generative models from six families scored against a fixed 23-statistic battery, with fitted/trained/predicted bookkeeping per statistic, page-half-sample CIs and z-scores | Adapted | same | Timm & Schinner 2020, Cryptologia 44(1), (generator, a program that writes text by a rule, vs a set of statistics) | Timm and Schinner, Gaskell and Bowern, Parisel, and Rozanova and Temerev all compared a generator with a set of statistics. Here 24 models are scored on one fixed battery, and a ledger records for each statistic whether the model was fitted to it, trained on it or predicts it. Only genuine predictions count. |
| 41 | Autocopy generators | Established | same | Timm & Schinner 2020, Cryptologia 44(1) | Timm and Schinner fix the copying kernels. Here they are fitted to the text. The family of generators is theirs. |
| 42 | Slot-template generators | Established | same | Rugg 2004, Cryptologia 28(1) (table-and-grille generation) | No material difference in kind |
| 43 | Word-level Markov chains | Established | Adapted | Parisel 2026, arXiv:2604.19762… | No material difference |
| 44 | Hierarchical category-code generator | New | same | none found | No earlier test was found that builds a Wilkins-style nested category code as a text generator and scores it against the text's statistics. Bowern and Lindemann discuss Friedman's philosophical-language idea, but no generator had been built from it. |
| 45 | Codebook/nomenclator generators | New | same | none found | No earlier work scored a rank-matched codebook text against the Voynich statistics. Kuykendall's annealer has a nomenclator mode, but as a solver, and Greshko's Naibbe cipher is a substitution cipher. Neither is a codebook generator. |
| 46 | Bias-corrected adjacent-word MI with an interior-only shuffle | Adapted | same | Rozanova & Temerev 2026, arXiv:2608.17096… | Rozanova and Temerev's null shuffles every word in the line. The null here keeps the first and last word of each line in place, because edge words are different kinds of words (once-only words are twice as common among them). With a whole-line shuffle the correction flips sign (-0.12 bits) on this data. The edge-fixed version is what separates the real text from every copy and template generator. Whether the whole-line version does the same is not reported in the preprint. |
| 47 | Simulated-annealing monoalphabetic solver | Established | same | Hauer & Kondrak 2016, TACL 4, (LM-based decipherment of the VMS as substitution/anagram over 380 languages) | Hauer and Kondrak, Kuykendall, and Rozanova and Temerev use the same approach. The coverage here (21 language models, 3 symbol views, both boundary modes) and the cost charged for surplus symbols are details. |
| 48 | Many-to-one (homophonic) solver | Adapted | same | Rozanova & Temerev 2026, arXiv:2608.17096… | Rozanova and Temerev, and Kuykendall, fix the number of classes in advance. Here the annealer chooses how many symbols share a letter, under a description-length penalty, and the collapse of the unpenalised objective to 'iiii...' is documented. |
| 49 | Verbose/null-cipher solver | Adapted | New | Rozanova & Temerev 2026, arXiv:2608.17096, (BPE-learned multi-glyph units mapped to letters | Rozanova and Temerev fix the segmentation first, by byte-pair encoding, and substitute afterwards. The solver here searches the segmentation, the nulls and the letters together, with a stated cost for each null, and it reports word-level recovery on verbose ciphers with a known key. |
| 50 | Meaningless controls for every solver run | Established | Adapted | Rozanova & Temerev 2026, arXiv:2608.17096… | Rozanova and Temerev use the same idea with an order-3 imitation. The imitation here is order 2, and the within-word shuffle is added as a second floor. |
| 51 | Held-out key-transfer test | Adapted | New | Parisel 2026, arXiv:2604.25979… | Parisel tested the same idea, that one key cannot serve both A and B, from the switching of character ratios. Fitting a real decipherment key on A, scoring it on B, and comparing the loss with a genuine Latin cipher (0.000) and a Markov control was not found. |
| 52 | Dictionary hit rate of decoded tokens vs the decoded Markov control | Established | same | Hauer & Kondrak 2016, TACL 4, (dictionary-based word accuracy of decodings) | Hauer and Kondrak measured the dictionary hits of their decodings. The comparison with a decoded control is an addition. |
| 53 | Segmentation-free repeated-substring test | New | same | none found | Zandbergen and Timm noted that the text repeats no long phrases. Counting repeated long strings with the spaces removed, so that the count cannot be changed by the cipher type or by errors in word division, was not found in earlier work. |
| 54 | Cipher transformations of Latin | Adapted | same | Bowern & Lindemann 2021, Annual Review of Linguistics 7 (bigraph/verbose encodings of Latin lower h2 | Bowern and Lindemann, and Rozanova and Temerev, enciphered known text and measured h2. The joint criterion, low h2 together with no long repeats, and the position-keyed variant are additions here. |
| 55 | Variant-collapse recount of repeated word 4-grams after merging unreliable or variable distinctions | Adapted | same | Timm 2014, arXiv:1407.6639, (the few repeated sequences differ by spelling variants and word order) | Timm and Zandbergen raised the idea that spelling variation hides repeated phrases. The test here merges the variable distinctions, recounts the repeats and compares them with a shuffled control. It was not found in earlier work. |
| 56 | Label-free unigram profile matching | Established | same | Zandbergen, voynich.nu character statistics | Comparing frequency profiles is standard, from Guy to Lindemann and Bowern. Reading the profile as a statistic that no transposition can change is a new framing. The measurement is not new. |
| 57 | Within-word bag-of-letters statistics | Adapted | same | Hauer & Kondrak 2016, TACL 4… | Hauer and Kondrak proposed that the words are alphabetically sorted anagrams. The test here counts precedence violations, with a lower bound that no ordering of the alphabet can beat, and measures types per bag and multi-order bags. The controls are Latin texts sorted or anagrammed with a known key. |
| 58 | Anagram-invariant substitution solver | Established | same | Hauer & Kondrak 2016, TACL 4… | Hauer and Kondrak solved the same problem with the same kind of objective. Added here are the Markov and shuffled controls, positive controls in five languages scored against held-out lexicons, and both mapping modes. |
| 59 | Route / columnar transposition test | New | same | none found | D'Imperio listed transposition among the possibilities without testing it, and Rozanova and Temerev put it outside their scope. No earlier work re-reads the glyph stream under route geometries and scores the readings on entropy and repeats with known-key controls. |
| 60 | Order-invariant phrase-repetition test | New | same | none found | Timm and Zandbergen recorded the missing repeated phrases. The statistic here counts repeated runs of letter bags, which anagramming within words and relabelling cannot change. It was not found in earlier work. |
| 61 | Entropy constraint under transposition | Adapted | same | Pelling 2018, Cipher Mysteries, (qualitative: adjacent-letter structure rules out arbitrary anagramming) | The argument is standard cryptanalysis, and Pelling made it in words for the Voynich text. The numbers here come from Latin of matched size transposed with known keys. They show the one exception, sorted anagrams, which keep h2 low. |
| 62 | Codebook / nomenclator simulation grid | New | same | none found | The rank-matched codebook generator (row E6) is extended into a grid of homophones, nulls and word splitting, scored on the word-order information as a budget and on the 1-edit rates. Pelling argued in words that the text has too few shapes for a nomenclator. No earlier quantitative simulation of a nomenclator against the Voynich statistics was found. |
| 63 | Words-as-numerals test via slot dependence | Adapted | New | Feng & Hu 2016 (supervisors Abbott & Ng), Adelaide project 'Cracking the Voynich manuscript code'… | Feng and Hu pursued the reading of the words as numbers by pattern matching. The test here measures how much one end of a word tells about the other, with generated numeral tables as positive controls. |
| 64 | Labelling-free leading-unit distribution test | New | same | none found | No earlier application of Benford's law to the text was found. The test turned out to have no diagnostic power, so nothing rests on it. |
| 65 | Ordered-table tests | New | same | none found | The ordered-table idea has been discussed on voynich.ninja. No earlier quantitative test of it with generated tables as controls was found. |
| 66 | Syllabic and letter-pair codes of real languages | Adapted | same | Stolfi 2002, 'Chinese theory redux', (syllable-inventory argument for an East Asian syllabic reading) | Stolfi, and Rozanova and Temerev, compared a real syllabic text, Chinese in pinyin, with the Voynich lexicon. Here European texts are written out in syllables and in letter pairs. The entropy and inventory comparison and the argument that a word-level code cannot change the repeated 4-gram count are added. |
| 67 | Derived-channel battery | Adapted | same | Matlach, Janeckova & Dostal 2022, PLoS ONE 17(1) e0260948… | Matlach, Janeckova and Dostal proposed one steganographic scheme and one diagnostic, the autocorrelation of symbol reuse. Here 29 derived channels are each tested for any learnable sequential structure, with held-out models and shuffle nulls. |
| 68 | Planted-message positive controls | Adapted | same | Matlach, Janeckova & Dostal 2022, PLoS ONE 17(1) e0260948… | Matlach, Janeckova and Dostal built a Voynich-like text from a message. Here a message is planted into one derived channel of the real word sequence. The power of the channel test to detect it is then measured as a function of length. |
| 69 | Chain-rule decomposition of the bias-corrected adjacent-word MI by slot | Adapted | same | Rozanova & Temerev 2026, arXiv:2608.17096… | Rozanova and Temerev, and Smith, measured the coupling between the glyphs at word edges. Here the whole word-to-word information is split by the chain rule into the parts that the prefix, the core and the suffix of the next word account for. |
| 70 | Word recurrence by physical distance class | Adapted | same | Montemurro & Zanette 2013, PLoS ONE 8(6)… | Montemurro and Zanette, Schinner and Timm established that words recur over long ranges. Here the distances are binned by the physical make-up of the book (line, page, leaf, quire, the gathering of leaves), with an exact expectation and a null stratified by hand and Currier language. The result is compared with fitted copying and page-subset generators. |
| 71 | Held-out within-page predictive gain | Adapted | same | Sterneck, Polish & Bowern 2021, arXiv:2107.02858… | Sterneck, Polish and Bowern ran topic models on the pages. Added here are a held-out predictive gain with a buffer between the halves and stratification by hand and Currier language. The tests also show that LDA, the topic model, has no power at this size even on natural-language controls. |
| 72 | PERMANOVA pseudo-F on page-by-page JSD | Adapted | same | Montemurro & Zanette 2013, PLoS ONE 8(6), (section vocabularies and links between sections) | Montemurro and Zanette, Zandbergen and Currier established that sections have their own vocabulary. Here the test is a PERMANOVA, a permutation test of group differences on a distance matrix. Its nulls preserve contiguity or are stratified by hand and Currier language, and the pages are subsampled first to correct the entropy bias. |
| 73 | Image-text association on the herbal pages | New | same | none found | Montemurro and Zanette, and Sterneck, Polish and Bowern, compared vocabulary with illustration themes at the section level. No earlier page-level statistical test links measured image features to page vocabulary with a planted positive control. |
| 74 | Function-word test | Adapted | same | Montemurro & Zanette 2013, PLoS ONE 8(6)… | Montemurro and Zanette applied the burstiness criterion for function words to the text. Added here are the contrast of the top-30 words with the mid-band under stratification by hand and Currier language, and the positional and neighbour-entropy features. A page-subset generator reproduces the Voynich pattern. |
| 75 | Label tests | Adapted | same | Pelling 2017, 'Voynich labelese' | Pelling and Zandbergen discuss how labels differ from running text. The shuffle and draw nulls for recurrence across pages, and the derangement null for similarity to the same page, are additions here. |
| 76 | Held-out Kneser-Ney cross-entropy against n-gram order 1-8 | Adapted | same | Lindemann & Bowern 2021, arXiv:2010.14697, (conditional entropy h2 across 316 texts) | Lindemann and Bowern, Parisel and Zandbergen stop at low orders and plug-in estimates. Here the gain beyond order 3 is cross-validated, and the learning curve against training size is compared with imitations of the text at matched order. |
| 77 | Word-level held-out bigram gain | Adapted | same | Rozanova & Temerev 2026, arXiv:2608.17096, (shuffle-corrected adjacent-token MI) | Rozanova and Temerev, and Parisel, estimate the same quantity by plug-in counts corrected with a shuffle. Here it is a held-out predictive gain. |
| 78 | PPMI-SVD word-embedding geometry | Established | Adapted | Perone 2016, 'Voynich Manuscript: word vectors and t-SNE', (word2vec + t-SNE on Voynich words) | Perone, Baum and colleagues, and an EPFL course paper built word embeddings of the text. Added here are matched meaningless controls, and the hubness result comes from them. |
| 79 | 1-edit word pairs as distributional neighbours | Adapted | same | EPFL MLO course paper 2022, 'Word embeddings for the morphosyntactic analysis of the Voynich manuscript'… | The EPFL course paper and Perone used embeddings to study word variants. The test here, 1-edit pairs against rank-matched pairs in nearest-neighbour lists, with split-half and context-restricted controls and generator baselines, was not found. |
| 80 | Unsupervised cross-lingual embedding alignment | Established | Adapted | Baum et al. (n.d.), voynich2vec (seminar project under C. Bowern)… | Baum and colleagues tried the alignment in a seminar project. The result here is negative. The known-answer and supervised controls show that the method has no power at this corpus size. |
| 81 | Extended language survey | Adapted | same | Lindemann & Bowern 2021, arXiv:2010.14697… | Lindemann and Bowern, Hauer and Kondrak, and Matlach and colleagues surveyed many languages on one statistic each (entropy, language-model likelihood, autocorrelation). Here the profile has several statistics that a fixed cipher cannot change, with a combined distance ranking, and it includes abjad and Indic corpora. |
| 82 | Offset sensitivity of the repetition counts | Established | Adapted | Rozanova & Temerev 2026, arXiv:2608.17096, (quire bootstrap and disjoint-half resampling) | A check of sampling variability. No new technique. |
| 83 | Known-key cipher positive controls in Hebrew, Arabic and Sanskrit and the solver grid | Established | same | Hauer & Kondrak 2016, TACL 4, (Hebrew | The substitution solver (row F1), which maps each symbol to 1 letter, and its meaningless controls (row F4) are run on more languages. The technique is the same. |
| 84 | Drifting-state generator | Adapted | same | Timm & Schinner 2020, Cryptologia 44(1)… | Timm and Schinner's generator clusters words by copying nearby ones, and other generators use fixed page or line vocabularies. The generator here makes the clustering a hidden state with fitted timescales. It is fitted only to the clustering statistics, and every other statistic is reported as an out-of-sample prediction. No latent-state generator for this text was found. |
| 85 | Generator fitting protocol | Adapted | same | Timm & Schinner 2020, Cryptologia 44(1)… | Timm and Schinner, Parisel, Rozanova and Temerev, and Rugg report which statistics their generators match. The protocol here (the rules the fitting follows) states the objective, its tolerances (the range a statistic wanders over between halves of the text) and the seed variance (the spread between runs with different random seeds), and it separates the statistics that were fitted from those that are predicted. |
| 86 | One-parameter verbatim pair memory | Adapted | same | Parisel 2026, arXiv:2604.19762, (slot-based parametric generator with 12 ablations | Parisel's word-Markov simulations fit the whole transition table. Here one memory rate inside the drift generator is fitted to one statistic, and the longer-range information and the held-out gain are the checks. |
| 87 | Section-base variant of the drift generator | Adapted | same | Montemurro & Zanette 2013, PLoS ONE 8(6), (section vocabularies and links between sections) | Montemurro and Zanette, and Sterneck, Polish and Bowern, measured the section vocabularies. Whether a fitted drift generator reproduces them, and at what cost to the other statistics, had not been tested. |
| 88 | Re-implementation self-check | Adapted | same | Rozanova & Temerev 2026, arXiv:2608.17096… | A verification step. It is not a new measurement, and it is listed because the generator's scores depend on it. |
| 89 | Variant genealogy along codex order | Adapted | same | Timm & Schinner 2020, Cryptologia 44(1), (network of similar words spanning the manuscript | Timm and Schinner infer the process from the network of similar words. The test here asks whether the codex shows the ordering that a vocabulary evolving during writing would leave. It uses page-order permutation nulls and positive controls for a global and a local genealogy. It finds no ordering. |
| 90 | Fit-to-space at the line end measured on the scans | New | same | none found | Vogt, Feaster and Smith measure word length and glyph forms by position in the transliteration, and Zandbergen describes the right margin in words. Measuring the physical space left at the margin, against a simulated copyist filling the measured width, was not found. |
| 91 | Conjugate-leaf vocabulary sharing | Adapted | same | Fagin Davis 2020, Manuscript Studies 5(1) (five scribes identified palaeographically | Fagin Davis, Zandbergen and Pelling raised the idea that the bifolium (a sheet folded once to give two leaves) was the unit of production, from palaeography and codicology. The test here uses the vocabulary, with controls matched by distance, hand and language, and with language texts re-flowed into the same pages. |
| 92 | Ink change points against vocabulary drift | New | same | none found | The ink has been described by eye on voynich.ninja and analysed chemically by McCrone Associates. A page-by-page ink series compared with a vocabulary-drift series was not found. |
| 93 | Transcription-free glyph segmentation of full-resolution scans | Adapted | New | Ponzi (n.d.), 'Neural-network handwritten text recognition for the Voynich manuscript', Medium/ViridisGreen… | Ponzi's pipeline is a supervised recogniser trained on transliterations. The pipeline here cuts units without any training and aligns them to the transliteration by dynamic programming. Every crop gets an EVA label, and the quality of the alignment is stated for each line. |
| 94 | Global shape clustering of aligned glyph crops | New | same | none found | No unsupervised clustering of glyph images against the transliteration alphabet was found for this manuscript. Confidence is low, because forum or blog experiments of this kind may exist. |
| 95 | Allograph-sequence channel test | New | same | none found | Transcribers have recorded the form variants, and Painter and Bowern and Fagin Davis used them to separate hands. Testing the sequence of variants for message-like structure, with shuffle nulls, the scribe's habits as covariates and a substituted-language control, was not found. |
| 96 | Whole-shape stream statistics | Adapted | same | Zandbergen, (entropy under different transcription alphabets | Zandbergen, and Rozanova and Temerev, measured entropy under alternative transliterations. Applying it to a symbol stream derived from the images, and reading the excess as clustering noise, is the extension. |
| 97 | Held-out information budget | Adapted | same | Zandbergen, (word entropy 9.9 bits at 8,000 words vs Dante 9.1, Pliny 10.6 | Zandbergen's estimate is a word entropy, and Bennett's, Lindemann and Bowern's and Parisel's are character entropies. The model here is a held-out structured model. Each stage prices one known regularity, and the residual is compared with generators of the same shape. |
| 98 | Residual-channel tests | New | same | none found | Matlach and colleagues, and Rozanova and Temerev, ran sequential tests on symbol streams. Testing the residual of a fitted predictive model as a channel, stream by stream, against generators of the same shape was not found. |
| 99 | Planted low-rate channel control and residual decipherment step | Adapted | same | Matlach, Janeckova & Dostal 2022, PLoS ONE 17(1) e0260948… | Matlach and colleagues planted a message in a text, and Hauer and Kondrak and Rozanova and Temerev validated solvers on known keys. Applying those controls to the 2 residual streams of a fitted model, prefix and suffix, with 1 planted letter per word, is the extension. |
| 100 | Naibbe cipher re-implemented from the published code with fixed seeds | Established | same | Greshko 2025, Cryptologia, (Naibbe cipher | Greshko compared the Naibbe ciphertext with the text on fewer statistics, and Rozanova and Temerev on a few more. Here the full battery, the content and embedding diagnostics and the order-2 Markov imitation are added. The extra statistics do not change the class. |
| 101 | Substitution solvers on Naibbe ciphertext | Adapted | same | Rozanova & Temerev 2026, arXiv:2608.17096… | Rozanova and Temerev score the Naibbe output with 1 language-match differential. Here the three substitution solvers and the key-transfer test are run on it, scored against the known plaintext, and the failure is stated as a limit of the solvers. |
| 102 | Homophone-exchangeability test | Adapted | same | Rozanova & Temerev 2026, arXiv:2608.17096, (Brown clustering of 88 glyph units into 23 classes by context) | Rozanova and Temerev, Matlach and colleagues, and Guy group glyphs or units by their contexts. The test here asks whether frequent word types are exchangeable in context, with a permutation null and a positive control of known homophone classes. |
| 103 | Spelling-variation sensitivity | Adapted | same | Bowern & Lindemann 2021, Annual Review of Linguistics 7… | Bowern and Lindemann, and Timm and Schinner, compare with real historical texts. Here synthetic variation is added in measured doses, and the repetition, entropy and family statistics are read together. |
| 104 | Transcription-collapse sensitivity on the Voynich side | Adapted | same | Zandbergen, (entropy under different transcription alphabets | Zandbergen, and Rozanova and Temerev, report entropy and repeat counts under merged alphabets. The progressive collapse here, with shuffled chance levels from 3 seeds, the 1-edit family collapse and the matched collapses of Latin and Italian, extends the variant-collapse recount. |
| 105 | Verbose-expansion matching | Adapted | same | Rozanova & Temerev 2026, arXiv:2608.17096, (variable 1-3-glyph groups in the calibrated attack | Tiltman raised the verbose-cipher idea long ago, and Rozanova and Temerev calibrated the group sizes. Reading the repetition threshold in plaintext-equivalent lengths under both size conventions was not found. |
| 106 | Eighteen deterministic state-dependent encodings of Latin in the skeleton | Adapted | same | Bennett 1976, Scientific and Engineering Problem-Solving with the Computer (character entropies | The entropy argument against polyalphabetic substitution is standard, from Bennett onward. Here it is applied together with the repeat, family and information-placement statistics to eighteen fully specified schemes, among them autokey, state machines and counter-chosen nulls. |
| 107 | Solver-detection check on state-dependent ciphers | Adapted | same | Rozanova & Temerev 2026, arXiv:2608.17096, (calibrated attack with Markov surrogates) | Hauer and Kondrak, and Rozanova and Temerev, validated their solvers on known keys. Running the detection protocol on ciphers outside the set of ciphers the solver can find, to state what it covers, was the added step. |
| 108 | Diplomatic medieval reference set | Adapted | same | Lindemann & Bowern 2021, arXiv:2010.14697, (18 transcribed historical texts in 8 languages | Lindemann and Bowern, Gheuens and Boxer compared a few historical texts, mostly on entropy. Here the transcriptions are diplomatic, with the abbreviation marks kept as symbols, and the same text appears at several levels of faithfulness. The full repetition, family and adjacency profile is measured with class ranges over windows. |
| 109 | Transcription-policy effect table on the reference side | Adapted | same | Lindemann & Bowern 2021, arXiv:2010.14697, (abbreviation and scribal-convention effects on entropy examined) | Zandbergen, and Rozanova and Temerev, examined the effect of transcription policy on the Voynich side (alphabets, separators), and Lindemann and Bowern examined it for entropy on the reference side. The step-by-step table over the whole profile of statistics is the extension. |
| 110 | Simulated scribal abbreviation and spelling variation as a rule model | Adapted | same | Lindemann & Bowern 2021, arXiv:2010.14697, and Bowern & Lindemann 2021, Annual Review of Linguistics 7… | Synthetic variation on reference texts is shared with the spelling-variation test (row R1). The calibration against the scribal inconsistency measured in real manuscripts is the added step. |
| 111 | Word co-occurrence network topology | Established | same | Amancio, Altmann, Rybski, Oliveira and Costa 2013, PLoS ONE 8(7) e67310, (EVA transcription) | Amancio and colleagues made the same measurement. Added here are a second null, the line-interior shuffle, which removes the assortativity excess, and 14 language chunks of matched size. Also added are 9 generators poured into the same page and line skeleton and a triad census verified against the reference implementation. |
| 112 | Intermittency of word recurrence | Established | same | Amancio et al. 2013 | Amancio and colleagues compared the text with one book in 15 translations. Here the reference set is 28 corpora of matched size, some of one genre and some of mixed topics. Each Currier language, each scribal hand and nine generators are scored separately. |
| 113 | Information over scales curve | Established | same | Montemurro and Zanette 2013, PLoS ONE 8(6) e66344 | Montemurro and Zanette made the same measurement. Added here are stratified, within-page and within-paragraph shuffles, which take the curve apart, and a moving-block bootstrap on the position of the peak (512 in 156 of 200 resamples). Nine generators are scored too, the fitted drift generator among them. |
| 114 | Keyword ranking by per-word information or block entropy with bootstrap significance, and the share of significant keywords | Established | same | Montemurro and Zanette 2013 | Montemurro and Zanette made the same measurement. Here the identical criterion is applied to 28 languages of matched size and nine generators. The counts are repeated per hand and per Currier language, and the resulting keyword list is tested for concentration in one section. |
| 115 | Paragraph-similarity network modularity against paragraph-shuffled text | Established | New | de Arruda, Marinho, Costa and Amancio 2018, Paragraph-based complex networks: application to document… | De Arruda, Marinho, Costa and Amancio compared the text with a paragraph shuffle. Here the reference set is 28 languages of matched size and nine generators poured into the same paragraph skeleton. A null stratified by hand and Currier language is added beside the paragraph shuffle. |
| 116 | Adjacent-page similarity | Established | same | Reddy and Knight 2011, What we know about the Voynich manuscript, LaTeCH at ACL… | Reddy and Knight made the same measurement. Added here are a free permutation null and one stratified by hand and Currier language, a tf-idf variant, two transliterations, 28 languages, nine generators and a planted positive control. |
| 117 | Two-state letter HMM and the a*b word grammar | Established | same | Reddy and Knight 2011 | Reddy and Knight, and Acedo, made the same measurement. Added here are three symbol views in place of one and random restarts beside the a*b initialisation. The within-word shuffle and the Markov-2 imitation are the nulls, and Latin without vowels, Hebrew and Arabic are the positive controls. |
| 118 | Word-class induction | Established | same | Reddy and Knight 2011 | Reddy and Knight made the same measurement. Added here are the interior-shuffled text as a floor for the gain, a word-Markov generator as a ceiling and five folds. The induced classes are compared with the function words of the Latin control. |
| 119 | Reading-direction cross-entropy asymmetry Delta_n = X_LTR - X_RTL | Established | New | Parisel 2025, arXiv:2509.10573, (Kaggle code linked from the paper) | Parisel made the same measurement. Added here are three transliterations, 15 languages, the within-word and character shuffles (both at zero) and five generators. Two smoothing families and three orders are run, and a worked reconstruction shows what an asymmetric edge convention would produce. |
| 120 | Boundary signatures | Established | New | Parisel 2026, arXiv:2604.19762 | Parisel made the same measurement. Added here are three transliterations, the within-word shuffle as a negative control, and the Markov-2 imitation and five generators scored on the same four signatures. The boundary information is bias-corrected by the line-interior shuffle. |
| 121 | Internality index of uncertain word separators and line breaks | Established | New | Rozanova and Temerev 2026, arXiv:2608.17096 | Rozanova and Temerev made the same measurement, but their description of the index admits two readings, and both are computed here. Added are four symbol views and two transliterations, 17 poured languages and a planted internal-boundary positive control. Three groups of Middle High German manuscripts whose scribal and editorial boundaries are known are added as well. |
| 122 | Physical gap width on the scans as a predictor of transcriber space certainty | Established | New | Rozanova and Temerev 2026 | Rozanova and Temerev made the same measurement from published word boxes. Here the gaps are measured from the scans. Added are blind predictions written down before the values were read, a within-line permutation null and a breakdown by hand. A positive control compares certain spaces with within-word gaps (AUC 0.913). |
| 123 | Dependence gap across BPE merge levels D_k = H | Established | New | Rozanova and Temerev 2026 | Rozanova and Temerev made the same measurement without controls. Added here are three space conventions, a page-block bootstrap, the Miller-Madow correction, nine languages, the within-word shuffle, the Markov-2 imitation and five generators. |
| 124 | Quire bootstrap and leave-one-quire-out confidence intervals for the headline statistics | Established | New | Rozanova and Temerev 2026 | Rozanova and Temerev used the same resampling design. Added here are a page bootstrap run beside the quire bootstrap for comparison, and matched PERMANOVA nulls without the two most influential quires. The tests also show that resampling with replacement corrupts every repetition and information statistic. |
| 125 | Table-and-grille generator | Established | same | Rugg 2004, Cryptologia 28(1):31-46, (abstract only) | The generator family is Rugg's. Added here are the full 23-statistic battery with page-half confidence intervals, extra repetition and boundary-information statistics, and 168 configurations over five seeds. A null device without locality isolates the effect of renewing the table. |
| 126 | Wheel (volvelle) grille generator | Adapted | same | Zandbergen 2021, The Cardan grille approach to the Voynich MS taken to the next level, arXiv:2104.12548 | Zandbergen's proposal is a sketch of a mechanism. Added here is a direct product-set test, which asks how much of the real vocabulary a bounded wheel product can cover. The full 23-statistic battery is run as well, with extra statistics for lexicon growth, the recurrence profile and boundary information, over six devices and five seeds. |
| 127 | Human-produced gibberish reference class and the 42-metric random-forest classifier | Established | same | Gaskell and Bowern 2022, Gibberish after all? Voynichese is statistically similar to human-produced samples… | Gaskell and Bowern's classifier and metrics are used unchanged. Added here are the two shuffles and the Markov-2 imitation passed through the same forest. The full battery is run on all 38 gibberish and 71 meaningful texts at their own 200-word scale, and a pooled profile is ranked against 29 languages. |
| 128 | Detrended fluctuation analysis, Hurst exponent and spectral slope of word-level series | Established | same | Schinner 2007, The Voynich Manuscript: evidence of the hoax hypothesis, Cryptologia 31(2):95-107… | Schinner, and Montemurro and Pury, report an exponent or an autocorrelation without taking it apart. Here three targeted shuffles (line-interior, line-order, page-order) locate the correlation between the line and the page, and 30 languages, six generators and the residual-surprisal series are added. |
| 129 | Spectral analysis of the spaceless glyph stream | Established | same | Landini 2001, Evidence of linguistic structure in the Voynich manuscript using spectral analysis, Cryptologia… | Landini and Arutyunov made the same spectral measurement. Added here are ten shuffle seeds, a white-noise shuffle, the within-word shuffle, the Markov-2 imitation and four languages. A second symbol view with merged glyphs is used, and the line-band component that the published abstract only asserts is tested directly. |
| 130 | Distribution of the distance to the next similar word and its test against the geometric law | Established | same | Schinner 2007 (abstract only) | Schinner and Timm tested the same distance law. Added here are a hazard-ratio decomposition by distance band, the text's own shuffle as the memoryless null, 30 languages and seven generators. A copying positive control and a page bootstrap on the short-distance excess are added as well. |
| 131 | Word and glyph mutual information across line and paragraph boundaries | New | same | Currier 1976 (line as a functional entity | No earlier measurement of mutual information across the line break was found. Currier called the line a functional entity, and Landini and Vogt took the idea up. The design here has three independent nulls: a same-page permutation, a generator with a line-start state as the negative control, and poured languages as the positive control (z 4-43). Two transliterations and a glyph-level counterpart are added. |
| 132 | Enochian | Established | same | Boxer 2022, Fingerprinting gibberish: a quantitative comparison of the Voynich and Sloane MS 3188, CEUR-WS… | Boxer's two statistics are used unchanged. Added here are the word-order shuffle of each text as a drift baseline, page-block resampling without replacement and the Markov-2 imitation. The comparison set has 29 languages, and the full battery is run at the Enochian sample size. |
| 133 | Polygraphia III style letter-to-invented-word cipher as a reference text | Established | same | Hermes 2022, Polygraphia III: the cipher that pretends to be an artificial language, CEUR-WS Vol-3313 paper 7 | Hermes compared the cipher with the text descriptively. Here the cipher class is generated at matched size over 14 table settings and three seeds. It is scored on the same 23-statistic battery as every other generator, with extra repetition and boundary-information statistics. |
| 134 | Non-linguistic symbol corpora | New | same | Sproat 2014, A statistical comparison of written language and nonlinguistic symbol systems, Language… | Sproat's and Nair's scorecards had never been applied to this text. Here they are, under two stated mappings (glyph in word, word in line), with permutation nulls for every scorecard statistic, the within-word and word-order shuffles and the Markov-2 imitation. The published calibration ranges were re-derived as a check. |
| 135 | Genuine historical ciphertexts | New | same | Knight, Megyesi and Schaefer 2011, The Copiale cipher | No genuine ciphertext had been compared with this text. The comparison here is made at matched size on symbol streams without spaces. A homophonic cipher of the same plaintext was built as a positive control, a symbol shuffle is the null, and 29 languages and three synthetic cipher families are measured alongside. |
| 136 | Index of coincidence, periodic IoC and Kasiski spectra at lags up to 200, and the cipher-type feature vector of Nuhn and Knight | Adapted | same | Nuhn and Knight 2014, Cipher type detection, EMNLP | The index of coincidence is a classical statistic, going back to Friedman. Here it is run at every lag up to 200 in three symbol views, with permutation nulls. A full set of synthetic positive controls is added (Vigenere of three periods, autokey, word-keyed, state-dependent, verbose, homophonic, transposition), and the index is measured page by page beside matched Latin. |
| 137 | Word-pattern (isomorph) fingerprint | New | same | Hauer and Kondrak 2016, Decoding anagrammed texts written in an unknown language and script, TACL 4:75-86 | Hauer and Kondrak use the word patterns as a decipherment tool for anagrammed text. Here the distribution of patterns is a language fingerprint, measured at matched size with a sampling floor, three symbol views, a monoalphabetic positive control, and generator and cipher controls. |
| 138 | Glyph-level homophone detection by hierarchical clustering of symbol contexts | New | same | Lehofer 2022, Applying hierarchical clustering to homophonic substitution ciphers using historical corpora… | Lehofer demonstrated the method on historical ciphers. Here it is applied across three symbol views, with two synthetic homophonic positive controls and natural-language negative controls. The within-word shuffle is a sensitivity check, and a token alignment identifies which fine distinctions are exchangeable. |
| 139 | Kober triplets and paradigm consistency across stems | Adapted | same | Kober 1948, The Minoan scripts: fact and theory, American Journal of Archaeology 52:82-103… | Kober's heuristic is qualitative. Here it becomes three quantitative statistics with a permutation null. They are applied to nine languages, three symbol views and five generators, with a monoalphabetic positive control and Latin written in syllables as a second positive control. |
| 140 | Scribe attribution by character n-gram classifiers and Burrows' Delta against the Fagin Davis hands | Established | same | Farrugia, Layfield and van der Plas 2022, CEUR-WS Vol-3313 paper 5, (ZL transcription) | Farrugia, Layfield and van der Plas used the same classifiers. Added here are a label permutation null within strata of Currier language and section, sub-tests within one language and within one section, a poured-Latin negative control and the drift generator. Two positive controls are added as well: one text in two simulated scribal habits (0.96-0.97) and two manuscripts by different scribes (1.00). |
| 141 | Rightwardness and downwardness scores of graphemic minimal word pairs | Established | same | Feaster 2022, Rightward and downward grapheme distributions in the Voynich manuscript, CEUR-WS Vol-3313 paper… | Feaster's minimal-pair test is used unchanged. Added here are a within-line permutation null, an interior-only variant that removes the known line-edge effect, and a control family of every other single-glyph substitution. Both transliterations are used, poured languages are scored on their own pairs, and word-final m and g are the positive control. |
| 142 | Per-scribe paragraph-initial habits and the interaction of within-word position with position in the paragraph | Established | same | Stafford 2022, Seven habits of highly eccentric paragraphs, CEUR-WS Vol-3313 paper 13 | Stafford described the same habits. Added here are exact counts for all 740 paragraphs with a page bootstrap, and two permutation nulls for the distance between hands (within section and across all pages). Poured Latin and Italian, the within-word shuffle and the Markov-2 imitation are the controls for the position interaction. |
| 143 | Segment information up to the uniqueness point as a function of word frequency | Established | same | Layfield, van der Plas, Rosner and Abela 2020, Word probability findings in the Voynich manuscript, LT4HALA… | Layfield and colleagues made the same measurement. Added here are the within-word shuffle, the Markov-2 imitation and six generators run through the identical trie pipeline, and 30 languages at matched size. A partial correlation on word length and a breakdown by frequency band are added as well. |
| 144 | Word length vs frequency | Adapted | same | Piantadosi, Tily and Gibson 2011, PNAS 108(9):3526-3529 | Piantadosi, Tily and Gibson used very large corpora and a surprisal model trained on them. Here both are estimated on matched samples of 34,780 tokens with a cross-validated surprisal, and the within-word shuffle, the Markov-2 imitation and six generators are scored alongside 30 languages. |
| 145 | Menzerath-Altmann law between word length in units and mean unit length | New | same | Menzerath-Altmann law (Altmann 1980, Glottometrika 2) | No value of this law had been published for the text. Here it is fitted through one segmentation pipeline applied identically to every text, with the within-word shuffle as the null. Three alternative segmentations, six languages, five generators and a bootstrap over types are added. |
| 146 | Unsupervised word segmentation of the spaceless glyph stream and its agreement with the transcribed spaces | New | same | Jin and Tanaka-Ishii 2006, COLING/ACL | Neither of the two segmentation methods, Jin and Tanaka-Ishii's and Goldwater, Griffiths and Johnson's, had been applied to this text. Both are run here at a rate-matched operating point, with the within-word shuffle as a floor and the Markov-2 imitation as a ceiling. Three languages of matched size are scored too, and recall is measured separately at the uncertain positions. |
| 147 | Long-context neural language model gain over Kneser-Ney n-grams | New | same | Standard method | No peer-reviewed neural comparison exists for this text. The design here fits the neural model and the n-gram baseline on exactly the same held-out folds. It adds the Markov-2 imitation as a zero-gain reference and a copying generator as a positive one, and it measures the gain as a function of context length. |
| 148 | Paired single-leg gallows on paragraph top lines | Adapted | New | Pelling, Cipher Mysteries blog posts of 2020-08-29 and 2025-07-12, (grey literature | Pelling's claim is grey literature with no statistics. Here it is given a binomial placement model, a within-line permutation null and a breakdown by hand. A Markov-2 imitation control and two placebo families in languages, chosen by the same rule, are added. |
| 149 | Stroke-level symbol view | Adapted | New | Cham 2014, Curve-line system, (grey literature) | Cham's proposal is grey literature with no numbers. Here a documented stroke table is written out and applied to the text, and to Latin and Italian with a comparable table. Two table permutation controls separate the effect of the recoding from the text. |
| 150 | Fractal dimension, sliding-window network centrality maps and visibility-graph statistics of text series | Established | same | Zelinka, Lara, Windsor and Lozi 2023, Softcomputing in identification of the origin of Voynich manuscript by… | Zelinka and colleagues made the same measurements. Added here are the text's own word shuffle and within-word shuffle as nulls, the drift generator and the Markov-2 imitation, and six languages. A language-to-language distance is the yardstick for differences between maps. |
| 151 | Two-language mixture test from the deviation of sorted letter frequencies from a logarithmic law, and spectral portraits of two-letter distributions | Established | New | Arutyunov, Borisov, Fedorov, Ivchenko, Kirina-Lilinskaya, Orlov, Osipov, Pyrkov, Shilin and Zeniuk 2016… | Arutyunov and colleagues argue from a fit statistic on sorted frequencies. Here the statistic is given single-language references of matched size and constructed bilingual references in the same page skeleton. A held-out two-component mixture model with a third-component check, restrictions to one hand and to one Currier language, and a spectral portrait of the bigram matrix are added. |
| 152 | Phonetic-prior decipherment into candidate known languages with a language-closeness score | New | same | Luo, Hartmann, Santus, Barzilay and Cao 2021, Deciphering undersegmented ancient scripts using phonetic… | Luo and colleagues' released model needs a GPU and phonetic transcriptions of every candidate vocabulary. The substitute here scores closeness as the gain of a validated homophonic solver on the text over the same solver on the text's own Markov-2 imitation. A monoalphabetic cipher is the positive control and a language isolate the negative control. |
| 153 | Autoencoder visual similarity of glyph shapes within the Voynich alphabet and to the letterforms of other scripts | Established | same | Zelinka, Lara, Windsor and Lozi 2023, open copy | Zelinka and colleagues made the same autoencoder comparison from font renderings. Here the glyphs are aligned crops from the scans. Added are a null script of random strokes and the spread within a class as the scale of 'same shape'. The positive control is a set of known alphabets rendered in other fonts and distorted, which must find their own script. |
| 154 | Unicity-distance table per cipher family | Adapted | same | Shannon 1949, Communication theory of secrecy systems, Bell System Technical Journal 28(4):656-715, (theory | Von zur Gathen computed the unicity distance for one historical cipher, the Zodiac-340, with one key count and one corpus entropy. The published solver calibrations report accuracy against ciphertext length for English and do not refer to the unicity distance. The one computation in a Voynich context is a rough forum figure for one proposed cipher. Here the distance is tabulated for every cipher family against the manuscript's own ciphertext lengths, in three symbol readings, per section and for the labels. The redundancy is the most conservative of four estimators over nine candidate languages. The Naibbe key is counted three ways, with its expansion and randomisation folded into an effective redundancy. The solvers are calibrated against the distance on Latin ciphers with a known key. Each solver failure is then sorted into one of two kinds. It is informative when the text is far above the distance and the key space was searched. It is uninformative when the family lies outside the searched key space, or when the text is at or below the distance. |
| 155 | Book of Soyga tables regenerated from Reeds's recurrence | Adapted | same | Daruka 2021, On the Voynich manuscript, Cryptologia 45(1):44-80 (online February 2020)… | Daruka made the one published comparison of the manuscript with a magical table text, Dee's Liber Loagaeth. How that text was generated is unknown, and the comparison reports statistical-linguistic matches at the whole-text level. The Voynich literature describes the Book of Soyga as algorithmically generated, but nobody had measured it as text. Here the tables are rebuilt from Reeds's published recurrence, so the generator is known exactly. They are scored on the full battery at matched size, in both reading directions and in a word view poured into the manuscript's page skeleton. The within-word shuffle, the Markov-2 imitation, three languages and a positive-control language are scored on the same statistics. |
| 156 | Hildegard of Bingen's Lingua Ignota scored as a type list | Adapted | New | veriarch.com 2025, The Abbess's Code: testing Hildegard's lingua ignota, unsigned web page dated 21 August… | The one earlier measurement found is on an unsigned web page: a normalised character entropy of the glossary against Latin, Middle High German and a random string. The published treatments of the Lingua Ignota, Higley's edition and Skowronska's classification, are qualitative. Boxer's quantitative comparison of an invented vocabulary with the manuscript concerns Enochian running text, on once-only words and repetition. Here the glossary is scored as a type list on the word-level statistics used throughout, beside size-matched type lists of the manuscript, three languages, Enochian and the Trithemius conjurations. The controls are the letters-within-word shuffle and a Markov-2 word chain fitted to the list. Two independent sources of the glossary are checked against each other, and a permutation test is run on its semantic ordering. |
| 157 | The conjurations of Trithemius's Steganographia | Adapted | same | Reeds 1998, Solved: the ciphers in Book III of Trithemius's Steganographia, Cryptologia 22(4):291-317… | Reeds, Ernst and Ponzi solved and described the Steganographia's ciphers. Boxer (2016) built tools that apply its techniques to any text and rank the outputs by letter-frequency divergence. Matlach and colleagues proposed a Trithemius-style scheme for the manuscript and tested it by symbol autocorrelation. Hermes profiled the output of Polygraphia III against the manuscript. Nobody had measured the conjurations as text, or planted the alternate-letter concealment described for them into a carrier of the manuscript's length. Here both are done, with the battery and the channel score. A Markov-2 imitation fixes the channel's baseline. An invented-word cover is generated from the conjurations, a wrong-phase control is added, and a detection curve is drawn from 500 to 5,000 channel symbols. |
| 158 | Known-plaintext and dictionary attack on the label sets | Adapted | same | Zandbergen, Some special properties of labels in the Voynich MS, voynich.nu, updated 11 September 2025… | The label literature reports frequency properties. Examples are the share of labels that occur once, ok and ot as the initials of 52% of zodiac labels, and labels beginning with o at probability 0.48. Bax, Sherwood, Tucker and Talbert offer crib readings of single labels without a test. The one dictionary attack found, a genetic-algorithm n-gram mapping of 111 herbal first words with 31 hits, has no null. Hauer and Kondrak's decipherment attack treats the whole text. Here the solvers are run on the label corpora alone, against their own Markov imitations, with a planted positive control at the same size and budget. The dictionary search reports the largest number of labels that one key maps onto a name list, beside imitation and running-text controls and a planted key. The first-glyph concentration is compared with real name lists, as a statistic that no one-to-one or homophonic relettering can change. The published identifications are scored under a null that shuffles the identifications. |
| 159 | Held-out model competition on random quire splits with the scoring set fixed in advance | Adapted | same | Timm and Schinner 2020, A possible generating algorithm of the Voynich manuscript, Cryptologia 44(1):1-19… | Timm and Schinner (relative errors), Rugg and Taylor (frequency and length features) and Parisel (four signatures) each compared one generator's output with the whole text on a list of statistics. In each case the generator was tuned on the same text. Cross-validated model selection is standard elsewhere but had not been applied to Voynich generators. Here every generator is refitted on half the quires and scored on the other half. The unit of distance is the between-quire jackknife error, and the floor is the text's own half-to-half distance. The scoring set was fixed before fitting, five splits give a stability measure, and the shuffles and the glyph imitation are placed on the same scale as the generators. |
| 160 | Held-out discriminator | Adapted | same | Gaskell and Bowern 2022, Gibberish after all? Voynichese is statistically similar to human-produced samples… | The published Voynich classifiers separate meaningful text from human gibberish, Currier A from B, or scribe from scribe on held-out folios. Gaskell and Bowern's random forest on 42 metrics, with the manuscript as a test case, is of the first kind. In other fields a classifier is used as a two-sample test to score generative models. Here the two classes are the manuscript's own pages and a generator's pages in the same skeleton. The split is by quire, so that no page's neighbours leak into the training set. The null is a label permutation refitted 100 times. A nested feature-removal analysis measures how much of the separation survives without the leading statistics. A Latin-versus-imitation positive control and an arbitrary-label negative control bracket the result. |
| 161 | Akshara rendering profile with the Kober positive control | Adapted | same | Lindemann and Bowern 2021, Character entropy in modern and historical texts: comparison metrics for an… | Lindemann and Bowern compared the text with texts as natively written in abugidas or syllabaries, by character-set size against h2. Other work induces units from the text's own glyphs. Nobody renders one language at several sign granularities, so that the same words are measured as letters, graphemes, consonant-vowel signs and aksharas. Here one abugida-written language is rendered by rule into four sign streams and profiled per sign on the battery. The battery covers h1 to h4, word-length shape in signs, one-edit share, repetition, and shuffle and Markov nulls. The Kober grid is given its first positive control from a real syllabary. The closeness score is calibrated with a renamed-sign control. That control shows the score to have no power for the syllabary question, while the excess over the real language after the best key does have power. |
| 162 | Matched-inventory PMI merging | Adapted | same | Rozanova and Temerev 2026, A glyph is not a letter, a token is not a word, a space is not a space: what the… | Rozanova and Temerev, and thevoynichproject, published merge trajectories that index the text and a few controls by the number of merges and report the dependence gap. Lindemann and Bowern, and the survey, vary the text's own character divisions and study the effect of set size. Bowern and Gaskell's fixed bigraph encodings go the other way, expanding letters into digraphs. Here 28 languages are brought to the text's own sign counts (24 and 38 signs for 99% coverage) by the text's own PMI unit-induction procedure. The entropy margin is then read at equal inventory, and the text's letters merged by the same code give a procedure-matched comparison. Finding 1 rests on the within-word shuffle gap at matched inventory, and not on the raw h2. Word length in signs is measured at the same targets. |
| 163 | Syllable canon and harmony test from the Sukhotin classes | Established | Adapted | Guy 1991, Statistical properties of two folios of the Voynich manuscript, Cryptologia 15(3):207-218… | Guy, Reddy and Knight, and Lindemann and Bowern published the Sukhotin classes. In the grey literature, thevoynichproject ran the syllable-shape shares and a harmony test with Turkish and Finnish as positive controls, using a bipartition pure-word statistic against vowel-resampled surrogates. Ponzi and Smith scored syllable-based word grammars for coverage. Here the canon is counted as pattern-coverage numbers and parsability shares against 28 languages. The harmony statistic is a mutual information across the consonant run. The null is the text's own Markov-2 imitation, which uses adjacency only, and it reproduces every canon statistic and all but 0.005 to 0.045 bits of the vowel association. The statistic and the null differ from the earlier work. The measurements and the conclusions are the same: a phonotactically unremarkable canon, and no harmony. |
| 164 | Clitic-stripping name test | New | same | Nearest, none making the measurement: Bowern and Lindemann 2021, The linguistics of the Voynich manuscript… | The lack of qo- on labels and the preference for o-, ok- and ot- are recorded in Bowern and Lindemann's peer-peer-reviewed survey and on voynich.nu. Sazonov counted prefixes with and without a root in the running text. Label attestation in the running text has been counted here and in the grey literature. No source strips the putative proclitic layer and re-measures the relation between labels and text, and none calibrates the operation on a language with real proclitics. The logic of the test is simple. If the prefixes are clitics, stripping them should bring text words closer to the labels, as it brings Hebrew text words closer to the names. Neither that logic nor the result, that stripping moves the text away from the labels and halves attestation, is in the searched literature. |
| 165 | Agglutinative slot, family and sequence controls with successor-variety signatures | Established | same | Reddy and Knight 2011, What we know about the Voynich manuscript, ACL LaTeCH workshop… | Reddy and Knight published unsupervised affix-signature extraction on the Voynich text, with Linguistica. Lindemann (2022) published the comparison with agglutinative languages on word statistics, and thevoynichproject made it in the grey literature on induced morphemes per word. The first part applies the three measurements of the earlier rows to five more languages, which adds new controls to the same measurement. The second part counts signatures at matched vocabulary with a successor-variety cut instead of Linguistica's minimum-description-length cut, and adds the Markov-2 imitation and the shuffle as nulls. The measurement, affix signatures found by unsupervised segmentation of Voynich words, is the published one. |
| 166 | Dense-abbreviation simulation at Cappelli density | Adapted | same | Bowern and Lindemann 2021, The linguistics of the Voynich manuscript, Annual Review of Linguistics 7:285-308… | Bowern and Lindemann measured real abbreviated texts: the entropy and character-set size of diplomatic against normalised transcriptions, and the abbreviated against the plain Secreta. In the grey literature, Edwards abbreviated Latin by rule and reported mean word length, and thevoynichproject reported entropy, mean length and a merge signature under a subsequent verbose encoding. Here the abbreviation is pushed to a stated density, 177 and 229 marks per 1,000 letters, which is the Cappelli-density target of the audit. It is applied consistently and inconsistently. The statistics at issue are the word-length variance and skew and the repetition counts (rep4 and rep30), which the mean cannot show. Consistent dense abbreviation matches the mean but moves the word-length shape and h2 the wrong way. Inconsistent abbreviation lowers repetition only by raising h2 further. |
| 167 | Mixture and coined-noun plaintexts | New | same | Nearest, none making the measurement: Arutyunov, Borisov, Fedorov, Ivchenko, Kirina-Lilinskaya, Orlov… | Mixed-language readings of the text exist as proposals: Arutyunov and colleagues fitted a two-language mixture from letter frequencies, and macaronic and bilingual readings appear on the forums. The text's repeat rates have been compared with single-language corpora. No source constructs sentence-level, word-level, macaronic or coined-vocabulary mixtures and profiles them against the text on entropy, phrase repetition, adjacent-word information and one-edit density. The result is not in the searched literature. Every mixture raises h2. Only an unnatural word-by-word alternation removes the repeated phrases, and it does so at the cost of the adjacent-word information that the text keeps. |
| 168 | Facsimile line-break re-parse of diplomatic manuscripts | Adapted | same | Vogt 2012, The line as a functional unit in the Voynich manuscript, self-published PDF… | Currier, Feaster, Stafford, and Rozanova and Temerev measured line position in the text alone. Vogt, Gnuchev, and Gaskell and Bowern compared it with natural texts cut at artificial widths. Vogt used 62 characters, Gnuchev wrapped prose, and Gaskell and Bowern a 60-character wrap where original breaks were unavailable. The diplomatic corpora with real manuscript line breaks have been used for entropy but not for line effects. Here manuscripts transcribed at the facsimile level, with allographs and abbreviation signs as symbols, are cut at their actual line and paragraph breaks. They are run through the text's own line-edge and paragraph-first statistics with a shuffle floor. That is the test of the scribal-convention reading (allographs at line edges, elongated first lines) which the artificial-wrap controls cannot make. Verse is included as the one genre with line-edge effects of the text's order. |
| 169 | Two-manuscript scribal calibration of the Currier A and B divergence | Adapted | New | Edwards 2025, Voynich reconsidered: scribes and languages, Medium, 7 February 2025… | Currier, Zandbergen, Lindemann and Bowern, and Parisel established the difference between A and B. The peer-peer-reviewed survey and voynich.nu say that the calibration against known-language variation has not been done. One grey source, Edwards, has since calibrated a glyph-frequency correlation against same-language and different-language document pairs. Here the calibration is against the specific case that the two-regimes reading (regime meaning one of the two writing styles) needs: two hands writing one text, and two collections of one genre. It is made at the diplomatic level, with the normalised form as the edited control, using the R1-A10 divergence and h2 statistics with a split-half null. It finds the letter divergence between A and B within a factor of two of that between two hands, and the h2 difference reproduced inside B alone. |
| 170 | List-genre repetition ladder | Adapted | same | Zandbergen, Voynich MS, Sentences etc., voynich.nu… | Zandbergen raised the list-genre explanation of the missing repeated phrases on voynich.nu. The grey literature tested it with other statistics, adjacent-repeat rates and long-range class information, on one recipe compendium and one Vulgate register. Peer-reviewed work compares the text's word-repeat rate and the 42-metric profiles of herbal texts, without a list genre or a phrase-repeat count. Here the genre is sampled as eleven entry-per-line texts of the period. They are measured at the text's own size in its own page skeleton, with the four-word phrase count that the verdict rows rest on and a word-shuffle chance level. The frequent-word head is the second reading, and positive controls and a window sweep are added. The finding is new: every bare list repeats at least 24 four-word sequences, and the two glossaries that reach the text's zero lack its frequent-word head. |
| 171 | Typological-extreme language profile | Established | same | Bennett 1976, Scientific and Engineering Problem-Solving with the Computer, Prentice-Hall, chapter 4… | Bennett (1976) compared the text's conditional entropy with Polynesian languages, Hawaiian first of all. The peer-peer-reviewed survey, the Lindemann and Bowern preprint (with Maori, Tongan, Samoan, Greenlandic, Inupiak and Nahuatl) and voynich.nu repeated the comparison, with Stallings's caveat about spelling. Here the same languages are read from nineteenth-century scripture in period spelling, which lowers the floor further (Maori 2.52). The instruments of the earlier rows are added: repetition counts, the 14-statistic distance, and the closeness score with known-key controls. The measurement is the published one on new corpora, and only the added legs are new here, so the technique is not new. |
| 172 | Labels against period name lists | Adapted | New | JB 2011, Naked ladies, Computational Attacks on the Voynich Manuscript, 31 May 2011… | The grey literature has compared labels with real name lists. One blog compared them with fifteenth-century female names on length and endings, and another source matched the Lapidario's stone names by structure. The initial-glyph and hapax profiles of the labels are published on voynich.nu and in the survey's remark on nouns. Here six real lists of the period are sampled at the label sets' sizes, with window resampling and whole-list initial concentration. The labels are placed on type, length, initial-letter, attestation, one-edit and numbering statistics together, against both the lists and the running texts of their languages. The finding is not in the searched sources. The labels have a list's type-token and singleton profile, but an initial concentration (54%) and a one-edit density (37%) that no real list shows. |
| 173 | Non-professional writer profile with the Paston letters | Established | Adapted | Lindemann and Bowern 2021, Character entropy in modern and historical texts, arXiv:2010.14697… | Lindemann and Bowern compared a private, non-professional early-modern hand transcribed as written, Napier's casebooks, with the text on conditional entropy. Early-spelling English is among the peer-reviewed comparison corpora on word-level metrics. Here a fifteenth-century family correspondence in original spelling is added. It is profiled on the instruments of the earlier rows: the one-edit family share and the phrase-repeat counts as well as h2. The corpus is new and the added statistics come from the earlier rows, so the measurement is the published one on another text. |
| 174 | Adversarial composite battery | Adapted | same | Greshko 2025, The Naibbe cipher: a substitution cipher that encrypts Latin and Italian as Voynich… | Encoded natural-language plaintexts have been scored against the text before, one encoding family at a time. Greshko scored the Naibbe cipher on seven prose plaintexts against 42 word-level metrics, and Bowern and Gaskell scored 22 manipulations by distance. Rozanova and Temerev measured unit statistics on Naibbe streams, and the earlier rows here scored fixed codes on Caesar. Here the plaintext side is chosen adversarially, from the list-like, low-repetition and low-entropy sources that the earlier items found closest, and crossed with the fixed and Naibbe encodings. The composites are poured into the text's page and line skeleton, so that the line-edge and page statistics can be scored. They are judged by a fixed 23-statistic pass count with Markov and shuffle nulls and a positive control. An oracle union bounds what any composite of these parts could match. The result, 6 of 23 at best and twelve statistics matched by none, is a bound on the whole family, which the earlier single-encoding comparisons could not give. |
| 175 | Segment isomorph scan | Adapted | same | Friedman and Callimahos, Military Cryptanalytics, Part II… | The classical isomorph attack of Friedman and Callimahos is a manual search for pairs of isomorphic sequences in a few messages, used to reconstruct alphabets once a pair is found. Here it is made a census over all 174,000 windows at every length from 12 to 30. The pattern counts and the far-pair counts are compared with Markov and shuffle nulls and with thirteen synthetic Latin ciphers at the text's size. The statistic therefore reports the size of effect each polyalphabetic scheme would leave (70 to 405 pairs) against the text's 0. The excess at 15 to 24 glyphs is attributed by word content to the known word series. Bowern and Lindemann's published argument against polyalphabetic ciphers rests on word recurrence and does not search for disguised repeats. |
| 176 | Mismatch-tolerant near repeats | Adapted | same | Baeza-Yates and Perleberg 1992, Fast and practical approximate string matching, Combinatorial Pattern… | The published near-repeat work on the Voynich text is at word level: edit distance between word types, their co-occurrence within three lines, similar-word networks and the self-citation model. The gapped n-grams of the Zodiac wiki are 2 to 4 symbols long. Here whole windows of 20 and 30 glyphs, across word boundaries, are compared exhaustively at Hamming distance at most k, by the k+1-block pigeonhole of Baeza-Yates and Perleberg. They are counted against Markov and shuffle nulls in three views. The count is calibrated with sparse-homophone (p = 0.05 and 0.10) and verbose-homophone Latin ciphers at the text's size. The target is a plaintext repeat disguised by a few homophones or nulls, which is a different thing from the word families that the word-level measures describe. |
| 177 | Zodiac ring depth test | New | same | Zandbergen, Some special properties of labels in the Voynich MS, voynich.nu, updated 11 September 2025… | Every source found compares zodiac labels across rings under the identity mapping, by duplicate counts, initial-glyph shares or edit distances, or reads them as a numbering without a test. None asks whether a single glyph permutation maps one ring's labels onto another's. Here the comparison is made invariant under relettering, by pattern overlap and pattern kappa at the best cyclic shift. A maximum bijection-consistent pairing is searched for all 66 ring pairs and the star pages, against three nulls (re-dealt, Markov-imitated and borrowed labels). A planted one-list-twelve-keys control at 20% damage (19.2 recovered against 3.65) and a twelve-lists-twelve-keys negative control state the power of the null result. |
| 178 | Same-alphabet chi and kappa tests between partitions of the text | Adapted | same | Kullback 1938, Statistical Methods in Cryptanalysis, revised edition, Register No. 193… | The published comparisons between Currier languages, sections, pages and hands measure frequency differences, correlations, Euclidean bigram distances, n-gram classifiers or a mixture model on substitution ratios. They ask whether the groups differ. The cross-coincidence chi, normalised by both indices of coincidence, asks the cryptanalytic question instead. Are two groups written in one alphabet, which gives a value near 1, or are they reletterings of each other, which gives 0.4 to 0.55? Here it is run on seven partitions, including three defined by markers that no source groups by: the paragraph opener, the line opener and the stretches between gallows. It has half-split baselines, permutation nulls of the minimum, Markov seeds, six keyed Latin controls at the text's size, and a per-page-pair distribution against two-key and five-key ciphers. A direct search for a glyph permutation between A and B is added, which no source found attempts. |
| 179 | Unit-level homophone exchangeability | Adapted | same | Lehofer 2022, Applying hierarchical clustering to homophonic substitution ciphers using historical corpora… | The forum measures whether word pairs that differ in one unit both exist, and how their frequencies compare. That shows alternation but not exchangeability of context. Greshko's Naibbe paper builds homophonic prefix and suffix units without testing the manuscript's own. Here Lehofer's context statistic, validated at glyph level in R4-Q28, is moved to the slot grammar's prefix, core and suffix units in a position-classed spaceless stream. It has a pooled-and-resplit null, two context definitions, and Markov and drift negative controls. The Naibbe ciphertexts are the positive controls, both with the key's own units (AUC 0.93 to 1.00) and with the same inferred segmentation as the text (AUC 0.45 to 0.66). That states the power limit of the null result. |
| 180 | Core-only repeat counts | New | same | Timm 2014, How the Voynich manuscript was created, arXiv:1407.6639… | No source counts repeated word or glyph sequences after the affixes are removed. The published and grey work describes affix variation between similar words, decomposes words into slots, or measures entropy after deleting glyph types. Here two grammars and six position reductions are applied to the text and to its shuffle and Markov imitations, each parsed with its own grammar. They are also applied to Latin with planted nulls at one or both word ends. The full repeat family (rep3, rep4, rep12 to rep30 and the longest repeat) is recounted on each. A fixed cipher of the cores with nulls or state-dependent affixes would then show as the Latin control does. There rep30 goes from 0 to 132 once the ends are dropped. The text's 0 is read at that power. |
| 181 | Key-table tests for the f57v ring, the f49v and f76r glyph columns and the f66r word column | Adapted | same | D'Imperio 1978, The Voynich Manuscript: An Elegant Enigma, NSA Center for Cryptologic History, section 4.3… | D'Imperio and later writers catalogue the four loci and note the f49v cycle and the near-exact f57v repetition qualitatively. Some read the ring as an instrument scale, Brumbaugh used the sequences inside a decipherment without a test, and the forums count f66r words elsewhere informally. Here each locus is put to a stated test with a null. The ring order is tested against 10,000 random orders on three correlates. The ring as an Alberti disk is scored over all shifts against 200 random rings and the Markov imitation. The columns' repeats are tested against 2,000 permutations, and the columns as running keys against 200 key permutations with two statistics. Agreement with line-initial glyphs is counted. The f66r attestation rate is tested against three draws, which turns the forum's counts into a comparison with the label classes. |
| 182 | Solver size and inventory ladder | Adapted | same | Ravi and Knight 2008, Attacking decipherment problems optimally with low-order n-gram models, EMNLP… | The published curves, Ravi and Knight's among them, report a solver's accuracy against cipher length (2 to 10,000 letters) for English or a few languages. For homophonic ciphers they report it against alphabet size (27 to 100 symbols). The Voynich decipherment papers validate at one benchmark size. Here the three solvers of the earlier rows are calibrated on Latin at the manuscript's token counts, 500 to 35,000 tokens and up to 247,760 letters. They are also calibrated on a wrong-genre plaintext and on a 144-symbol inventory whose frequencies copy the manuscript's finest transliteration at its own glyph count. The true key's objective is scored beside the found key's, so that each miss is attributed to the search or to the objective. That turns the question of size and inventory into a stated operating range for the verdict rows. |
| 183 | Drawing-interruption edge test | Established | New | MarcoP 2019, Lines interrupted by drawings, voynich.ninja thread 2945, started 25 September 2019… | MarcoP made the same two measurements on the voynich.ninja forum: the edge-glyph distributions and the glyph association at image breaks, set against line edges and interior boundaries. Here the design differs. A positional null of random interior boundaries in the same lines, and a line-edge reference, place each statistic on a 0-1 edge index. Jensen-Shannon divergence replaces raw shares. A bias-corrected mutual information with a within-page permutation null replaces a summed deviation from independence. There are size-matched in-line and line-break references, per-hand strata (the effect belongs to hand 1, the first of the five scribal hands, with an index of 0.76 against 0.15 for hand 2), and a page bootstrap. Poured Latin and Markov-2 controls in three seeds show that the information test has power. The forum's result, that association at image breaks is as weak as at line breaks, becomes a partial retention of 15-20% against none at line breaks. |
| 184 | Leftover to the drawing outline | New | same | none found | The literature, Zandbergen among others, states the fit qualitatively, or measures the lengths of the words beside drawings in the transcription. Here the space is measured on the scans in glyph units for 1,442 segments and compared with a simulated copyist who fills the same width with the same words. The design is that of the line-end test in R3-P2 (row 90), moved from the paragraph margin to the drawing outline, and so to hand 1. A page bootstrap and a per-bifolium table are added. |
| 185 | Word-gap profile along the line from full-resolution word boxes | New | same | none found | The published gap measurement on the scans compares classes of separator and does not look at position along the line. Whether the line is fitted by spacing or by the choice of word had been argued from transcribed word lengths. Here the physical gaps and word widths of 544 lines are profiled by position, with a within-line permutation null, per hand. They are tied to the measured leftover at the margin, which separates the two mechanisms directly. |
| 186 | Physical-boundary step test | Adapted | same | Parisel 2026, A quantitative confirmation of the Currier language distinction, arXiv:2604.25979… | Parisel types folio transitions by language and quire and compares the jumps in character-pair ratios. The precedents in manuscript studies locate a change of hand or spelling and note that it falls at a sheet or quire boundary. Here the unit is the page, so the turn of a leaf, the centre of a quire, the change of sheet and the change of quire are separated. A folio-level design cannot see the turn of a leaf. The comparison is restricted to boundaries inside one hand and one section. The statistics are the letter-distribution divergence and the vocabulary overlap of matched samples, with a label-permutation null. Seven poured and generated controls are written out into the same pages. The ink change points and the drift generator's step unit are tested against the same typing. |
| 187 | Mixed-quire natural experiment | Adapted | same | Currier 1976 (as G30 | Currier observed that the herbal B bifolia keep their own statistics inside herbal A quires. That observation is the basis of the scribal attribution by bifolium. Parisel's published test compares the jumps between consecutive folios by transition type. Here the twelve bifolia of the mixed quires are compared pairwise, in four classes of hand by quire, on seven diagnostics and on vocabulary overlap. The contrast of hand against quire is given a permutation null that shuffles the hand labels within each quire. So the attribution is tested from the text alone, and the quire is shown not to be the unit. |
| 188 | Bifolium variance of held-out surprisal | Adapted | New | Pelling 2022 (as G33 | Pelling's bifolium-level consistency is a visual map of three glyph-pair densities in one quire. The published per-page statistics (bigram vectors, character-ratio switches) are not decomposed by bifolium. Here a held-out per-token predictability is averaged per page and its variance is partitioned by hand, section and bifolium. The bifolium effect is isolated inside hand-by-section strata with a permutation of pages to bifolia. Three generators and four poured real texts in the same skeleton are scored the same way. That turns the observation into a test with a null and a control set. |
| 189 | Physically-above word test | Adapted | same | Timm 2014, How the Voynich manuscript was created, arXiv:1407.6639 (v3, December 2015)… | Timm counted vertical similarity by line index, as the same position in a previous line or any word one to three lines up. Otherwise it has been seen by eye on a page image. Here the word physically above is found from measured word boxes and set against the same-index and same-fraction candidates on the same records. The test has a page-shuffle null, a fourth candidate from another line pair and a sign test on the discordant subset. A planted 10% physical-copy control fixes the detectable copy rate. The immediate precedent is the index- and fraction-matched test of R1-C6 (rows 28 and 41). |
| 190 | Line-resolution ink runs | New | same | voynich.ninja thread 4377, Ink/pen dynamics and the rhythms of writing, September-October 2024… | No source measures the ink darkness of a manuscript line by line as a series, or tests its persistence against a line-order null. None gates change points by permutation with a false-discovery rate from permuted pages, or asks whether the shifts fall on paragraph starts. The Voynich sources, such as the voynich.ninja thread on ink dynamics, describe the ink variation by eye and explain it by re-inking. The ink models of the adjacent field concern single strokes or ink identity. The nearest earlier measurement is the page-level change-point row R3-P4, with six shifts over 171 pages. The line series refines it to 130 shifts, with a breakdown by hand and a paragraph-start test. |
| 191 | Gregory's-rule side check | Adapted | same | Gregory 1885, Les cahiers des manuscrits grecs, Comptes rendus des seances de l'Academie des Inscriptions et… | In codicology Gregory's rule is applied by reading the sides by eye. On the Voynich manuscript the one remark found is that the sides can barely be told apart. No side sequence has been recorded and no image statistic has been tried. Here the rule is turned into a test on the scans. It uses the recto-verso background difference per leaf and its sign alternation between consecutive leaves of a quire, against the 50% chance level. The ink darkness index is checked for the same asymmetry. That sorts a recto-lighter imaging bias from a parchment-side signal and clears the ink index of it. |
| 192 | Layout regularity | Adapted | same | De Stefano, Fontanella, Maniaci and Scotto di Freca 2011, A method for scribe distinction in medieval… | The absence of ruling in the manuscript is a qualitative observation in the grey literature. De Stefano and colleagues use interlinear spacing as one of several layout features for telling scribes apart on a ruled book. Here the spacing and the baselines are measured as regularity statistics: the pitch coefficient of variation at two resolutions, the within-line baseline slope, page skew and the pitch trend. They are compared with ruled and freehand benchmarks to ask whether the book was ruled. A permutation null within sections tests the spread between hands, and the quire effect is tested within hand-by-section strata. The measurement's own noise floor on a ruled page remains uncalibrated, because no ruled scan was on disk. |
| 193 | Letter-form hand metrics from the aligned glyph crops | Adapted | same | Fagin Davis 2020, How many glyphs and how many scribes? Digital paleography and the Voynich manuscript… | Image-based writer identification with statistical validation is published for other manuscripts. The Voynich hands were assigned by eye, by Fagin Davis, and tested only with text n-grams. The one forum project to put numbers on the letterforms published none. Here palaeographic quantities are measured from the crops of the automatic alignment and averaged per page. They are the height of the gallows and of sh, slant, descender length, stroke width and pen lifts. They are classified by leave-one-page-out against a null that permutes the hand labels within a section. That separates the shape signal from the section signal, which dominates the text classifiers. Hands 2 and 3 remain confounded with section, and hand 5 cannot be tested. |
| 194 | Census of the transcribers' comments and uncertainty codes | Adapted | same | D'Imperio 1978, The Voynich Manuscript: An Elegant Enigma, NSA, section 4.2… | The Voynich sources make the observation by eye: Nill, reported through D'Imperio, and Zandbergen. Pelling estimated an uncorrected-error rate on one paragraph, and the forum disputes the claim. The survey in the adjacent field counts corrections in manuscripts of a known language. Here every correction, insertion and uncertainty mark recorded by transcribers who examined each glyph is counted as a rate per 1,000 letters, by hand, section and locus type. The spread between hands is tested by permuting pages to hands within a section. The marked words are tested for line-edge position, and darkness outliers of the line series are listed for a manual pass. The result is one suspected correction per 9,000 letters in every hand, 20 to 50 times below a copyist's rate, with the marks clustered at the line edges. |
| 195 | Set-off contrast | Adapted | same | Pelling, Voynich codicology, Cipher Mysteries (after The Curse of the Voynich, 2006)… | Pelling read contact transfers in the manuscript by eye, to argue about the original order of the bifolia and the timing of the painting. One forum poster overlaid a flipped page to test a single mark. Document image analysis registers the verso's mirror image in order to remove it. Here the mirror-image transfer is measured as a contrast statistic for 106 pairs, scored against a null of random same-section masks. It is calibrated by show-through as the positive control, with a detection rate of about 40% at 1600 px. It is applied to the three pair classes that the folding and the binding predict to differ: conjugate inner faces, conjugate outer faces and the facing pages of the binding. |
| 196 | Battery-wide generator search | Adapted | same | Timm and Schinner 2020, A possible generating algorithm of the Voynich manuscript, Cryptologia 44(1):1-19… | Published Voynich generators, such as Timm and Schinner's, are tuned by hand or swept over a parameter grid and scored on the whole text. The one evolutionary search on the manuscript optimises a decipherment fitness. Fitting a simulator by matching summary statistics, and checking it by recovering known parameters, are methods of the adjacent field that had not been applied to Voynich generators. Here a global optimiser is run over every free parameter of four families against all 23 battery statistics together. The tables are built from training quires and the optimum is re-scored on held-out quires. A recovery control and a sensitivity scan say which parameters the battery can identify and which it cannot. It can identify the inclusion probabilities, but not the rates of change, the recency copy or the section base. The failure of every family to reach the held-out floor is therefore charged to the families and not to local fitting. The held-out competition of row G6 refitted generators by coordinate descent. This row asks whether a global search changes that answer, and it does not. |
| 197 | Shared-escape prequential class contest | Adapted | same | Dawid 1984, Present position and potential developments: some personal views: statistical theory: the… | Prequential codelength comparison, two-part codes, class n-grams and cache models are standard in the adjacent field. On the Voynich text, the published work induces word classes and measures bigram predictability, and the grey neural experiments score n-gram reproduction or a single model's loss. Here the three accounts of the text (a language with word classes, a self-citing procedure, and a word codebook into a known language) are each given their strongest likelihood-bearing member. All run under one shared spelling model on the same page folds, so that only the sequence prediction differs. An oracle-table codebook control fixes the cost of the codebook class even when the table is known. A class-detection control is run on each account's own kind of text. It shows that the class instrument fails its own positive control at this text size, and that the codebook loss says nothing about hidden content. |
| 198 | Gain anatomy of the character LSTM | Adapted | same | Khandelwal, He, Qi and Jurafsky 2018, Sharp nearby, fuzzy far away: how neural language models use context… | The interpretation tools are published for natural-language models by Khandelwal and colleagues, Hupkes and colleagues and others. They are context ablation by shuffling, retraining on permuted text, and diagnostic classifiers on hidden states. The only Voynich neural interpretation found is a grey saliency map on a small GPT. Here the quantity taken apart is the recurrent model's advantage over the best n-gram. It is split by position in the word and by word length, against four controls on their own models. The real-trained model is cross-scored on three generator outputs and a generator-trained model on the real text, with the per-line difference as a discriminator. Retraining on a word-order shuffle measures how much of the gain needs the order of the words. A word-identity model tests whether that order is held in word identities. Probes with majority and shuffled-label baselines complete the set. Together they locate the gain in the interior glyphs of the current word. It is drawn from the glyphs of the preceding words and from the page's regime, and not from word identities or long words. |
| 199 | Planted-signal power ladder | Established | Adapted | Cohen 1988, Statistical Power Analysis for the Behavioral Sciences, 2nd ed., Lawrence Erlbaum, reissued… | The grey 2026 repositories plant a signal at two or three sizes into one of their own tests (page-level anchors, planted numerals, attribute-label bindings) and report whether it is recovered. Here the same design is run as a ladder of six sizes with ten replicates, over fourteen null results already published in the earlier rows. Each has a planting recipe that copies the text's own dependence law, so that the planted size eps is measured in units of the in-line or adjacent level. The tests keep their own permutation nulls, with matched nulls for the mutual-information and glyph-model tests. An interpolated minimum detectable effect at 80 percent power is given for every null. That turns each null verdict into a bound. Cross-line dependence is ruled out above 0.2 of the in-line level, hierarchical glyph structure above 2 percent of tokens, and a wrong key entry at 1 of 23. It also identifies three nulls with no useful power (MIx3, Q2 and Q18). Against the textbook form of Cohen, power is estimated by planting into the real text with test-specific recipes, and not by an effect-size convention or a parametric simulation. |
| 200 | Half-sample calibration against seed spread | Adapted | same | Politis and Romano 1994, Large sample confidence regions based on subsamples under minimal assumptions… | Politis and Romano validated subsampling intervals by theory, and the literature otherwise validates them by Monte Carlo coverage on a known process. Here each generator is treated as the known process. The text-side half-sample estimator is applied to single realisations and divided by the generator's seed-to-seed standard deviation. That gives a calibration factor per statistic and per generator: 1.05 for a stateless page process, and 0.44 for a process with state that spans pages. A quire block bootstrap is added as an upper scale. The battery's within-tolerance counts are recomputed on five scales, so that the published rule can be labelled conservative or liberal for each generator. The grey Voynich battery of 2026 reports a folio bootstrap and generator seeds side by side without comparing them, and the 2026 preprint reports seed ranges only. |
| 201 | Leave-one-seed-out Mahalanobis joint test | Established | Adapted | Mahalanobis 1936, On the generalised distance in statistics, Proceedings of the National Institute of… | Three grey 2026 repositories place the real text in a Mahalanobis region of generator seeds with a regularised covariance, which is the same measurement. Here the covariance is shrunk by the Ledoit-Wolf estimator over 23 statistics and 50 seeds. The reference distribution is the seeds' own leave-one-out distances, in place of a chi-square or an empirical 95 percent region. The calibration is checked with held-out seeds of the same generator (one rejection in ten) and seeds of another generator (five of five). The test is run for five generators and repeated on a second transliteration. The design matches the leave-one-out goodness-of-fit statistic of the approximate Bayesian computation literature more closely than the textbook Mahalanobis test. |
| 202 | Energy-distance separation of generator clouds | New | same | Szekely and Rizzo 2004, Testing for equal distributions in high dimension, InterStat, November (5), cited… | The energy test is applied in its textbook form, as a two-sample E-statistic with a permutation reference, without change. What is specific is the object: clouds of 23-statistic battery vectors produced by generator seeds and by quire resampling of the text. There are three uses: seed-cloud homogeneity, pairwise generator separation, and separation of the text from each generator. No earlier Voynich use of any multivariate two-sample test between generator output clouds was found. The classification is therefore NEW as a standard method applied to this text for the first time, and not as a new statistic. |
| 203 | Max-T over half-sample vectors | Established | New | Westfall and Young 1993, Resampling-Based Multiple Testing: Examples and Methods for p-Value Adjustment… | Two grey 2026 repositories apply a Westfall-Young max-statistic permutation to their own families of tests on the Voynich text. Under the conservative rule that makes the method's use on this text established. Here the family is the 23-statistic generator battery. The resampling unit is the page half-sample of the real text, in place of a permutation of labels. The output is a family-wise critical |z| of 3.16 against 1.96, applied to every generator's misfit count on two variance scales. The count is re-derived on a second transliteration. The design is the textbook max-T with the battery's own resampling scheme substituted for the permutation scheme. |
| 204 | Multiplicity ledger of harvested p-values | Established | New | Holm 1979, A simple sequentially rejective multiple test procedure, Scandinavian Journal of Statistics… | Several grey 2026 projects apply Bonferroni or Benjamini-Hochberg over their own search space. One of them, the automated exploration of April 2026, corrects over its whole ledger of findings, which is the same measurement in kind. Here the ledger is built mechanically from every saved p-value of a 153-technique examination, and not from nominated findings. The correction is run three ways: Holm within family, Holm across all, and Benjamini-Hochberg. Every demoted effect is named against its threshold. The floor set by permutation resolution, which makes Holm across all unattainable, is stated. The techniques are counted per verdict row, so that each row's evidence base is sized. Against the textbook procedures nothing is changed. |
| 205 | Held-out transliteration protocol | Adapted | same | Nosek, Ebersole, DeHaven and Mellor 2018, The preregistration revolution, PNAS 115(11):2600-2606… | That single statistics hold across transliterations is established: entropy across transcription systems in the peer-reviewed literature, and a locked ten-metric equivalence panel across editions in the grey 2026 battery. Preregistration, as Nosek and colleagues describe it, is a published practice for new data. Here the two are combined, so that a second transliteration of the same manuscript is the confirmatory sample for a generator-comparison battery. The decision rules (per-statistic tolerance, family-wise misfit count, joint leave-one-out p, and the distinction between fitted and predicted statistics) and the sampling scales are written down before the run. The verdict is then re-derived, and not only the statistics compared. The object confirmed is the ledger of generator verdicts, and not the values of the statistics. |
| 206 | Bias-corrected entropy and complexity profile with quire jackknife | Adapted | same | Bennett 1976, Scientific and Engineering Problem-Solving with the Computer, Prentice-Hall (record at | The conditional glyph entropies are established, by Bennett (1976), Stallings (1998), Lindemann and Bowern (2021) and voynich.nu. Gaskell and Bowern (2022) have a compression metric on excerpts, and grey sources report a Miller-Madow-corrected h1 and a bare LZ complexity. The published Voynich figures use plug-in counts at one length without bias treatment, and the sample-size question raised in the forum literature is answered nowhere. Here the same quantities are placed in one profile, all on the same 30 language samples and generator seeds. The profile has four estimators and their disagreement as a function of block length, an excess-entropy curve, a normalised LZ76 ratio and a match-length rate. A delete-one-quire jackknife gives a bias estimate as well as a spread, and a page half-sample interval is added. An estimator self-test at the text's length fixes the usable block length (L at most 4 or 5) and the match-length bias. The finding is that the text has the lowest E(4) of all the languages, and that it is more compressible and more predictable than its own drift generator. |
| 207 | Skeleton model | Adapted | same | Timm and Schinner 2020, A possible generating algorithm of the Voynich manuscript, Cryptologia 44(1):1-19… | Published and grey generators handle line lengths in one of three ways. Timm and Schinner fix them by a character limit. Rozanova and Temerev, and the borrowed skeleton of the earlier rows, copy the manuscript's own line template. The 2026 automated exploration draws words per line independently from the empirical histogram. The model here is hierarchical (section, page width class, paragraph, line class, and the previous line's length) and fitted. It is checked against the real distributions with KS distances and the lag-1 correlation of consecutive line lengths. It is used for a controlled comparison of the same generator seeds in the borrowed and the fitted skeleton. That measures how much of a generator's battery result the borrowed layout is responsible for: two statistics of twenty-three. |
| 208 | Line-length regression with a within-page permutation null | Adapted | New | Vogt 2012, The line as a functional unit in the Voynich manuscript, self-published PDF, 27 November 2012… | The coupling between words per line and word length is a grey measurement. Vogt (2012) inspected the distributions, and the 2026 automated exploration used Pearson correlations per position class against a simulation. Word-length selection near drawings and line edges is measured in a preprint and a workshop paper. Under the conservative rule the fit-to-space component alone would be established. Here the statistic is a held-out R squared decomposition over three nested term groups: mean word length, paragraph position class, and a vocabulary term that cannot encode the count. It is run on the real text against two controls poured into the same skeleton, each line given the real line's number of words. Such a pour leaves a control with no dependence by construction, so the two controls set the floor of the statistic and cannot fail. The vocabulary gain, which asks whether word identities beyond their lengths predict the line's width, is tested against a within-page permutation null in two seeds. No precedent was found for the vocabulary test or for the poured-control design. |
| 209 | Paragraph topic gain against the drift generator with a codebook control | Adapted | same | Sterneck, Polish and Bowern 2021, Topic modeling in the Voynich manuscript, arXiv:2107.02858, and the Yale… | Sterneck and colleagues (2021) ran LDA on the Voynich text, and Reddy and Knight (2011) measured page topicality. Montemurro and Zanette (2013) measured section keyword structure, and the grey 2026 multiscale audit measured held-out cache gains at several scopes. An earlier row here measured within-page cache and cross-page LDA gains. Here the document-completion evaluation of the topic-model literature is moved to the paragraph half. The statistic is the topic gain beyond what the paragraph's own words explain (LDA plus cache, minus cache). The null is a stateful generator without topics, run in the same strata with five seeds, in place of a shuffle. A codebook over English text is planted as the positive control, and it is shown to be detectable only through far recurrence of page-dominant topics. The result is a stated power limit: topics of English strength in the Currier B strata, and nothing weaker. |
| 210 | Verdict statistics re-measured inside each Currier regime against samples of the same length | Adapted | same | Currier 1976, Some important new statistical findings, seminar paper, consulted through the summaries… | Earlier work measured the two regimes to show that they differ. Here each regime is scored on the full verdict battery against reference samples cut to its own length and against imitations of its own glyph statistics, to test whether any verdict drawn from the whole text depends on pooling A and B. None reverses, and the repeated-phrase argument is found to need the whole text because the A pages alone are too short for it. |
| 211 | A hand procedure of the period, simulated and scored on the full battery | Adapted | same | Rugg 2004, An elegant hoax? A possible solution to the Voynich manuscript, Cryptologia 28(1)… | Rugg's table and grille and Timm and Schinner's self-citation are hand methods, and each was shown to reproduce some properties of the text. Here a procedure restricted to devices in use before 1400 (tables, lots or dice, a ruler and the page, with no Cardan grille and no cipher disk) is simulated with every random act counted, its rates fitted, and its output scored on the 23-statistic battery, the extras and the recurrence profile beside the published generators on the same footing. Ablations tie each device to the statistics it buys, two table sets (one set of tables for each writing style) test the Currier regimes, and the cost per word is stated. The cost accounting, the per-device ablation on a common battery and the regime test have no precedent found for this text. The search that followed has none either: some forty families of smaller procedure, the kinds of smaller procedure only, on four independent lines under a bound on the apparatus, the count of cells (table entries) and sides beside every score, held-out quires with the tuned splits named, a page classifier, and the derivation of every word from what was in sight with description length as the judge. Greshko's Naibbe cipher (2025) is the one published key of the period's kind whose sign variants are chosen by chance, and the dice-read tables here are its relatives without a plaintext. |
The techniques that are new under both views.
- Sequential vs hierarchical within-word prediction test. This test measures how much the first glyph of a word tells about its later glyphs once the adjacent glyph is known. A nested category code, in which the first sign names a class and the later signs narrow it down, would keep that information, and a left-to-right template would not. The text keeps none of it, where Latin, Italian and Hebrew keep 0.3 to 0.6 bits.
- Hierarchical category-code generator. This generator tests the idea that the text is a philosophical language of the kind Wilkins proposed. In such a language a word is built by choosing a category, then a subcategory, and so on. The generator arranges the attested words in a branching tree (a trie) and makes a choice at every branch that is biased by frequency rank. It is fitted and scored on the full battery of statistics.
- Codebook/nomenclator generators. These control texts test the idea that each Voynich word is a code group standing for one word of a real text. Every word of a Latin, Italian or Finnish text is replaced by the Voynich word of the same frequency rank. The control therefore has no free parameter. Such texts fail on the statistics of similarity between neighbouring words.
- Segmentation-free repeated-substring test. This test removes all the spaces and counts the strings of 12 to 30 glyphs that occur more than once. The count does not depend on where the word breaks are. A fixed substitution cipher cannot change it, and neither can a verbose cipher, which writes each letter as a group of signs. The text has no repeated string of length 30, where the sixteen core texts have 28 to 3,899.
- Route / columnar transposition test. A transposition cipher keeps the signs and changes their order, for example by writing the message in columns. This test re-reads the glyph stream under nine page geometries (reversed lines, boustrophedon lines that alternate direction, columns, blocks and diagonals). It scores each reading on conditional entropy, which is how much surprise each glyph holds once the glyph before it is known, and on long repeats. Controls with a known key check that the test finds a transposition when there is one.
- Order-invariant phrase-repetition test. This test looks for repeated runs of words after each word is reduced to its bag of letters, with the order inside the word ignored. Anagramming the words cannot hide such repeats. Anagrammed Latin keeps 111 of them, and the text has 1 to 5.
- Codebook / nomenclator simulation grid. A codebook cipher replaces each word of a text with a code group. This grid builds 620 such texts from five plaintexts. It varies the homophones (several code groups for one word), the null groups (code groups that mean nothing) and the split words. Each text is scored together on the word-order information budget, the vocabulary statistics and the near-copy statistics.
- Labelling-free leading-unit distribution test. If the words were measurements written as numbers, their leading digits should follow Benford's law, under which small leading digits are commoner than large ones. This test compares the distribution of the leading units of the words with Benford's law and with a uniform distribution, in which every leading unit is equally common. It found no evidence either way, and nothing rests on it.
- Ordered-table tests. If the text were a table of numbers, neighbouring entries would often share their leading units and would often run in one direction. These tests measure such statistics of adjacent words (shared leading units, monotone runs) without needing to know what any sign means. They were calibrated on generated numeral tables (Roman, positional, ephemeris-like, sorted and accounting-like).
- Image-text association on the herbal pages. This test asks whether the words on a herbal page are linked to its picture. This examination computed descriptors from the scans of the herbal pages (pigment shares, plant shape, roots and flowers) and compared them with the page vocabulary by a Mantel test. That test asks whether pages with similar pictures have similar words. As a positive control, it planted a vocabulary linked to the pictures, which the test had to find.
- Fit-to-space at the line end measured on the scans. This measurement takes the space left at the right margin of 974 line ends on 42 pages from the scans. It compares that with a simulated copyist who fills the same lines with the same words and breaks the line when the next word will not fit. The text ends its lines within half a glyph of the margin and the copyist within one. The text is tighter on 13 of 15 pages with a clean margin.
- Ink change points against vocabulary drift. This examination measured the darkness and hue of the ink page by page from the scans. It then compared the places where the ink changes with the places where the vocabulary shifts. The two coincide no more often than chance would predict.
- Global shape clustering of aligned glyph crops. This examination cut small images of single glyphs from the scans without using a transcription. It grouped them by shape alone and compared the groups with the alphabet of the transcription. The grouping gave 30 clusters with a purity of 45 percent, purity being the share of each cluster that its commonest transcribed glyph makes up. Most of the mixing came from noise in the crops. Confidence that no forum experiment of this kind exists is low.
- Allograph-sequence channel test. Many glyphs have shape variants (allographs). This test asks whether the sequence of variants along the text is more predictable than a shuffled sequence. It allows for the scribe's hand, the page, the neighbouring glyph and the position. As a positive control, this examination replaced the variant labels with the vowels of a Latin text, to check that the test finds a hidden message when there is one.
- Residual-channel tests. A hidden message could live in the part of each word that the best predictive model cannot predict. These tests take that leftover as six streams (surprisal, rank, prefix, suffix, core and generating component). Each stream is tested for sequential structure, compressibility, repeats and position effects, against generators of the same shape. As a positive control, this examination planted a message of one letter per word, which the tests had to find.
- Word and glyph mutual information across line and paragraph boundaries. Across the line break the words share 0.008 bits of information, against 0.129 inside the line, and the glyphs share -0.003. Languages keep a third to two thirds of their in-line value across the break.
- Non-linguistic symbol corpora. The text was placed on a scorecard built for non-linguistic symbol systems. At the glyph level it scores a positional rigidity of 0.249 and an h2 of 2.07. Those values put it among the emblem systems and away from letters in words.
- Genuine historical ciphertexts. The Copiale cipher, a genuine 18th-century ciphertext, was measured as a reference class. Its h2 is 4.81 bits against 2.36 for the text, and it has no repeated run of 20 symbols.
- Word-pattern (isomorph) fingerprint. A word pattern records which positions in a word hold the same symbol, so no substitution cipher can change it. Only 15-16% of the merged-glyph tokens contain a repeated symbol, against 24-72% in alphabetic languages. The nearest scripts are the consonantal ones.
- Glyph-level homophone detection by hierarchical clustering of symbol contexts. The clustering of glyph contexts recovers artificial homophone keys exactly. On the text it finds no exchangeable pair of glyphs. The only candidates are variants that the standard transliteration already merges.
- Menzerath-Altmann law between word length in units and mean unit length. The Menzerath-Altmann law says that words with more units have shorter units. On minimum-description-length units the exponent is -0.59, inside the language range. The shuffled text and every generator reproduce it.
- Unsupervised word segmentation of the spaceless glyph stream and its agreement with the transcribed spaces. Word boundaries were recovered from the glyph stream with the spaces removed, at an F1 of 0.585 against 0.505 for Latin. The shuffled control recovers nothing, so the shape of the words is what marks the spaces.
- Long-context neural language model gain over Kneser-Ney n-grams. A character-level recurrent network was scored against n-gram models on the same folds. It beats the best Kneser-Ney model by 0.129 bits per glyph on the text, by 0.035 on the Markov imitation and by nothing on Latin. The gain lies inside the current word and the few words before it.
- Phonetic-prior decipherment into candidate known languages with a language-closeness score. The closeness score was calibrated with a positive control from the same language family. The text scores 0.05 or less against seven candidate languages. A Latin cipher scores 1.00 against Latin and 0.14 against its relative Italian.
- Clitic-stripping name test. The manuscript's labels are tested as running-text words minus a layer of proclitics. Stripping qo-, o-, y-, d- and s- from the running text moves its initial-glyph distribution away from the labels, by a divergence of 0.155 to 0.441 bits. It also halves the share of label types found in the text. The same operation on the Hebrew Bible brings text words to within 0.03 bits of the names that follow ben (son of), while leaving them attested in the text.
- Mixture and coined-noun plaintexts. Mixtures of Latin with Middle High German or Italian are constructed in four ways. They mix by sentence blocks, by word-level alternation, by replacing rare Latin words with vernacular ones, or by replacing them with coined strings. They are profiled at the manuscript's size. Every mixture raises the conditional letter entropy above both of its sources, to 3.36 to 3.43 bits against 3.11 to 3.30. Only word-by-word alternation removes the repeated four-word phrases, and it does so by destroying the adjacent-word information that the manuscript keeps. Coined vocabulary moves the one-edit family share away from the manuscript's value.
- Zodiac ring depth test. The twelve zodiac rings are tested for being one list of names written under twelve different alphabets. The largest set of labels that a single relettering maps from one ring onto another averages 3.6, against a chance level of 3.4 to 4.0. A planted list damaged by 20 percent gives 19.
- Core-only repeat counts. Repeated sequences are recounted on the word cores alone, after the prefixes and suffixes are removed by two grammars and by position. No reduction yields a repeated 25- or 30-glyph string. Latin with a null glyph at each word end recovers its 132 repeated 30-glyph strings when the ends are dropped.
- Leftover to the drawing outline. The space left before a drawing that interrupts a line is measured on the scans. On the herbal pages of the first hand the median is 0.8 glyph, against 2.7 for a copyist who breaks before the word that does not fit. Half the segments end within one glyph of the drawing.
- Word-gap profile along the line from full-resolution word boxes. Gaps between words and word widths are measured along 544 lines. The last gap narrows by about a tenth of a glyph and does not track the space left. The last word is 5 to 9 percent narrower than the line's mean. So the line end is fitted by the choice of word, and not by squeezing.
- Line-resolution ink runs. The darkness of the ink is measured line by line down each page of the manuscript. It is persistent, with a lag-1 autocorrelation of 0.44, and three quarters of the pages are individually beyond chance. It steps at 0.7 places per page, where a shuffled page steps at 0.02. On the pages aligned to the transcription, the steps fall on the first line of a paragraph three times more often than chance. The ink series records runs of the kind that dips, re-cuts and pauses produce. They are most frequent in hand 2 and rarest in hand 1.
- Energy-distance separation of generator clouds. Whole clouds of generator output, with each seed a vector of twenty-three text statistics, are compared with one another and with quire resamples of the text. The comparison uses the energy distance (Székely and Rizzo 2004) with a permutation reference. Every pair of generators separates at the smallest attainable p, and so does the text against every generator. The two halves of one generator's seeds do not.
The techniques that become new once the unreviewed work is set aside. These twenty-three have a preprint, a forum thread or a student project as their nearest precedent. The peer-reviewed literature has no precedent for them.
- Numbering-system test for labels. If the star and zodiac labels were numbers or names in a fixed order, consecutive labels on a page would resemble each other. They would be more alike than random pairs of labels from the same page. This test makes that comparison. Neighbouring labels are no more alike than random pairs. Forum posts have looked at adjacent labels, but none has used this test.
- Verbose/null-cipher solver. A verbose cipher writes one letter as a group of signs and may add nulls, signs that mean nothing. This solver searches at the same time for which symbols are nulls, which symbol pairs stand for one letter, and which letter each unit is. On Latin ciphers of that kind it recovered 91 to 97 percent of the words across the validation runs. On the text it behaves exactly as it does on the meaningless control. The nearest published work discusses verbose ciphers or builds one. It does not solve one.
- Held-out key-transfer test. This test fits a decipherment key on the Currier A pages and scores it on the Currier B pages, and the other way round. A genuine Latin cipher loses nothing in the transfer. The text loses 0.15 to 0.72 bits per character, which is what a meaningless control loses when its two halves are generated separately. Earlier work argued the same point from letter ratios and did not fit a key.
- Words-as-numerals test via slot dependence. This test measures how much the last two units of a word tell about its first two. The text gives 1.37 bits, where a positional number system gives nothing at all and Roman numerals give 0.2 bits. Generated numeral tables were the positive controls. The idea that the words are numbers had only been pursued by pattern matching before.
- Transcription-free glyph segmentation of full-resolution scans. This examination cut the full-resolution scans into lines, words and glyph-sized pieces without using a transcription. It then matched the pieces to the transcription by dynamic programming, a method that finds the best alignment of two sequences. Every crop then has a label and every line an alignment score. The only earlier work found is a recogniser trained on existing transliterations and described in blog posts.
- Paragraph-similarity network modularity against paragraph-shuffled text. The network of similar paragraphs was measured again with a matched reference set. Its modularity is 0.462 against 0.244 for the shuffle, inside the range of 28 languages in the same paragraph skeleton. The hand and Currier blocks explain 78% of the excess.
- Reading-direction cross-entropy asymmetry Delta_n = X_LTR - X_RTL. The reading-direction test was made symmetric in its treatment of the two directions. Held-out cross-entropy forward minus backward stayed under 0.001 bits per character for all 30 texts. The published figure is 0.065.
- Boundary signatures. The four boundary signatures were measured again with controls. The text scores 3 of 4, and so does its meaningless Markov-2 imitation. The criterion cannot tell the two apart.
- Internality index of uncertain word separators and line breaks. The index measures how strongly the glyphs on either side of a separator are associated. Uncertain spaces keep 0.29-0.43 of the within-word association, certain spaces 0.09-0.14, and line breaks none.
- Physical gap width on the scans as a predictor of transcriber space certainty. The widths of the gaps between words were measured on the scans. Spaces the transcribers marked uncertain are half as wide as certain ones (0.55 against 1.10 glyph widths) and ten times as wide as a gap inside a word.
- Dependence gap across BPE merge levels D_k = H. The dependence curve over merge levels has its minimum of 1.053 bits at 64 merges, which matches the published 1.045. Every language has its minimum at zero merges.
- Quire bootstrap and leave-one-quire-out confidence intervals for the headline statistics. Every headline number was given a quire-level interval. Dropping any single quire leaves h2 at 2.10-2.15 and the repeated 4-grams at 0-1. The spread between quires is 2-3.5 times the page-level noise.
- Paired single-leg gallows on paragraph top lines. The pattern of paired gallows on top lines was tested and is absent. Two or more such words appear on 56.6% of first lines, fewer than the 61.5% that random placement predicts, and they lie left of centre.
- Stroke-level symbol view. The alphabet was recoded into strokes. In that view h2 is 1.68 bits against 1.56 for Latin, and permuting the stroke table moves it by 0.21, so the view depends on the table.
- Two-language mixture test from the deviation of sorted letter frequencies from a logarithmic law, and spectral portraits of two-letter distributions. A mixture model asks whether the pages come from two sources. Its gain of 0.019-0.021 bits is 5-10 times that of any single language. Yet the gain persists inside each hand, and a third source still pays, which a real mixture of two languages does not show.
- Hildegard of Bingen's Lingua Ignota scored as a type list. The one medieval invented vocabulary of known meaning is scored beside the manuscript's word types. It shares the narrow ending system and the uninformative first glyph. But it keeps a language-like glyph entropy, 3.0 bits against 2.1 to 2.5, and few one-edit neighbours, about 20 percent against 50 to 87. So the first-glyph test cannot separate the manuscript from a constructed vocabulary, while the entropy and the neighbour density can.
- Two-manuscript scribal calibration of the Currier A and B divergence. The difference between the manuscript's two writing regimes, Currier A and B, is calibrated against two manuscripts of one text and two charter collections transcribed as written. Two hands copying one text differ in their letter distributions by 0.011 to 0.012 bits, four to seven times their split-half floors. A and B differ by 0.0196 bits, within a factor of two of that. The 0.19-bit difference in conditional entropy between A and B is matched by contiguous halves of B alone.
- Labels against period name lists. The manuscript's labels are set against six real name and word lists of its period, at matched size with resampled windows. The lists are litany saints, martyrology names, a monastic name list, a freemen roll, a plant synonymy and a glossary. The labels have a list's type-token ratio and singleton share. But 54 percent of them begin with one glyph, where no list puts more than 16 percent on its commonest initial. And 37 percent of their types lie one edit from a commoner type, against 7 to 14 percent in the lists. So they have the vocabulary shape of a list and the form of a template.
- Drawing-interruption edge test. The line-start rule and the loss of glyph association are measured where the pen restarts after a drawing inside a line. On the herbal pages of the first hand, the word after the restart is three quarters of the way to a line-initial word. The glyph association across the restart keeps 15 to 20 percent of its in-line value, where a true line break keeps none.
- Bifolium variance of held-out surprisal. Per-page predictability is partitioned by the physical make-up of the book. Inside hand and section strata the bifolium explains 56 percent of the variance of the held-out surprisal per page, against 22 percent for pages shuffled among bifolia. None of three generators in the same page skeleton reproduces the effect. Of four real texts poured into it, the German and the English show 0.44 at p 0.005, the Latin 0.29 and the Italian 0.20.
- Max-T over half-sample vectors. The twenty-three comparisons between the text and each generator are given one family-wise tolerance. It is the maximum absolute z over two hundred page half-samples of the text, a critical value of 3.16 in place of 1.96. That lowers the drift generator's misfit count from twenty to fifteen of twenty-three and leaves the joint rejection intact.
- Multiplicity ledger of harvested p-values. All 445 p-values saved for the text across a 153-technique examination are corrected together, by Holm within and across the five tool families and by Benjamini-Hochberg. Every demoted effect is named against its threshold. The 203 raw discoveries become 185 under the false-discovery correction and 137 under the family-wise one. The resolution floor of permutation p-values is shown to make the strict family-wise count a lower bound.
- Line-length regression with a within-page permutation null. Whether the number of words on a line depends on which words they are is measured by a held-out regression. It predicts the word count from the line's mean word length, its place in the paragraph and a vocabulary term. It is run on the manuscript and on Latin and generator text poured into the same layout, each line given the manuscript's own number of words. Mean word length explains 5 percent of the variance in the manuscript, against under 1 percent in the poured texts, but the poured texts are a null by construction, since the pour fixes the number of words on a line whatever the words are. The paragraph position term explains about 14 percent in every text alike, poured or not. Word identities add a further 1.4 percent over a within-page permutation null. That is a fit-to-space coupling, and no poured text could show one, so this test has had no control that could fail. A pour that fills each line to the manuscript's own width could fail, and it has not been run.
The adapted techniques that matter most. Two of them, the verbose solver and the key-transfer test, count as new once the preprints are set aside.
- Bias-corrected adjacent-word MI with an interior-only shuffle. Rozanova and Temerev's null shuffles every word in the line. The null here keeps the first and last word of each line in place, because edge words are different kinds of words (once-only words are twice as common among them). With a whole-line shuffle the correction flips sign (-0.12 bits) on this data. The edge-fixed version is what separates the real text from every copy and template generator. Whether the whole-line version does the same is not reported in the preprint.
- Verbose/null-cipher solver. Rozanova and Temerev fix the segmentation first, by byte-pair encoding, and substitute afterwards. The solver here searches the segmentation, the nulls and the letters together, with a stated cost for each null, and it reports word-level recovery on verbose ciphers with a known key.
- Many-to-one (homophonic) solver. Rozanova and Temerev, and Kuykendall, fix the number of classes in advance. Here the annealer chooses how many symbols share a letter, under a description-length penalty, and the collapse of the unpenalised objective to 'iiii...' is documented.
- 24 generative models from six families scored against a fixed 23-statistic battery, with fitted/trained/predicted bookkeeping per statistic, page-half-sample CIs and z-scores. Timm and Schinner, Gaskell and Bowern, Parisel, and Rozanova and Temerev all compared a generator with a set of statistics. Here 24 models are scored on one fixed battery, and a ledger records for each statistic whether the model was fitted to it, trained on it or predicts it. Only genuine predictions count.
- Ordered-slot grammar learned by greedy set cover. Stolfi and Zattera built their slot grammars by hand. The grammar here is learned automatically. It is scored by a curve of rule budget against coverage, by the productivity gap and by the shuffled-word selectivity test, and the same scoring is applied to Latin, Italian and Hebrew.
- Held-out key-transfer test. Parisel tested the same idea, that one key cannot serve both A and B, from the switching of character ratios. Fitting a real decipherment key on A, scoring it on B, and comparing the loss with a genuine Latin cipher (0.000) and a Markov control was not found.
- Cipher transformations of Latin. Bowern and Lindemann, and Rozanova and Temerev, enciphered known text and measured h2. The joint criterion, low h2 together with no long repeats, and the position-keyed variant are additions here.
What goes beyond the 2025 and 2026 preprints
Four preprints posted between September 2025 and August 2026 cover part of the same ground. Parisel wrote three of them (arXiv 2509.10573, 2604.19762 and 2604.25979) and Rozanova and Temerev wrote the fourth (arXiv 2608.17096). None has been peer reviewed. This examination therefore treats them as unreviewed claims. It does not treat them as a benchmark. The first count above includes them as prior work and the second sets them aside. Everything said about the preprints here comes from their own texts, and this examination did not run anything taken from them. The table lists, area by area, what they report doing and what this examination does that they did not.
| Area | What the preprints did, by their own account | What this examination adds |
|---|---|---|
| Word-order information | Rozanova and Temerev measured the information between adjacent words as the excess over the mean of 100 shuffles, each of which moves every word of a line, with the vocabulary capped at 2,000 types. They found 0.066 bits. Parisel measured the glyph-level information across the word boundary. | This examination corrects the estimate for bias and compares it with a shuffle that keeps the first and last word of every line in place, so that line-edge effects do not count as word order. That gives 0.113 bits. Splitting the same information by word slot places it at the boundary between words. This examination also measures a held-out bigram gain with an open vocabulary and the character information at distances of 1 to 20 glyphs. It re-measures the preprints' own boundary signatures and merge-level dependence gap with a Markov imitation and a within-word shuffle as controls, which the preprints lack. It measures the information across line and paragraph breaks as well. |
| Substitution-cipher attacks | Rozanova and Temerev ran one calibrated attack. It merged 88 units into 23 classes, used a near one-to-one mapping and seven trigram language models, and scored held-out lines. Parisel ran none. | This examination ran three kinds of solver: one-to-one, many-to-one with a description-length cost, and verbose with nulls, which the Rozanova and Temerev paper puts out of scope. Each solver ran over 21 language models of order 4 and order 5, which predict each letter from the last few letters, and over three symbol views. Each was validated on ciphers with a known key. Every run had a meaningless control. This examination also ran the key-transfer test between Currier A and B, measured dictionary hit rates against the decoded control, and added 19 further corpora with known-key checks in Hebrew, Arabic and Sanskrit. |
| Transposition | Rozanova and Temerev put transposition out of scope. Parisel tested the reading direction only. | This examination ran six tests. They include a solver that anagramming cannot defeat, checked on known keys in five languages, nine re-readings of the glyph stream by route and by column, and a count of repeated letter-bag sequences, which anagramming cannot hide. It also re-ran the reading-direction test with a convention that treats both directions alike, and under that convention the published asymmetry vanishes. |
| Codes and numerals | Rozanova and Temerev used five glyph-level homophonic set-ups, and took the Naibbe cipher, a Linnaeus catalogue and Chinese in pinyin as comparison texts. Parisel did nothing in this area. | This examination built a word-level codebook grid of 620 texts. It ran the words-as-numerals slot test, a leading-unit test against Benford's law and the ordered-table statistics. It also tested syllabic and letter-pair codes of real languages. |
| Hidden channels | Neither preprint. | This examination tested 29 derived channels for a message, with planted messages as positive controls. |
| Text generators | Rozanova and Temerev ran the Timm and Schinner copying generator and an order-3 glyph generator, on a few statistics. Parisel built a slot-based generator with 12 ablations and a Cardan-grille generator. | This examination scored 24 generators from six families on a fixed battery of 23 statistics, with a ledger of what each generator was fitted to. It added a table-and-grille generator over 168 configurations and a wheel generator. It counts the repeated strings without word segmentation, which no generator is fitted to. It runs the first-glyph test, which separates one kind of nested category code from a left-to-right template. It scores a recurrent model against the n-grams on the same folds. |
| Line and page layout | Rozanova and Temerev measured the divergence of line starts with a relabelling null, stratified by Currier language, and measured the enrichment of gallows. Parisel measured positional polarisation and a boundary anomaly. | This examination pours reference texts into the actual skeleton of pages, paragraphs and lines and uses that as the null. It measures the line-edge derivation rates, and the recurrence of words by physical distance with a null stratified by hand and Currier language. It runs a permutation test of the section vocabularies. It tests the claim of paired gallows on top lines against a placement model, and it measures the rightward and downward drift of minimal pairs with nulls. |
| Transcription reliability | Rozanova and Temerev used a second reader, measured agreement on separators, validated the uncertain spaces from glyph coordinates and ran a blind ink audit. Parisel ran cross-transcription checks. | This examination aligns the Zandbergen–Landini and Takahashi files at the level of text unit, word and character, with a dispute rate for every symbol. It recomputes every headline statistic under three symbol views and two word-boundary conventions. It crops the disputed glyphs from the scans. It measures the widths of the uncertain spaces on the scans against blind predictions, and it recovers the spaces from the space-free glyph stream with an unsupervised segmenter. |
| Currier A and B | Parisel built a classifier that predicts held-out folios at 89 percent and found near-categorical switching of four letter pairs. Rozanova and Temerev only used A and B to stratify their measurements. | This examination runs the key-transfer test, which asks whether one key explains both dialects. It classifies pages with one page left out at a time. It measures the differences between A and B against half-splits of real texts. It attributes scribes by character n-grams and by Burrows' Delta and compares the result with the palaeographic hands. It fits a two-source mixture model with single-language and bilingual references in the same page skeleton. |
| Meaning, pictures, labels, embeddings | Neither preprint. | This examination runs a picture-to-word test on 119 herbal pages with a planted control, a function-word test, and label recurrence and numbering tests. It fits word embeddings and aligns them across languages, with a power check on known language pairs. It builds the paragraph-similarity network, the keyword ranking and the information-over-scales curves, with 28 languages and nine generators poured into the same skeleton. It adds gibberish, glossolalia, ciphertext and symbol systems as reference classes. |
| Verification | Rozanova and Temerev resampled by quire, left one quire out at a time and compared disjoint halves. | Separately written code re-derives every headline number. The repetition counts are re-run over 140 contiguous samples, and the repeated 4-grams are recounted after the transcription variants are merged. Every headline statistic gets a leave-one-quire-out interval. A demonstration shows that a quire bootstrap with replacement corrupts the repetition and information statistics. |
The preprints also did some things that this examination did not, so this section is not a superset of them. Rozanova and Temerev ran a blind ink audit and used a catalogue and a pinyin text as comparison texts. Parisel measured the switching of four letter pairs between Currier A and B and built a Cardan-grille generator in Parisel's own form. A Cardan grille is a card with holes in it, laid over a table so that the holes pick out the parts of a word. The table-and-grille generator tested here is Rugg's. Four of the preprints' measurements were re-run here with controls: the validation of uncertain spaces from glyph coordinates, the reading-direction test, the byte-pair dependence gap and the resampling by quire. The section on the data and the sections each measurement bears on say what came of it.
The findings, old and new. About thirty of the results in this report reproduce published ones. Among them is the low glyph predictability, which is the low entropy of a glyph given the one before it. So is its cause, the fixed order of glyphs inside a word. The Zipf-like vocabulary, with a few common words and a long tail of rare ones, was known, and so were the narrow spread of word lengths and Sukhotin's vowel classes. The line was known to be a functional unit, with words behaving differently at its two ends. The absence of repeated phrases was known, and so were the networks of near-copy words. Currier A and B were known to follow the scribal hand. Earlier work had ruled out simple substitution, in which each sign stands for one letter, and had found that verbose encodings, in which one letter becomes several signs, survive. It was known that meaningless generators reproduce most of the single statistics. Rare words were known to come in bursts within a page, and each section was known to have its own vocabulary. The modularity of the paragraph-similarity network, the two-state letter grammar and the long-range correlations between words had all been reported. So had the low glyph-level rigidity on the symbol-system scorecard and the width of the uncertain spaces on the scans.
The search did not find the following results anywhere else. With the spaces removed, the text has no repeated string of 30 glyphs. Once the adjacent glyph is known, the first glyph of a word tells nothing about its later glyphs, which counts against a nested category code. On word-order information measured against the edge-fixed null, a shuffle that leaves the first and last word of each line in place, the real text separates from every meaningless generator. A cipher key fitted on Currier A loses information when it is scored on Currier B, where a Latin cipher loses nothing. Rank-matched codebook texts fail on the statistics of similarity between neighbouring words. The word-order information lies at the boundary between words. The line-edge derivation rates were quantified.
The search did not find these results elsewhere either. One slowly drifting state reproduces the whole multi-scale clustering of the vocabulary. The one-edit variants of words show no genealogy along the codex, that is, no line of descent as the pages go on. The line ends are fitted to the margin. The held-out predictive model leaves 11 bits per word unpredicted and gains 0.02 bits from word order. The Naibbe cipher separates from the text on vocabulary and page structure. Its agreement with the text on entropy and repetition is in the preprint by Rozanova and Temerev, whose Naibbe numbers agree with the ones found here. The argument from missing repeats holds when spelling variation is allowed for. No information crosses a line break. The claim of paired gallows on the top lines of paragraphs fails. The recurrent model gains over the n-gram models, and the gain lies within the word and the few words before it. The calibrated closeness score picks out no candidate language.
The later comparisons add more results that the search did not find elsewhere. The unicity distance of each cipher family, which is the length of ciphertext needed to pin down its key, was set against the length of the text. The Book of Soyga, the Lingua Ignota and the Steganographia were scored beside the manuscript as texts of known status. The first-glyph test does not separate a constructed vocabulary from the text. The generators competed on held-out quire halves, and a page classifier tells every generator from the text. The entropy margin below the nearest languages was measured at a matched sign inventory. The list genres of the period were compared at matched size. There are no isomorphic repeats (two strings with the same pattern of repeated signs) at 25 and 30 glyphs under any relettering. The lines of the first hand fit the drawings by word choice, and the line rule was tested again where the pen restarts after a drawing. The bifolium, a sheet folded once to give four pages, is a unit of the writing on two tests. Within one hand and section the letter distribution steps more at a change of sheet than at the turn of a leaf, and none of seven poured and generated controls in the same pages shows that step. And a bifolium of one hand bound in a quire of another keeps its own hand's statistics, not its quire's. Every generator is rejected jointly, with a calibrated tolerance. The minimum detectable effect of each null test was measured.
Five points differ from published readings. The size of the word-order signal depends on the null model. Rozanova and Temerev find 0.066 bits with a whole-line shuffle, and this examination finds 0.113 bits with the edge words held fixed. The signature of copying is diffuse here, where Timm and Schinner's generator copies locally. Landini and Montemurro read the long-range glyph correlations as linguistic organisation. This examination attributes them to the word template plus repetition. Parisel's reading-direction asymmetry does not survive a convention that treats the two directions alike. Arutyunov and colleagues inferred a mixture of two languages from the letter frequencies. This examination reproduces that gain, but the gain persists inside each scribal hand. So it reads as variation from page to page. It does not read as two languages.