What is and is not established
Which findings about the text are established, and how firmly? Not established: whether the text has meaning. That stays undecided. Not established either: that a hand procedure with tables of words and dice wrote the text. That is the best guess, not confirmed. A sign is one letter of the manuscript's own alphabet. Established with high confidence: phrases scarcely repeat. Also with high confidence: the next sign is easier to guess than in thirty-one languages. Established with moderate confidence: some words, a minority, were copied from a nearby word with one sign changed. Most were spelled out fresh.
| Claim | Basis | Confidence |
|---|---|---|
| The stream of glyphs (the signs of the script) is far more predictable than any language sampled, and the extra predictability lies in the order of the glyphs inside words | Conditional entropy, the bits of surprise in the next glyph once the previous one is known: 2.11 in the text in its two EVA transcriptions (2.107 and 2.112), and 2.53 in the finer v101 alphabet, against 3.03–3.61 in the sixteen core reference texts. Across all forty texts of the verdict table the range is 2.52 to 3.69. Shuffling the glyphs inside each word removes the gap | high |
| Each word fits a small template of slots in a fixed order, and each glyph depends only on its neighbour. The first glyph of a word adds at most 0.05 bits about the later glyphs beyond what its neighbour already tells | A grammar of 60 rules covers 90% of the words. The information the first glyph adds is at most 0.05 bits held out (measured on pages the estimate was not fitted to) and 0.12 by a plug-in estimate (counted from the whole text). A category code, a made-up language in which each word spells out a class and its sub-classes, built from the text's own statistics, scores 0.08 held out. Languages score 0.3–0.6. Read from the other end of the word, the last glyph adds as little about the earlier glyphs. That tells against nested codes with unequal, re-used sub-classes, though not against a complete product code, in which every class has the same sub-classes. The last basis is the density of the edit network, the web that joins words differing by one glyph | high for the template; moderate for neighbour-only dependence, which rests on a margin of a few hundredths of a bit over one weak control (a text of known origin run through the same test) |
| The stock of words comes back the way a language's does: the Zipf slope (how fast word frequencies fall away from the commonest word), the Heaps growth (how the count of distinct words grows with the text) and the share of words used once are all inside the language range | comparisons with 35 texts cut to the same size | high |
| Runs of several words almost never come back, whether the spaces between words are counted or ignored | The text has 1 repeated 4-gram (a run of four words), or at most 45 with every disputed glyph distinction merged. The 140 language samples have ≥ 67, or ≥ 38 merged the same way. With the spaces removed, the text has 0 repeated 30-glyph strings, and 131 of the 140 samples have ≥ 28 (four samples of Dante have 0; five of Montaigne, Cervantes, modern Greek and Welsh have 4 to 26) | high |
| A word is a near-copy of its neighbour, or of the word above it, more often than chance allows. Every prose reference text shows the opposite | edit distance (the number of single-glyph changes that turn one word into another) and repeat rates, compared with the same page shuffled | high for the observation; medium for reading it as copying |
| The choice of glyph depends on where the word falls in the line and the paragraph. The effect reaches one word deep at each edge, and the edge forms are made from interior words | Jensen–Shannon divergence, a measure of how far two distributions of glyphs differ: 0.181 / 0.082 for the two edges, against ≤ 0.002 when a language text is broken into lines at arbitrary points. Tests that strip the edge words away, or swap them | high for the effect; medium for the derivation of the edge forms |
| Currier A and B, the two writing styles Prescott Currier identified, are a two-way split in the glyph statistics, and the split follows the scribal hand and not the illustrations | 99% of pages classified correctly from their glyphs; the table of scribal hand against regime (the name used here for the A or B style) | high |
| No fixed cipher or code of any sampled language produced the text, where fixed means that a unit of the original is always written the same way, no matter what stands around it | the invariants above, the properties that relettering cannot change; the repetition of phrases and of glyph strings; the solvers, code-breaking programs that were first checked on ciphers with known keys | high for the sampled languages. The argument from repetition reaches further, to any original that repeats phrases as every sampled text does, under any fixed mapping with fixed unit boundaries, but it does not reach encodings whose spelling varies by chance or with an internal state. Among those, the published Naibbe cipher (Greshko's card cipher, which encodes Latin by drawing cards) is ruled out on other statistics, and its family is unlikely. The solvers' silence tells nothing there, since they score Naibbe ciphertext the way they score the meaningless control |
| Whether the text has meaning | These measurements cannot decide it, because the surviving explanations include generators (programs that write text by a set procedure) with meaning and generators without, and the content tests, which detect content of particular kinds and sizes, found none of those | not established |
| Which of the surviving processes produced the text | see the generative models, the search for a smaller procedure, the derivation from the page and the drifting-state generator | not established |
Four caveats limit everything above. The first concerns the reference set, which is forty texts in thirty-one languages. It includes no syllabic or logographic script (one sign per syllable, or one per word) as written, because Chinese and Japanese had no usable romanised source, so Sanskrit rendered in aksharas, its syllable signs, stands in for a syllabary. The second is that the word-level statistics depend on genre, as the Bible and the Kalevala show, and the reference texts are mostly running prose and verse. The exceptions are the eleven list-like texts of the period: glossaries, martyrologies, litanies, freemen rolls, a plant synonymy and Pliny. Where they are long enough to sample at the text's size, each of them repeats four-word sequences at least 24 times. So the verdicts that rest on phrase repetition do reach the genres with one entry per line, as far as the sample goes, and they assume only that the original was not a list unlike any of those. Even so, the range of genres is narrower than the count of languages suggests, since nine of the nineteen further texts are single scriptures in translation. The third caveat is that the reference set holds no transcribed fifteenth-century Latin or Romance manuscript, so a simulation stands in for the scribal abbreviation of those languages and the measurements include no real sample of it. The fourth is that the two transcribers read about 2% of the characters differently, which is why nothing here rests on the a/o, r/s or stroke-count distinctions between glyphs.