Reading the procedure off the page
A sign is one letter of the manuscript's own alphabet. Was each word copied from one on its page, or spelled out sign by sign? We rewrote each page in a shorthand with three moves. The moves are to copy an earlier word, to copy one with a sign changed, or to spell it out. The shorthand always takes the shortest move. So if most words were copies, it would be far shorter than spelling every word out. But the manuscript's shorthand is only 1.3 per cent shorter than that. Imitations that copy words from the page shorten by 2.7 to 6.8 per cent. So most words were spelled out, with moderate confidence.
The question. The generators above run forward: a procedure is designed, its rates fitted, and its output compared with the text. This test runs the other way. It starts from the manuscript and asks, for each of its 34,780 running words, what the scribe could have been looking at when the word was written, and which is the shortest account of it. Five accounts are allowed. A fresh word is spelled from the text's own habit of which glyph follows which. A copy is taken from a word already on the page: from the line above, from an earlier line, from the word to the left or from the previous page. A copy can have one glyph changed, dropped or added. A word can be made by joining the front of one visible word to the back of another. Or a word can come from a small stock of the commonest words. The judge is description length. Under each account a word costs a certain number of bits to write down. Spelling a word fresh costs its glyphs, about eleven bits on average in this text, while copying a word that stands three lines up costs only a pointer to it. The mix of accounts that makes the whole book shortest is the estimate of what the scribe did, read page by page and hand by hand, with fourteen candidate derivations weighed for every word. The same test was run on texts whose making is known. These were Latin, Italian and English prose poured into the manuscript's own pages and lines, and five of the imitations above, whose real copying rates are known from their code. Three scrambled copies of the manuscript were added, with its words shuffled within each page, its lines shuffled within each page, and its words shuffled across the book. The shuffles say how much of what the test finds comes from the words a page uses and how much from where on the page they stand.
What the page says. The manuscript is spelled far more than it is copied. Allowing every kind of copy from the page shortens its description by 1.3 per cent under the strongest spelling account. Its own words shuffled within each page shorten by the same 1.3 per cent. Every imitation shortens by more, 2.7 to 6.8 per cent, because each of them really does copy words from the line above and vary them. So the words on a page are related to one another almost entirely through the vocabulary the page uses, not through the positions of the words. Three things follow. First, there is no ruler. When a word has an exact copy on the line above, the copy stands within one position of it 37 per cent of the time, against 30 by chance. But shuffling the page's lines leaves 34, so nearly all of the small excess is a habit of position within lines. The imitations that copy the word above show a ratio to chance of 1.13 to 1.52 against the text's 1.24. Their derivations from the line above sit within two positions of the word four to nine times in ten. The manuscript's do so twice in ten and are otherwise spread along the whole line. Second, the variation has a shape. Where a once-only word is one glyph from a page word, the exchange or drop falls at the first glyph in 61 per cent of the exchanges. The pairs are o and y, d and s, k and t, ch and sh, a and o, r and l. A once-only word is one glyph longer than a page word only 10 per cent of the time, exactly the shuffled rate, against 20 to 56 per cent in the imitations. Third, the rare words are spliced or spelled. Of the once-only words, 30 per cent can be made from the front of one page word and the back of another, and the shuffled page gives 29. So the join is a property of the vocabulary and not of the page, and the rest look like fresh spellings under the text's own glyph habits. The edges of lines show the pen habits in derivation terms. A paragraph-first word has an exact copy on its page 4.3 per cent of the time against 23.5 for a word in a random position, because it is a gallows written before a stock word. A line-first word is a letter before a word attested five or more times. The hands differ in the fresh-word rate, not in copying. The share of once-only words is 0.159 on the Currier A pages and 0.124 on the B pages (A and B being the two writing styles Prescott Currier identified). By hand it is 0.155, 0.100, 0.146, 0.208 and 0.187 for hands 1 to 5 (the five scribal hands Lisa Fagin Davis identified), so hand 3 writes new words half again as often as hand 2 within the same regime (writing style). The labels show no excess derivation from their own page's text over a random page's.
The candidates under the same derivation. The derivation was then run, with the same code, on the output of seven candidates of the first line (the first of the four separate searches for a small procedure on the previous page) at their fitted points. Each was regenerated and checked against its family's own record, and each was also shuffled within its pages. The table gives the twelve counts that separate the manuscript from every imitation, with the manuscript, chance and each candidate. Two further rows give the saving that survives the shuffle, which measures the shape of the vocabulary alone.
| Count | Manuscript | Chance | H3 | R8c | C6e | SC2 | G3 | PW6 | H4b | DR |
|---|---|---|---|---|---|---|---|---|---|---|
| Description saved by copying, two glyphs of context | 0.013 | 0.013 | 0.041 | 0.022 | 0.054 | 0.024 | 0.042 | 0.034 | 0.058 | 0.030 |
| Description saved by copying, one glyph of context | 0.046 | 0.040 | 0.096 | 0.059 | 0.111 | 0.058 | 0.068 | 0.092 | 0.118 | 0.057 |
| Copies from the line above within one position, ratio to chance | 1.24 | 1.07 | 1.37 | 1.22 | 1.15 | 1.12 | 1.25 | 1.10 | 1.27 | 1.15 |
| Derivations from the line above within two positions, inner words | 0.196 | 0.071 | 0.838 | 0.444 | 0.665 | 0.619 | 0.861 | 0.264 | 0.714 | 0.096 |
| Once-only words one glyph longer than a page word | 0.100 | 0.102 | 0.300 | 0.175 | 0.447 | 0.198 | 0.077 | 0.134 | 0.401 | 0.082 |
| Once-only words one glyph from any page word | 0.194 | 0.194 | 0.398 | 0.268 | 0.563 | 0.270 | 0.230 | 0.234 | 0.515 | 0.208 |
| Once-only words spliceable from two page words | 0.296 | 0.293 | 0.406 | 0.304 | 0.596 | 0.321 | 0.377 | 0.331 | 0.555 | 0.295 |
| Exact copy on the page, paragraph-first words | 0.043 | 0.235 | 0.268 | 0.055 | 0.143 | 0.061 | 0.097 | 0.091 | 0.149 | 0.042 |
| Once-only share, Currier A pages | 0.159 | 0.159 | 0.130 | 0.141 | 0.129 | 0.171 | 0.151 | 0.110 | 0.144 | 0.129 |
| Once-only share, Currier B pages | 0.124 | 0.124 | 0.124 | 0.134 | 0.116 | 0.140 | 0.124 | 0.096 | 0.127 | 0.116 |
| Once-only share, hand 2 | 0.100 | 0.100 | 0.125 | 0.135 | 0.118 | 0.145 | 0.123 | 0.096 | 0.128 | 0.102 |
| Once-only share, hand 3 | 0.146 | 0.146 | 0.126 | 0.135 | 0.115 | 0.135 | 0.129 | 0.100 | 0.127 | 0.130 |
| Saving that survives shuffling within pages, two glyphs of context | 0.013 | . | 0.026 | 0.017 | 0.034 | 0.018 | 0.025 | 0.029 | 0.039 | 0.021 |
| Saving that survives shuffling within pages, one glyph of context | 0.040 | . | 0.077 | 0.050 | 0.084 | 0.048 | 0.047 | 0.084 | 0.094 | 0.046 |
Chance is the mean of three copies of the manuscript with the words shuffled within each page. The tolerance for each count is three times the seed-to-seed spread of the multi-seed runs (a seed being the starting point of the random numbers). By that rule every candidate of the first line misses nine or ten of the twelve counts, and the misses sort into four faults, each tied to a mechanism. The last column, DR, is the recipe read off this table and built as a candidate, described at the end of this subsection. It misses five. The first is the ruler. Only the pointer walk, which never copies from the line above, is within tolerance on the within-two-positions share and on the positional excess, the part of the saving that shuffling removes. R8c and SC2 pass the alignment ratio yet fail the share, at 0.44 and 0.62, because a window along the line above keeps every copy aligned. The second is adding a glyph. Only the glyph chain, whose rare words are spelled glyph by glyph with no add rules, is below the text on the one-glyph-longer count, at 0.077. The copying frames sit at 0.447 and 0.401. The third is too much copying. Every candidate saves more than the text at both orders, with R8c and SC2 closest at 0.022 and 0.024 against 0.013. The part of the saving that survives shuffling, the vocabulary's shape, is 0.017 to 0.039 for all of them against 0.013. The fourth is one fresh-word rate for both hands. SC2 and the glyph chain reproduce the difference between the A and B pages, but no candidate makes hand 3 fresher than hand 2, the one fault no mechanism on any line touches. The four-figure summary below puts the rows of every line on one footing, with the second line's best row scored by its own run of the same code.
| Text | Saving by copying | Share of words spelled fresh | Bits a word to spell every word fresh | Once-only words one glyph from a word of their page |
|---|---|---|---|---|
| Manuscript | 0.013 | 0.820 | 11.11 | 0.194 |
| Manuscript, words shuffled within each page | 0.013 | 0.861 | 11.11 | 0.196 |
| H3, the procedure above | 0.041 | 0.672 | 11.17 | 0.398 |
| R8c | 0.022 | 0.786 | 11.21 | 0.268 |
| R5 | 0.026 | 0.737 | 11.24 | 0.293 |
| SC2 | 0.024 | 0.799 | 11.37 | 0.270 |
| G3 | 0.042 | 0.740 | 11.23 | 0.230 |
| PW6 | 0.034 | 0.645 | 11.14 | 0.234 |
| C6e | 0.054 | 0.646 | 11.02 | 0.563 |
| C5 | 0.068 | 0.622 | 11.02 | 0.661 |
| H4b | 0.058 | 0.625 | 11.31 | 0.515 |
| Hybrid J3 | 0.041 | . | . | 0.259 |
| refined_s4 (second line) | 0.054 | 0.653 | 11.13 | 0.323 |
| p2best (second line) | 0.041 | 0.689 | 11.01 | 0.361 |
| throw_s18start (second line, its own run) | 0.076 | 0.54 | 11.67 | 0.36 |
| refined_s18 (second line, its own run) | 0.059 | 0.61 | 11.08 | 0.314 |
| refined_s19 (second line, its own run) | 0.077 | 0.55 | 11.32 | 0.362 |
| ho4_fc1 (second line of the search, tuned on the splits, the divisions of the quires into a tuning half and a scoring half) | 0.029 | 0.68 | 11.69 | 0.260 |
| lo1_final_tail (second line, split 1 left out) | 0.024 | 0.71 | 11.44 | 0.198 |
| v1b_tail (second line, clean protocol: tuned on other halves and reported on this one, split 1 reported) | 0.032 | 0.67 | 11.46 | 0.260 |
| v2_tail (second line, clean protocol, split 2 reported) | 0.034 | 0.67 | 11.71 | 0.270 |
| v3_tail (second line, clean protocol, split 3 reported) | 0.026 | 0.72 | 11.59 | 0.195 |
| v4_tail (second line, clean protocol, split 4 reported) | 0.032 | 0.66 | 11.39 | 0.287 |
| v5b_tail (second line, clean protocol, split 5 reported) | 0.026 | 0.71 | 11.41 | 0.188 |
| DR, the recipe read off the page | 0.030 | 0.781 | 11.08 | 0.208 |
The saving is the share of the description length that copying from the page removes, under a glyph model with two glyphs of context. The fresh share is the share of words the shortest account spells fresh. The third column is the cost of spelling every word fresh, which measures how regular the words are. The last is the share of once-only words that stand one glyph from a word written earlier on their own page, where the chance rate is the same page shuffled, and it is by the same definition for every row. Of the rows at the floor (a row being one candidate procedure, and the floor the distance between the text's own two halves), throw_s18start is the least like the page on three of the four figures. It copies most and spells least, and its words are less regular than the text's, at 11.67 bits a word against 11.11. That is the "pointing beats spelling" difference in numbers. The whole-text row refined_s18 spells its words as regularly as the text, at 11.08 bits a word, but it still saves four times as much by copying and spells fresh 61 per cent of its words against the text's 82. The recipe read off the page is nearer the text than R8c on the regularity of its words and on the once-only words' relation to the page, and further from it on the copying and the fresh share. The second line's tuned rows, which carry the new-word step, bring the once-only words' relation to the page to the text's value, 0.202 against 0.194 for the row with a split left out. They still copy twice as much as the text, and they spell fresh 68 per cent of their words against 82.
The step that makes the once-only words. About one word in seven in the manuscript, 13.6 per cent, occurs nowhere else in the book, and these are the long words. A word of seven glyphs is once-only 43 per cent of the time, a word of eight 75 per cent, a word of nine or more 95 per cent. Derived from what was in sight, they are not made from the page at all. A once-only word is one glyph from a word already on its page 19 per cent of the time, exactly the rate of the same page shuffled. It can be assembled from the front of one page word and the back of another 30 per cent of the time, against 29 shuffled. It is two whole page words joined 5 per cent of the time, against 6. What the once-only words are related to is the book's vocabulary as a whole. That is what any well-formed word of this text looks like, because the spelling habit is so regular. A glyph chain trained on the repeated words alone produces new words that can be cut into two attested words at the same rate. Two details fix the shape of the step. The length is settled before the glyphs are. The two halves into which a once-only word can be cut are strongly and negatively correlated in length, so the whole comes out six or seven glyphs long far more consistently than two independently drawn parts would. And the new word is not two whole words glued together. Glyph pairs that are common across a word boundary but rare inside a word appear in 46 per cent of the once-only words of seven or more glyphs. Joining two whole words at random would put one in 94 per cent, and the text's own repeated long words carry one in 12. The mean length of the once-only words is 5.83 glyphs against 4.10 for the rest, and every imitation has the same, so the count and the length are not what separates them; the relation to the page is. Written as a step, the scribe decides by lot at each word whether it is to be a new word. The rate depends on the hand and the position: 0.48 for a paragraph-first word, 0.17 for another line-first word, 0.11 inside the line and 0.23 at the line end. Over all positions it is 0.155, 0.100 and 0.146 for hands 1, 2 and 3. The scribe then draws the length from a table for the hand's regime and fills it in one of two ways that consult nothing on the page. One way spells it glyph by glyph from the succession table the procedure already uses. The other draws a front and a back from two lists of word parts, the front shorter than the length and the back of exactly the remaining length. On held-out quires (the gatherings of folded sheets that the tables were not built from) both ways pass on the count and the length of the new words. The second is closer on the two measures that separate them, the boundary pairs and the distance from existing words. The two ways also combine, and the combination is the step as it is specified. The long new words, of seven or more glyphs, are made the second way, from a front and a back. Of the short ones, half are made that way and half by exchanging one glyph of a word from the stock, the 894 words that occur five times or more. The exchange falls at the first glyph more often than at any other single position. It comes from a table of 25 exchanges such as y for s, e for sh or y for d. Spelling a new word from the succession table alone does not serve. A table built from the stock makes a word new to the book only 4 per cent of the time, and those words recur too often and sit too close to the stock. The exchange supplies the recurrence of the short new words and their closeness to the stock, and the join supplies their diversity and their boundary pairs. A mixture between one half and seven tenths of joins matches every measure within about 0.06. The rate of the step by hand and by length, and the sizes of its lists, 606 fronts, 664 backs, the 25 exchanges and the length table, under 2,500 entries in all, are given in the derivation's own file. Halves drawn independently of each other give twice the length spread, which is the failure every candidate met. This step is stated here from the page evidence; the candidates built with parts of it are described in the subsection above.
The recipe, tested. The four faults name the mechanisms to remove, and the step above names the one to add. Put together, they make a candidate. Its frame is the copying frame with no ruler and one lot a word. Variation is an exchange or drop at the first glyph in place of an added glyph. New words have the length thrown first and are spelled to it from the glyph rows. The paragraph gallows and the line letter are written at the text's rates, and the fresh-word rate is set by hand. It is the row DR of the table in the subsection above. On the whole text it reaches 12 statistics within tolerance and 22 within three at 2,983 cells (table entries) and 4.75 acts a word. Under this derivation it misses five of the twelve counts, where the candidates above miss nine or ten. The ruler fault is gone. Its derivations from the line above stand within two positions of the word 10 per cent of the time against the text's 20, and its alignment ratio is 1.15 against 1.24. Its paragraph-first words have an exact copy on the page at the text's rate, 0.042 against 0.043. Its once-only words are spliceable from two page words at 0.295 against 0.296 and one glyph from a page word at 0.208 against 0.194, and hand 2 writes new words at the text's rate. What remains is the copying, at 0.030 against 0.013 saved, twice the text, and the one-glyph-longer count at 0.082 against 0.100, below the text where every earlier candidate was above it. The hands differ in the right direction for the first time, hand 3 fresher than hand 2 by 0.028 against the text's 0.046. But the A pages and hand 3 are still short of new words, at 0.129 against 0.159 and 0.130 against 0.146, because the rate by hand was applied to the sheet alone. Its join share is 0.43 against the text's 0.51. Held out by quire it scores 1.25 against the floor of 1.29, with its rates fixed and nothing tuned on any split, ranked first of nine generators in every split. The share of once-used words is its one statistic beyond three tolerances, on four of the five splits, at three to six short, and the line-end statistic is the next miss. The page classifier (the program that learns to tell real pages from generated ones) tells its pages from the real ones 82.5 per cent of the time. One thing about the fitted row has to be said. Its first page is the manuscript's own first page, copied as the working sheet's first filling, and every later page is written by the recipe. Written by the recipe as well, the first page moves the whole-text score from 0.93 to 0.91, and the held-out runs write the held-out half's first page like every other. The companion page Writing Voynich by the Recipe Read off the Page lets a reader throw the lot word by word through the recipe's nine steps and watch a page take shape. The next rung, a fresh-word rate by hand on every source of new words and not on the sheet alone, held out at 1.26 and closed no gap on the whole text. So the recipe removes the fault that every earlier candidate shared and halves the misses on the page counts, and the copying and the fresh-word rates by hand are what is left to fix.
What it means. The verdict on the imitations is firm on two points and moderate on the shares. There is no ruler. Copying from the line above by position is not in the page evidence, and the candidates that use one, the procedure above among them, put the source beside the word two to four times as often as the text. Variation is an exchange or a drop at the start of the word, not a glyph added. The text copies less than any imitation built so far. Its rare words are spliced from the vocabulary or spelled fresh, with their length settled first. The two hands of the B regime differ in how often they write a new word, not in how much they copy. So an account of the hands that gives them different table sets is the wrong shape, and an account that gives them different fresh-word rates is the right one. The exact split between spelled and copied words depends on a bit or two per word. It is given under two spelling accounts, a table of which glyph follows which, which a scribe could own, and a model with two glyphs of context. The conclusions rest on the model-free counts, which the shuffles and the seeds reproduce.
Earlier work. Timm and Schinner (2020) proposed that the text was written by copying words already on the page and altering them, and their account is what this derivation tests. The page evidence gives copying a smaller part than any form of that account, and it finds the alterations of a different shape. Rugg (2004) proposed tables read through a grille, which leave no trace on the page for this test to find. Zandbergen (n.d.) and Timm (2014) described the near-copies between neighbouring words that the copying account was built to explain. Those near-copies are confirmed here as a property of the vocabulary that shows on the line and hardly at all down the page. New here is the derivation of every word from what was in sight with description length as the judge, with the controls (texts of known origin run through the same test) of shuffled and poured texts and of imitations with known rates. New too are the finding that there is no ruler, the shape of the variation, the fresh-word rate by hand, and the step that makes the once-only words with its held-out check. So is the recipe read off the table, built and scored as a candidate.