The search for a smaller procedure
Lot books were books of answers picked by dice. One of the time held about a thousand answers. The procedure's large tables hold 18,000 entries. Can far smaller tables match the manuscript? On the whole manuscript they can. One procedure has 2,870 entries, each a word or word part. It matches all 23 measurements. But we fitted its rates to the whole manuscript. So that match is partly what it was made to give. A harder test fills the tables from half the manuscript only and asks the procedure to write the other half. There every small procedure writes too few words that appear only once, where the manuscript has many. So the best guess stays not confirmed.
Walk-through pages, on which the reader can run the procedures described here: ‘Writing Voynich by Hand’, ‘Writing Voynich by the Recipe Read off the Page’, ‘Four Dice, One Tablet’.
The question and the bound. The procedure above shows that tables, lots, a ruler and a working sheet can write a text with most of the manuscript's properties. It also shows the price, about 18,000 table cells, which is more than a scribe is likely to have written out. The question that follows is whether a much smaller set of tables and devices can do the same work. A bound was set before the search began. The apparatus may hold a few thousand cells at most, about ten page sides at 300 cells a side. Each word may cost about four random acts, a throw of a die or a draw of a lot. And every step has to be one that a scribe of the years 1404 to 1438 could carry out. The bound is not arbitrary. Johannes Bolte's 1903 survey of the lot books gives the sizes of the books of the period whose text was chosen by a wheel or by dice. Checked against the editions and catalogues, those books run from 784 to 1,296 rhymed answers on 28 to 72 page sides, at about 32 lines a page. The Vienna book of about 1370 holds 1,296 couplets on 36 sides. The Heidelberg book of about 1430 holds 1,024 on 32 sides. So a candidate with about a thousand entries on seven to ten sides is the size of an ordinary lot book of the decades before the manuscript. The 27,000 cells of the largest candidate tried here are twenty times any attested book. The working sheet has a period form too. Thomas Wozniak's 2025 study of wax tablets gives their capacity as 3.4 to 4.9 characters a square centimetre, so a sheet of 30 words fits one leaf of a school tablet. The cycle of drafting in wax, copying out and wiping is documented for a note-taker at Siena in 1427. That a scribe used a tablet for a word table is an inference from those sizes, not a documented practice. No slate with chalk is documented in Italy or the Empire before 1450.
The families tried. The search ran on four independent lines with the same battery (the set of 23 statistics) and the same five-seed scoring (a seed being the starting point of the random numbers, so five seeds are five independent runs) in the manuscript's own page and line skeleton. Every row (one candidate procedure, with its scores on one line of the table) carries the same count of cells, entries, page sides and random acts. One line designed families by hand and fitted their rates. It tried the two table sets with lots of the procedure above and their repairs, lot books of word classes read by throws, and marked wheels of word fragments. It tried a chancery nomenclator applied to Italian, a hand-sized letter-permutation table of the zairja kind, syllable tables, and abbreviated Latin in an invented alphabet. It tried copying with variation as the whole procedure, knucklebones and dice tables, a rhyme table walked with a finger, linked rows of word parts, hybrids of these, a glyph-by-glyph chain and a finger walking a stock list by dice. A second line worked the same way on families of its own. Those were layout copying, formula frames, word columns, a stencil grid, and linked part tables with a working sheet. They also included a search over every device together with no size limit, the same held under the bound, and lot-book rows of sixteen entries read by three dice. The rest were a wheel-and-grid device, stems and endings, geomantic figures, a word mill, a paradigm table, a repertory of pieces, and an Ogam-tract transformation of Latin. A third line searched automatically over programs. A program is an ordered list of rules of the form "in this situation, this often, do this", where the situations are things a scribe sees without counting and the actions are the steps every family uses. Twelve mechanisms drawn from the practices of the period were among the families: the lot books, the letter arts of Lull's wheels, the geomantic tables and the copyists' name lists. A fourth line went the other way and asked for the smallest device that explains most of the measurements, a single sheet of glyph groups read by one die, and built a ladder upward from it. Last, the recipe read off the page in the next subsection (the procedure derived word by word from the manuscript's own pages) was built and scored as a candidate of its own. The table below gives each family's best row, and the chart below sets each family's best row against the size of the tables it needs.
Every procedure family scored against the size of its tables
One row for each family's best five-seed row, sorted by the number of table cells its apparatus needs, on a logarithmic scale. The number beside each mark is how many of the 23 battery statistics the row reproduces within three tolerances (a tolerance being the range a statistic wanders over between random halves of the text's pages). The dashed line is the bound of 3,000 cells, ten page sides at 300 cells a side. Colour is the outcome band, not the family.
| Procedure | What the scribe does | Within tolerance | Within three tolerances | Cells | Entries | Sides | Acts a word |
|---|---|---|---|---|---|---|---|
| Tables and lots with a die on the ruler and three redraw rules (R9Y) | Two table sets read by lots, boundary rows, a rule table, the four pen habits, a splice; the die sends the copy to the word directly above half the time; the paragraph base, a varied q-word and the last word of a line are drawn again under stated conditions | 8 | 23 | 21,305 | 6,231 | 71.0 | 6.08 |
| Tables and lots with a die on the ruler (R9A) | The same without the redraw rules | 13 | 23 | 21,305 | 6,231 | 71.0 | . |
| Tables and lots with a check grid and a three-line ruler (R5) | The procedure above with a grid of allowed glyph pairs, 600-cell line-end rows, mixing between the table sets, a length guard and a splice | 12 | 23 | . | . | . | 5.32 |
| Linked part tables under the bound (interim_s11) | A beginning, a middle from its row, an ending from the middle's row, a stock of 200 words, 30 rules, a sheet, page joins capped at eight glyphs | 10 | 23 | 2,825 | 1,034 | 9.4 | 7.65 |
| Copying frame with glyph rows (hybrid J3) | Two seed tables, glyph rows as the source of fresh words, no rule row, no ruler, the word just written varied three times in a hundred, a stop die from the seventh glyph | 10 | 23 | 2,909 | 1,359 | 9.7 | 9.65 |
| Linked part tables with one lot a word, searched on the whole text (refined_s18) | The one-lot-a-word tables searched again on the whole text and no held-out split, the best of the top fifteen points re-scored on five seeds; over the bound by this report's count | 14 | 23 | 3,402 | 1,518 | 11.3 | 3.37 |
| The same, a second search point (refined_s19) | A second point of the same search, with the recurrence profile six of six on the whole text; the whole-text match under the bound | 13 | 23 | 2,870 | 1,002 | 9.6 | 3.33 |
| Linked part tables with the length drawn first and the new-word step, tuned on the quire splits (ho4_fc1) | throw_s18start's tables with a length row, lists of fronts and backs and an exchange table; a built word's length thrown first half the time; a new word made from a front and a back at the text's rates, or a stock word with one glyph exchanged; the rates tuned on the five quire splits in rotation | 12 | 23 | 2,944 | 1,263 | 9.8 | 3.68 |
| The same, tuned with one split left out (lo1_final_tail) | The same tables tuned under the scaled counts on four of the five splits, split 1 never used, the join tails written as one row for each hand; the point at the end of the search | 12 | 23 | 2,968 | 1,310 | 9.9 | 3.85 |
| The same under the clean protocol (tuned on some halves, reported on a half never seen), split 1 reported (v1b_tail) | Tuned on splits 3, 4 and 5 and stopped on split 2, the reported point of the search; split 1 never seen | 9 | 23 | 2,988 | 1,228 | 10.0 | 3.51 |
| The same under the clean protocol, split 2 reported (v2_tail) | Tuned on splits 1, 3 and 4 and stopped on split 5, the reported point of the search; split 2 never seen | 10 | 21 | 2,950 | 1,203 | 9.8 | 3.95 |
| The same under the clean protocol, split 3 reported (v3_tail) | Tuned on splits 1, 2 and 5 and stopped on split 4, the reported point of the search; split 3 never seen | 11 | 22 | 2,900 | 1,317 | 9.7 | 3.55 |
| The same under the clean protocol, split 4 reported (v4_tail) | Tuned on splits 1, 2 and 5 and stopped on split 3, the reported point of the search; split 4 never seen | 12 | 23 | 3,000 | 1,254 | 10.0 | 3.66 |
| The same under the clean protocol, split 5 reported (v5b_tail) | Tuned on splits 1, 2 and 3 and stopped on split 4, the reported point of the search; split 5 never seen | 12 | 23 | 2,872 | 1,234 | 9.6 | 3.70 |
| Linked part tables with one lot a word (throw_s18start) | The same tables with one throw a word, the next cell taken instead of a re-throw, beginning and ending rows shared by the two hands, the paragraph letter at the text's rate | 13 | 22 | 2,100 | 899 | 7.0 | 3.24 |
| The recipe read off the page (DR) | The copying frame with no ruler, new words with the length thrown first and spelled to it from the glyph rows, an exchange or drop at the first glyph, the paragraph gallows and the line letter at the text's rates, a fresh-word rate by hand, one lot a word | 12 | 22 | 2,983 | 1,395 | 9.9 | 4.75 |
| Tables and lots with the four pen habits (R8c) | R5 with the line letter, the paragraph gallows, the line-end stroke change and the loops on top lines in place of the line tables | 12 | 22 | 21,305 | 6,231 | 71.0 | 5.86 |
| Linked part tables with a capped page join (refined_s9) | Part tables per regime (writing style), 20 rules, a page join with a cap, a 24-cell sheet | 12 | 22 | 2,745 | 920 | 9.2 | 6.30 |
| Lot book of word classes (LB3c) | Word classes read by throws, a chain of references from class to class, chapters, two books for the two regimes | 12 | 22 | . | . | . | 5.66 |
| Every device, no size limit (p2best) | Linked tables with whole-word rows, glyph rules, joins of two words' halves and a length cap, found by a search | 9 | 22 | 242,048 | 9,960 | 806.8 | 5.57 |
| The procedure above (H3) | Two word tables of 600 words a regime read by lots, a rule table, line tables, ruler copying, a renewed sheet | 9 | 21 | 18,776 | 7,309 | 62.6 | 4.89 |
| Copying frame with part rows (H4b) | Seed tables, a splice, the four pen habits, part rows as a third source, a tail rate by ending class | 8 | 21 | 2,838 | 1,214 | 9.5 | 6.76 |
| Part rows found by the automated search, the third line (parts_g1) | Part rows by word class, a copy from the line above, a front-and-back join of words in sight, a part swap after a copy, the four pen habits | 7 | 21 | 2,670 | 1,136 | 8.9 | 5.46 |
| Linked part tables, unlimited (parts/d) | Three linked tables, ruler copies, a sheet | 9 | 21 | 55,236 | 3,822 | 184.1 | 4.58 |
| Small recipe with the pen habits (refined_s4) | The recipe held to 3,000 cells with the paragraph letter, line letter, loops and tails as habits | 13 | 20 | 2,785 | 947 | 9.3 | 7.64 |
| Linked part rows refitted (SC2) | Part rows of 18 cells, a length throw, quota filling, 36-cell start rows, 80 rules a regime, a forbidden-pair list | 10 | 20 | 2,409 | 888 | 8.0 | 12.76 |
| Marked fragment wheels (W3) | Three wheels with head and tail rings, ruler copying, a rule wheel and a working wheel, two wheel sets | 8 | 20 | . | . | . | 6.08 |
| Copying shrunk to the bound (C6e) | Two seed tables, 100 shared one-glyph rules, a forbidden-pair list, the four pen habits, a splice, a sheet | 8 | 20 | 2,828 | 1,335 | 9.4 | 4.46 |
| Hand-sized zairja (Z6) | Glyph chain rows of 36 cells read by two dice, with copying and a sheet | 5 | 20 | . | . | . | 4.97 |
| Glyph chain (G3) | Start rows, boundary rows and successor rows keyed by the last glyph and its position, a sheet, ruler copying, a splice, the habit rows | 10 | 19 | 2,838 | 1,054 | 9.5 | 6.30 |
| Small lot book (LB5) | One shared class word table, chain rows keyed by the last glyph, line-head and line-end rows, a sheet, rules, a forbidden-pair list | 7 | 19 | 8,594 | 2,747 | 28.6 | 5.57 |
| Syllable tables (Y3) | Beginning, middle and ending part tables, copying, a sheet | 5 | 19 | . | . | . | 4.21 |
| Copy and vary as the whole procedure (C5) | A seed table, rules and line tables, most words copied from the page and changed | 4 | 19 | 5,700 | 1,705 | . | 3.30 |
| Linked part rows (SC) | Part rows of 18 cells, 80 rules a regime, a forbidden-pair list, the sheet, ruler copying, the four pen habits | 12 | 18 | 2,037 | 668 | 6.8 | 7.40 |
| Wheel and grid | A wheel-and-grid device of the lot-book kind | 6 | 18 | 2,419 | 919 | 8.1 | 2.66 |
| Part rows tuned on held-out quires (ho_parts_g1) | The automated search's part rows with the splice raised for the once-only words and no glyph added | 8 | 17 | 2,236 | 988 | 7.5 | 5.46 |
| Part rows with the length drawn first (fit_g1) | The automated search's part rows with each word's length thrown first and its parts filled to it; a draw that misses the length is rejected and thrown again | 6 | 17 | 2,742 | 1,152 | 9.1 | 9.34 |
| Pointer walk (PW6) | A finger walking a stock list by two dice and a coin, a splice from the line above, rules, the four pen habits, a sheet | 6 | 17 | 2,868 | 1,624 | 9.6 | 5.77 |
| Wheels under a hard bound (W4) | Prefix, core, suffix and stop-chain wheels, edge-habit rows, ruler copying, a working wheel | 8 | 15 | 2,696 | 1,008 | 9.0 | 11.14 |
| Chancery nomenclator on Italian | A codebook with homophones, nulls, a syllabary, rule cells and a repeat-null | 8 | 15 | 1,386 | 736 | 4.6 | 1.96 |
| Word columns | Each word from a column for its position in the line | 4 | 15 | . | . | . | 3.34 |
| Formula frames | Fixed line frames filled with words from lists | 6 | 14 | . | . | . | 2.31 |
| Stencil grid | A grid of stems walked with one part swapped a step | 6 | 13 | . | . | . | 2.18 |
| The new-word step from the page evidence (newword_g1) | The same base with a front and a back cut from the once-only words at a length drawn first, no glyph added, no ruler-tight copy; retried until the length matches | 4 | 13 | 2,346 | 1,256 | 7.8 | 10.30 |
| Two sheets, one die (L5_two) | The one-sheet design below with a shape row and eight columns for each hand | 3 | 13 | 2,426 | 423 | 8.1 | 4.97 |
| Layout copying | The previous page's line shapes copied, each word from a glossary by length | 8 | 12 | . | . | . | 1.10 |
| One sheet, one die (L4_len2) | A row of word shapes with two length classes over eight columns of glyph groups, one throw a filled column; two pen rules, the four pen habits, a look-back to the lines above and the pages before | 4 | 12 | 1,638 | 228 | 5.5 | 5.01 |
| Lot-book dice rows | Every row held to sixteen entries read by the sum of three dice | 2 | 11 | . | . | . | . |
| Rhyme table (R5 rimarium) | Words grouped by ending, entered by dice and walked with a finger, a rare list, edge habits, a sheet, a ruler | 5 | 9 | 3,727 | 2,236 | 12.4 | 4.50 |
| Abbreviated Latin | Latin abbreviated with a suspension and sign table, in a fitted invented alphabet | 5 | 9 | 368 | 249 | 1.2 | 1.70 |
| Paradigm table | Words inflected through a paradigm | 1 | 7 | 2,228 | 415 | 7.4 | 4.31 |
| Knucklebones (A6) | Two throws of five astragali read on sheets of fronts and backs, one sheet per five pages, rules of the hand | 4 | 7 | 4,587 | 629 | 15.3 | 2.28 |
| Word mill | A card's table read as a mill | 3 | 7 | 35,172 | 34,888 | 117.2 | 1.60 |
| Stems and endings | Two short tables | 1 | 5 | 1,089 | 263 | 3.6 | 2.26 |
| Repertory | Pieces combined by throws | 3 | 5 | 1,294 | 327 | 4.3 | 5.43 |
| Dice tables (B7) | Two throws of three dice read on 56-cell tables, ten table pairs | 2 | 4 | 1,979 | 276 | 6.6 | 3.10 |
| Geomancy | Geomantic figures cast and read off a table | 0 | 4 | 772 | 180 | 2.6 | 2.07 |
| Ogam transformation | An Ogam-tract transformation applied to every word of a Latin text | 1 | 2 | 1 | 1 | 0.0 | 0.00 |
Each row is the family's best five-seed row, scored as in the figure above. The columns give the number of the 23 statistics within the text's tolerance interval and the number within three times it. Then come the table cells, the distinct entries, the page sides at 300 cells a side, and the random acts a word as the family counted them. A dot means that the family's results file carries no count. Seven families of the first line were counted only in their reports, and those counts are not given here; the chart draws them as hollow marks at the count of their scoreboard row. The H3 row counts its cells as fitted, with the sheet and the rule table, which is why it exceeds the 17,976 of the description above. The held-out figures of every row scored on the quire halves are in the table further down, with what each row was tuned on. The interim_s11 row is by its own count; counted with every row the generator draws from, as this report counts, it is 3,725 cells and 1,471 entries. The measures of the gap between Currier A and B (the two writing styles Prescott Currier identified) are given in the text where a row is discussed, with the rule that produced them.
What each mechanism buys, and where it fails. The families sort into three groups by what they can and cannot do. Devices that pick from a small stock and add nothing fail on the vocabulary at the first test. The dice tables, the knucklebones, the geomantic figures, the repertory and the stems-and-endings tables are of this kind. A stock of a few hundred pieces cannot produce 6,983 distinct words with the manuscript's long tail of words used once. The lot-book rows of sixteen entries lose the spread of each row. The wheels under 3,000 sectors lose the once-only words by 14 tolerances. Copying from the page with variation, the mechanism of Timm and Schinner's account, makes the vocabulary but bursts it wrongly. As the whole procedure it misses the burstiness of words by seven tolerances. Shrunk to the bound and given a splice and the pen habits, it reaches 20 statistics within three tolerances at 2,828 cells. The families that build words from parts do best under the bound. Linked rows of beginnings, middles and endings, a glyph-by-glyph chain and the hybrids reach 19 to 23 statistics within three tolerances at 2,000 to 2,900 cells. The automated search, which started from no design, arrived at the same part rows on its own, with a copy from the line above, a join of two words in sight and a part swap. The ladders of each family show which device pays for which statistic. Rows of glyphs as the source of fresh words restore the vocabulary without a rule row. In the copying frame, dropping the ruler brings the recurrence profile from three of six classes to six of six. The recurrence profile is the set of six measures of how often a word comes back across the page, the leaf and the gathering. The word just written, written again with one glyph exchanged or dropped three times in a hundred, is the only device that brings the adjacent one-edit pairs to the real value. A stop die thrown from the seventh glyph onward repairs the spread of word lengths. The first line's tables with lots reach 23 within three tolerances only at 21,305 cells, and their shrink ladder shows why. The score is bought with the two stocks and the twelve boundary rows. From R9A the rungs go 8 in tolerance and 22 within three at 17,305 cells, then 8 and 16 at 9,305, then 5 and 15 with one shared table set at 4,033. No rung under 3,000 cells keeps 20 of the 23.
The simplest device, and what each further device buys. The fourth line of the search asked the opposite question: not which apparatus reaches the most measurements, but which is the smallest that explains most of them, and what each further measurement costs. Its answer is one sheet. At its head is a row of word shapes, and under it eight columns, each a short row of glyph groups such as qok, ch, ok, aiin or dy, the seventy or so commonest groups of the book. A word is one throw on the shape row, which says which columns to fill, and one throw in each filled column, read left to right. That sheet alone, at 811 cells and three throws a word with no rules and no copying, makes an open vocabulary of the right size, with the right share of once-only words and the right density of one-glyph neighbours. No design of two throws or of one grid reaches those at any size, and the rows of a table whose cells are word parts, not glyph groups, give an open vocabulary only through their variation steps. Four additions bring it to 4 of the 23 statistics within tolerance and 12 within three, at 1,638 cells and 5.0 acts a word. They are two pen rules, the four pen habits, a look-back that writes a quarter of the words again from the lines above, and a length class in the shape row. A sheet for each hand adds the gap between Currier A and B, at 2,426 cells. The ladder from the bare sheet upward says which device pays for which measurement. The vocabulary measures come from the sheet itself, and the word-edge measures from the pen rules and habits. The neighbour and repeat measures come from the look-back, the length spread only from a length class in the shape row, and the gap between the hands only from a second sheet. What no rung reaches is a short and stable list. The burst of words on a page and the far classes of the recurrence profile need a vocabulary that belongs to the page and the quire, which is what a working sheet supplies. The once-only share falls as soon as a quarter of the words are copies. And the link across a word boundary is not in the shape of a word. Held out by quire the sheet stands at 2.53, about 1.2 above the floor (the distance between the text's own two halves), with nothing tuned on any half and no count to scale. Its once-only share is short by 9 to 15 tolerances on every half, twice the size of the same wall on the rows at the floor. A page classifier (a program that learns to tell real pages from generated ones) tells its pages apart 97 per cent of the time, on the absence of one-glyph words. So the rows at the floor are not carrying more than they need, and the table below sets out what each carries, from the sheet to the largest tables.
| Row | Devices | What is written down | Rules the scribe applies | Acts a word | Cells | Whole text | Held out |
|---|---|---|---|---|---|---|---|
| One sheet, one die (L4_len2) | one sheet, one die or a bag of lots | two shape rows of 300, eight columns of glyph groups, the four habit rows | two pen rules, the four habits, the look-back | 5.0 | 1,638 | 4 / 12 | 2.53, nothing tuned |
| Two sheets, one die (L5_two) | two sheets, one die | a shape row and eight columns for each hand, the habit rows | the same | 5.0 | 2,426 | 3 / 13 | 2.49, nothing tuned |
| One lot a word (throw_s18start) | part tables, lots, a working sheet, a ruler | beginning, body and ending rows, a stock of 200 words, a list of 30 glyph rules, the habit rows, the sheet | one lot a word read against its cut points, the swap of a middle, the join, the sheet's renewal, the four habits | 3.24 and 0.27 checks | 2,100 | 13 / 22 | 1.44 fixed counts, 1.21 scaled; nothing tuned |
| The recipe read off the page (DR) | lots, a working sheet, no ruler | a table of 200 words for each hand, rows of glyphs, a length row, an exchange table, the habit rows, the sheet | nine steps, set out on its companion page | 4.75 | 2,983 | 12 / 22 | 1.25, nothing tuned |
| Length first and the new-word step (ho4_fc1) | as throw_s18start | as throw_s18start, with a length row for each hand, lists of 300 fronts and 300 backs and an exchange table of 50 | as throw_s18start, with the length thrown first and the new-word step | 3.68 and 0.23 checks | 2,944 | 12 / 23 | 1.26 fixed counts, 1.07 scaled; tuned on all five splits |
| The same with a split left out (lo1_final_tail) | as above | as above, the join tails as one row for each hand | as above | 3.85 and 0.35 checks | 2,968 | 12 / 23 | split 1 left out: 1.43 scaled, 2.09 fixed counts |
| The same under the clean protocol, split 1 reported (v1b_tail) | as above | as above | as above | 3.51 and 0.20 checks | 2,988 | 9 / 23 | split 1 reported: 1.65 scaled, 2.37 fixed counts; tuned on splits 3, 4 and 5, stopped on split 2 |
| The same under the clean protocol, split 2 reported (v2_tail) | as above | as above | as above | 3.95 and 0.24 checks | 2,950 | 10 / 21 | split 2 reported: 1.12 scaled, 1.09 fixed counts; tuned on splits 1, 3 and 4, stopped on split 5 |
| The same under the clean protocol, split 3 reported (v3_tail) | as above | as above | as above | 3.55 and 0.38 checks | 2,900 | 11 / 22 | split 3 reported: 1.08 scaled, 1.31 fixed counts; tuned on splits 1, 2 and 5, stopped on split 4 |
| The same under the clean protocol, split 4 reported (v4_tail) | as above | as above | as above | 3.66 and 0.28 checks | 3,000 | 12 / 23 | split 4 reported: 0.95 scaled, 1.04 fixed counts; tuned on splits 1, 2 and 5, stopped on split 3 |
| The same under the clean protocol, split 5 reported (v5b_tail) | as above | as above | as above | 3.70 and 0.33 checks | 2,872 | 12 / 23 | split 5 reported: 0.98 scaled, 1.10 fixed counts; tuned on splits 1, 2 and 3, stopped on split 4 |
| Tables and lots with three redraws (R9Y) | two table sets, lots, a die, a ruler, a working sheet | two word tables for each regime, twelve boundary rows, a rule table, the habit rows, the sheet | the ruler's copy, the splice, three redraw rules, the four habits | 6.08 | 21,305 | 8 / 23 | 1.29, nothing tuned |
Each row is described by what a scribe would need at the desk: the devices, the written apparatus, the rules kept in mind, the random acts a word and the cells, by this report's count. The whole-text column is the number of the 23 statistics within tolerance and within three tolerances. The held-out column names the cut and what the row was tuned on, since a figure tuned on the splits is not a clean one. The one-sheet rows come from the fourth line's files, the others from the files named in the table above.
The kit, timed by hand. The final row of the tuned family was also run as a scribe would run it, with a kit built for it. The kit is four dice, eleven sides of dice-shaped tables and a wax tablet of six columns by three cells for each hand, and it writes the text the row writes. Over five whole-book runs its pages come within three tolerances on all 23 statistics, with 10 within tolerance, three of the six recurrence classes within tolerance and three of the four gap measures within tolerance and all four within three. Counted as the rows are counted, its tables, lists and tablet hold 2,712 cells, against the row's own 2,968, the difference being the kit's rounding of every row to a square or strip a die can address. Written out in full, with the lot's cut points and rates and the check's grid of allowed glyph pairs made into cells a scribe can read, the kit holds 3,276. So the rules a scribe follows cost a further 564 cells, about a fifth again, and the whole kit stands over the bound. The row's own five runs give 12 within tolerance and 23 within three at a misfit of 26.3, so the rounding costs two statistics within tolerance and half a unit of misfit. Every act was timed from a measured human speed where one exists and from the keystroke-level model of Card, Moran and Newell where none does. A word takes 33 seconds at the central rates and 21 to 52 at the low and high ones, for 3.65 throws and 3.9 cells read. About a second of that goes to rebuilding the tablet's struck cells when it is renewed. A line of eight or nine words takes 4.5 minutes, a page of 168 words about 1.5 hours, and the book of 34,780 words 320 hours, which is 53 six-hour days, or 33 to 84 across the range. Writing alone would take 6.7 seconds a word and eleven days. A third of the time goes to throwing, a quarter to finding leaves and cells, a sixth to reading the dice and the cells and a fifth to writing. A fair copy of the draft, a scribe reading it and writing it clean at the plain writing rate, takes 7.2 seconds a word and 71 hours, twelve six-hour days. So the draft and the fair copy together take 391 hours, 65 six-hour days, or 39 to 107 across the range. A professional copyist of the period, at 2.36 folios a day, would manage about 793 words a day of a book written as sparsely as this one, and the kit gives 651 in six hours. The fair copy alone would take such a copyist 44 calendar days. The kit and its test are on the page Four Dice, One Tablet, every throw of five pages is on The Dice Ledger, and the one-page instruction sheet is The Scribe's Instructions. The kit itself runs in the browser, with real or virtual dice and a stopwatch, on Try the Four Dice.
Four habits read off the text. Every family that scores well carries the same four things, and each was read off the manuscript here, not taken from a prior account. The first is where the rare words stand. In the text 17 per cent of line-opening words and 21 per cent of line-ending words occur once in the whole manuscript. Inside a line the share is 9 per cent, and on the top lines of paragraphs 23. The second is the line letter. About half of the line-opening words are a single letter, one of y, s, d, o, t or l, written in front of a commoner word. Which letter, and whether at all, depends on how that word begins, and it is never written before q. The third is the paragraph gallows. In 82 per cent of paragraphs the first word is a tall letter, one of p, t, k or f, written before an ordinary word. Four fifths of the text's p and f stand on the top lines of paragraphs. The fourth is the line-end tail. Of the line-final words 17.5 per cent end in m or g, and 12 per cent are a commoner word with its last stroke finished as a tail. A page classifier is a program shown real pages and generated pages and asked to tell them apart. When a candidate lacked the paragraph gallows, the classifier keyed on it first: 0.83 of paragraphs in the text against 0.15 in the first recipe without it. Written as a habit, a gallows glyph before a stock word as the first word of a paragraph, the rate comes out at 0.81 to 0.90 with no cost to the score. The candidate that draws its paragraph word from a table without the habit sits at 0.38. The classifier then keys on what is left. In throw_s18start, the row at the floor from the second line (one of the four independent lines of the search), the line letter is written before nearly every line-first word. So lines opening with a plain ch word are 0.4 per cent of lines, against 3.9 in the text. Lines opening with a plain sh word are 0.2 per cent against 5.4. Words beginning with d are 0.8 per cent of glyph pairs against 2.1. The d-initial deficit has another cause, found on the second line. As a share of words, 9.4 per cent of the text's words begin with d, 15 per cent in Currier A and 6.7 in B, against 5.0 per cent in that line's tuned rows. Its word-cutting model assigns d as a beginning to under 2 per cent of the text's words, so the d-words live in the head of the frequency curve. There daiin alone is 2.4 per cent of all words, and dar, dy, dal and dain are 6 per cent together. That family builds its head words by copying what the page already has, so which words become the commonest is left to chance, and its daiin comes out at 0.6 per cent. The missing device is a short list of the commonest words at their rates. A probe of one on the row tuned on all five splits, 100 cells drawn one time in ten, moved the classifier from 0.903 to 0.873 and daiin to 1.1 per cent, by that line's account. Lowering the habit's rate from one and a half times the text's to the text's own leaves roughly a third of line openers bare. The automated search's part-table row counts the text's ch and sh openers at 3.5 and 4.7 per cent of lines. There the lower rate moves the candidate's openers from near zero to 5.8 and 2.4 per cent, at a small cost to the line-opener statistics. The tables with lots write these openers at the text's rate already, with plain ch and sh openers at 3.6 and 5.7 per cent of lines against 3.9 and 5.4.
The join. A fifth finding fixes what the rare words are made of. The manuscript has 4,743 words that occur once. Of these, 53.5 per cent are two attested words written together, each half of two or more glyphs and attested twice or more, and the halves are short, about three glyphs each. By the same strict rule Caesar's Latin gives 5 per cent, the Aeneid 6, Dante 13 to 14 and the King James Bible 3 to 5. A control that shuffles the letters gives 3 to 8 anywhere. For 57.5 per cent of these words one half stands earlier on the same page, against 45.8 for a control matched in word counts. For 14.5 per cent both halves do, against 6.9. A looser rule, any front of a common word plus any back of another, gives 77 to 81 per cent for the manuscript and the same for Latin, so only the strict figure means anything. The join could in principle be a copyist's slip, since copyists of name lists do fuse neighbouring names. The real rate of that slip was measured on the copies of the Sworn Book of Honorius, a magic book whose lists of non-words were recopied through the fourteenth and fifteenth centuries. A copyist alters about one name in ten to twelve per copy, and 70 per cent of the alterations are a single letter. Two whole names are fused in 0.8 per cent of names. A true join of the front of one name to the back of another occurs once in 2,899 names. The text needs the join for about one word in ten. So the join step of every candidate here is a rule the scribe applied, not drift in copying. The candidates that carry it, the copying frames, the part rows and the tables with a splice, write the join at 0.40 to 0.53 of their once-only words. The text's figure is 0.51 to 0.535 by the two definitions used.
The held-out checks. A match on the whole text is not enough, because every rate was fitted to the whole text. Three further checks were applied to the leading rows, and they decide the outcome. The first fills the candidate's tables from half of the quires, the gatherings of folded sheets that make up the book. It then generates the other half and scores 27 statistics against that half's own values, over five such splits. The floor is 1.29, the distance in the same units between the text's own two halves. The second is the page classifier described above. Its null, the accuracy a classifier reaches on two halves of the real text, lies near 0.61. The third is the set of gap statistics. Two of them matter most here: a joint test of all 23 statistics against the cloud of the candidate's own seeds, and a test of whether a line's words carry information on its length, as the text's do. A held-out figure depends on what the candidate was tuned on, and a figure tuned on a fixed subset of the splits overstates. The automated search tuned one row on three of the five splits and reached 1.70 on those three but 1.78 on all five. Honest tuning rotates the splits, and the honest figure for a tuned row is the one split its tuning never saw. The table further down gives every row scored on the quire halves, with what it was tuned on, and every figure in it is a mean over five seeds a split. The rows tuned on no split come first. R4, the first repaired table set, scores 1.82 with its rates refitted from the training half and 1.54 with them fixed. R8c scores 1.25, R9A 1.24 and R9Y 1.29. The recipe read off the page, described in the next subsection, scores 1.25 with its rates fixed, ranked first of nine generators in every split, and the hybrid J3 it was built on scores 1.36. Its tables are rebuilt from each training half by the code that built them from the whole text, and they come out at 2,917 to 3,004 cells against 2,983. The second line's rows were fitted on the whole text and scored on all five splits, with the tables filled from the training half and the rates fixed: refined_s9 at 1.51, interim_s11 at 1.48 and throw_s18start at 1.44. Its two whole-text rows score 1.49 for refined_s18 and 1.55 for refined_s19, on the same footing. The automated search's parts_g1 scores 1.81, and its two evidence rows with the length drawn first, fit_g1 and newword_g1, score 1.74 and 2.05; their files record no tuned split. On the page classifier every row is separable from the real pages, at 0.80 to 0.92 against the null of 0.61. The lowest are R9Y at 0.800, the hybrid J3 at 0.808, refined_s18 at 0.820 and the recipe read off the page at 0.825. R4 stands at 0.894, R8c at 0.859 and R9A at 0.835. On the second line throw_s18start stands at 0.880, interim_s11 at 0.887, refined_s9 at 0.890 and refined_s19 at 0.850, and the automated search's parts_g1 at 0.872. The rows scored on the classifier alone stand at 0.916 for the part rows, 0.895 for the hybrid H4b, 0.844 for the pointer walk and 0.832 for the glyph chain.
A rule of thumb that starves the tables. One step of the second line's table building turned out to depend on the size of the text it is built from, and the finding changed several held-out figures. That line cuts each word of the text into a beginning, a middle and an ending. It keeps a beginning in its tables only if at least 50 distinct words carry it, and an ending only if at least 40 do. The two counts were set on the whole text, where they keep 34 beginnings and 20 endings. Built from half the quires, the same counts keep 19 and 11 on the smallest training half, because a half carries about two thirds of the book's distinct words. The built words then come out shorter while the joined words stay long, and the spread of word lengths on the held-out half widens. That is a fault of the check and not of the procedure, since the scribe of the account has tables built from the whole. The repair scales the two counts to the training half's share of the distinct words, so that a half keeps 29 to 32 beginnings and 17 to 20 endings, and it changes nothing else. Under the scaled counts throw_s18start scores 1.21 held out in place of 1.44, and refined_s18 1.33 in place of 1.49. The spread of word lengths comes within three tolerances on every split for the first and on two of five for the second. About half of that line's length-spread miss was the check, and the once-only miss was not: the share of once-used words stays beyond three tolerances on three of the five splits under either cut. The automated search's part rows carry the same counts, and by that line's own account its parts_g1 moves from 1.81 to 1.67 under the scaled ones, with its once-only miss unchanged at seven to nine tolerances. The scaled counts were chosen after it was seen that they help. So both figures are given for every row that carries the counts, and no ranking below rests on one of them alone. An audit of every absolute count in the table building of every line found no other count that starves a table. The first line's caps are full on every half, since a half still attests more entries than any cap keeps, and its presence tests cannot be scaled without abolishing them. The recipe read off the page has one floor, an observation count of 20 on its pen-habit rows, which lose a sixth to a quarter of their cells on a half. Scaled to the half, that floor makes the recipe's figure worse, 1.34 against 1.25, so its figure is quoted as run. The audit also reproduced throw_s18start's two figures to the third decimal from the second line's files, and found interim_s11 moving from 1.48 to 1.19 and refined_s9 from 1.51 to 1.44 under the scaled counts.
| Row | Line | Tuned on | Standard cut | Scaled cut | Within two tolerances, of 27 | Beyond three tolerances in every split | Left-out split | Page classifier |
|---|---|---|---|---|---|---|---|---|
| R4 tables, rates refitted | first | rates refitted on each training half | 1.82 | no cut step | 16.6 | none | 0.894 | |
| R4 tables, rates fixed | first | no split | 1.54 | no cut step | 18.6 | S6_ | 0.894 | |
| R8c tables with habits | first | no split | 1.25 | no cut step | 22.2 | none | 0.859 | |
| R9A tables, die on the ruler | first | no split | 1.24 | no cut step | 22.4 | none | 0.835 | |
| R9Y tables, three redraws | first | no split | 1.29 | no cut step | 21.8 | none | 0.800 | |
| hybrid J3 (DR0) | first | no split | 1.36 | no cut step | 19.8 | none | 0.808 | |
| DR_final recipe off the page | first | no split | 1.25 | no cut step | 22.2 | none | 0.825 | |
| DR8 recipe, fresh-word rate by hand | first | no split | 1.26 | no cut step | 21.0 | none | ||
| refined_s9 part tables | second | no split | 1.51 | 1.44 | 21.0 / 20.8 | S2_ | 0.890 | |
| interim_s11 part tables | second | no split | 1.48 | 1.19 | 20.0 / 22.2 | S3_ | 0.887 | |
| throw_s18start, one lot a word | second | no split | 1.44 | 1.21 | 19.6 / 22.4 | S2_ | 0.880 | |
| refined_s18 whole-text row | second | no split | 1.49 | 1.33 | 20.6 / 21.0 | S2_ | 0.820 | |
| refined_s19 whole-text row | second | no split | 1.55 | not run | 19.8 | S2_ | 0.850 | |
| every device, no size limit | second | no split | 2.01 | not run | 17.0 | S2_ | 0.907 | |
| ho4_fc1, length first and new-word step | second | rotating splits, all five used | 1.26 | 1.07 | 21.2 / 24.0 | none / none | 0.903 | |
| lo1_final_tail, split 1 left out, end of the search | second | splits 2 to 5; split 1 left out | 1.31 | 1.01 | 21.6 / 24.0 | none / none | split 1: 1.43 scaled, 2.09 standard (floor 1.18) | 0.890 |
| lo1top_fc1, commonest-words row, first full check | second | splits 2 to 5; split 1 left out | 1.26 | 1.12 | 21.8 / 23.6 | none / none | split 1: 1.46 scaled, 1.69 standard (floor 1.18) | 0.877 |
| lo1top_final, the same at the end of its search | second | splits 2 to 5; split 1 left out | 1.20 | 1.01 | 21.8 / 23.6 | none / none | split 1: 1.72 scaled, 1.92 standard (floor 1.18) | 0.861 |
| v1_tail, clean protocol, reported point | second | splits 3, 4 and 5; stopped on split 2; split 1 left out | 1.25 | 1.10 | 21.2 / 23.6 | none / none | split 1: 1.86 scaled, 2.03 standard (floor 1.18) | 0.852 |
| v1b_tail, clean protocol, re-run under the fixed cap, reported point | second | splits 3, 4 and 5; stopped on split 2; split 1 left out | 1.35 | 0.97 | 21.0 / 23.8 | none / none | split 1: 1.65 scaled, 2.37 standard (floor 1.18) | 0.864 |
| v1_tail_end, clean protocol, end point | second | splits 3, 4 and 5; stopped on split 2; split 1 left out | 1.26 | 1.11 | 21.0 / 23.6 | S3_ | split 1: 1.81 scaled, 1.93 standard (floor 1.18) | 0.852 |
| v1b_tail_end, clean protocol, re-run under the fixed cap, end point | second | splits 3, 4 and 5; stopped on split 2; split 1 left out | 1.35 | 0.97 | 21.0 / 23.8 | none / none | split 1: 1.65 scaled, 2.37 standard (floor 1.18) | 0.864 |
| v2_tail, clean protocol, reported point | second | splits 1, 3 and 4; stopped on split 5; split 2 left out | 1.35 | 1.13 | 20.6 / 22.0 | none / none | split 2: 1.12 scaled, 1.09 standard (floor 1.64) | 0.863 |
| v2_tail_end, clean protocol, end point | second | splits 1, 3 and 4; stopped on split 5; split 2 left out | 1.40 | 1.16 | 21.0 / 22.8 | none / none | split 2: 1.16 scaled, 1.25 standard (floor 1.64) | 0.886 |
| v3_tail, clean protocol, reported point | second | splits 1, 2 and 5; stopped on split 4; split 3 left out | 1.37 | 1.22 | 20.8 / 22.2 | S2_ | split 3: 1.08 scaled, 1.31 standard (floor 1.14) | 0.933 |
| v4_tail, clean protocol, reported point | second | splits 1, 2 and 5; stopped on split 3; split 4 left out | 1.26 | 1.07 | 21.2 / 22.6 | none / none | split 4: 0.95 scaled, 1.04 standard (floor 1.33) | 0.913 |
| v4_tail_end, clean protocol, end point | second | splits 1, 2 and 5; stopped on split 3; split 4 left out | 1.43 | 1.11 | 19.6 / 22.6 | S2_ | split 4: 0.99 scaled, 1.13 standard (floor 1.33) | 0.906 |
| v5_tail, clean protocol, reported point | second | splits 1, 2 and 3; stopped on split 4; split 5 left out | 1.35 | 1.16 | 19.8 / 22.6 | none / none | split 5: 0.93 scaled, 1.13 standard (floor 1.18) | 0.917 |
| v5b_tail, clean protocol, re-run under the fixed cap, reported point | second | splits 1, 2 and 3; stopped on split 4; split 5 left out | 1.31 | 1.09 | 20.6 / 23.2 | none / none | split 5: 0.98 scaled, 1.10 standard (floor 1.18) | 0.879 |
| v5_tail_end, clean protocol, end point | second | splits 1, 2 and 3; stopped on split 4; split 5 left out | 1.35 | 1.16 | 19.8 / 22.6 | none / none | split 5: 0.93 scaled, 1.13 standard (floor 1.18) | 0.917 |
| v5b_tail_end, clean protocol, re-run under the fixed cap, end point | second | splits 1, 2 and 3; stopped on split 4; split 5 left out | 1.34 | 1.23 | 21.2 / 22.8 | S2_ | split 5: 1.15 scaled, 1.19 standard (floor 1.18) | 0.874 |
| lo3_final, split 3 left out | second | splits 1, 2, 4 and 5; split 3 left out | 1.41 | 1.16 | 21.0 / 22.0 | none / none | split 3: 1.04 scaled, 1.29 standard (floor 1.14) | 0.874 |
| parts_g1, automated | third | no split | 1.81 | not run | 19.8 | S2_ | 0.872 | |
| ho_parts_g1, held-out tuned | third | splits 1 to 3 | 1.78 | not run | 20.0 | S2_ | ||
| fit_g1, length first | third | not recorded in its file | 1.74 | not run | 19.6 | S3_ | ||
| newword_g1, new-word step | third | not recorded in its file | 2.05 | not run | 16.4 | S6_ | ||
| combo_g1, both steps | third | not recorded in its file | 1.79 | not run | 16.4 | none | ||
| c1_tail, clean protocol with the copy stage in the grammar, reported point | second | splits 3, 4 and 5; stopped on split 2; split 1 left out | 1.40 | 1.16 | 20.0 / 21.6 | none / none | split 1: 1.51 scaled, 2.01 standard (floor 1.18) | 0.911 |
| c2_tail, clean protocol with the copy stage in the grammar, reported point | second | splits 1, 3 and 4; stopped on split 5; split 2 left out | 1.24 | 1.17 | 21.8 / 22.0 | none / none | split 2: 1.22 scaled, 1.24 standard (floor 1.64) | 0.911 |
| c3_tail, clean protocol with the copy stage in the grammar, reported point | second | splits 1, 2 and 5; stopped on split 4; split 3 left out | 1.09 | 1.01 | 23.2 / 24.0 | none / none | split 3: 0.83 scaled, 0.89 standard (floor 1.14) | 0.875 |
| c4_tail, clean protocol with the copy stage in the grammar, reported point | second | splits 1, 2 and 5; stopped on split 3; split 4 left out | 1.29 | 1.20 | 22.0 / 22.0 | none / none | split 4: 1.12 scaled, 0.98 standard (floor 1.33) | 0.916 |
| c5_tail, clean protocol with the copy stage in the grammar, reported point | second | splits 1, 2 and 3; stopped on split 4; split 5 left out | 1.52 | 1.25 | 19.2 / 21.2 | none / none | split 5: 1.13 scaled, 1.23 standard (floor 1.18) | 0.913 |
Every row scored on the quire halves, from the files of the line that built it. The standard cut is each family's table building as written, and the scaled cut applies only to the rows whose part tables use the fixed counts; the floor is 1.29 under both. The count within two tolerances is over the 27 held-out statistics, averaged over the five splits, and the next column names the statistics beyond three tolerances in every one of the five. The left-out split is the one figure of a tuned row that its tuning never saw, given with that split's own floor. The page classifier is the accuracy of the logistic classifier on the row's first seed, against a null near 0.61.
The best row under the bound, and the closest at any size. By the whole-text count alone, four rows reach 23 of 23 within three tolerances: the hybrid J3, interim_s11, and two later search points of the second line, refined_s18 and refined_s19. Counted as this report counts, with every row the scribe draws from, two of them stand over the bound, interim_s11 at 3,725 cells and refined_s18 at 3,402. The whole-text match under the bound is refined_s19. On the whole text it lands within tolerance on 13 of the 23 statistics and within three tolerances on all 23, with a misfit of 25.0, and all six recurrence classes are within tolerance. It needs 2,870 cells, 1,002 entries and 9.6 sides, and it costs 3.33 random acts a word. The page classifier tells its pages from the real ones 85 per cent of the time. But held out by quire it scores 1.55 against the floor of 1.29. The spread of word lengths and the share of once-used words are beyond three tolerances in every split. refined_s18, the row with the lowest misfit of the four at 21.9, does no better held out. It scores 1.49, with 20.6 of the 27 statistics within two tolerances, and its words spelled fresh cost 11.08 bits a word against the text's 11.11, where throw_s18start's cost 11.67. J3 scores 1.36 held out, 0.07 above the floor, with the once-used share short on four of the five splits. So a whole-text match under the bound exists, and it does not carry to pages its tables were not built on. The held-out halves decide, and this account ranks by them first and by the page classifier second. Two rows under the bound stand at the floor held out, and neither was tuned on any split. One is the recipe read off the page, the row DR, described in the next subsection. It scores 1.25 against the floor of 1.29, with 22.2 of the 27 statistics within two tolerances, and the once-used share is its one statistic beyond three, on four of the five splits. The other is throw_s18start, found on the second line. It scores 1.44 with its tables cut by the fixed counts and 1.21 with the counts scaled to the half, with 19.6 and 22.4 of the 27 within two. The two rows are 0.19 apart one way under the fixed counts and 0.04 apart the other way under the scaled ones. So the held-out check does not separate them, and the page classifier does. It tells the recipe's pages from the real ones 82.5 per cent of the time and throw_s18start's 88 per cent, on the plain ch and sh line openers and the d-initial words named above. The best procedure under the bound is therefore the recipe read off the page, with throw_s18start beside it, and each has a companion page. The scribe of the recipe copies from the page with no ruler and throws one lot a word. A new word has its length thrown first and is spelled to it from rows of glyphs. A word is varied by an exchange or a drop at its first glyph, and the next subsection sets out the whole of it. The scribe of throw_s18start builds each word from linked part tables with one lot a word, taking the next usable cell instead of throwing again. The beginning and ending rows are shared between the two hands, the swap of a middle is read off the body row of the same beginning, and the paragraph letter is written at the text's rate. On the whole text it lands within tolerance on 13 of the 23 statistics and within three tolerances on 22, with a misfit of 23.9. The one miss is the entropy (bits of surprise) of the first glyph of a line's words. Four of the six recurrence classes are within tolerance. The four measures of the gap between Currier A and B are within a factor of two of the text's, which is the rule the second line used, though none is within the tolerance this report uses elsewhere. The values are a letter divergence of 0.0249 against the text's 0.0168, a gap in glyph predictability of 0.201 against 0.147, a shared vocabulary of 0.124 against 0.159, and a length gap of 0.39 against 0.30. It needs 2,100 cells, 899 distinct entries and 7.0 page sides, and it costs 3.24 random acts a word plus 0.27 checks. The companion page Writing Voynich by Hand lets a reader throw the lot word by word and watch a line of this text take shape under these rules. The page Writing Voynich by the Recipe Read off the Page does the same for the recipe. Held out by quire under the fixed counts, throw_s18start's spread of word lengths is beyond three tolerances in every split and its share of once-used words in four of five. Under the scaled counts the spread is within three everywhere, and the once-used share stays short by three to five tolerances on three splits. The rows tuned with the length drawn first and the new-word step reach lower figures on the splits they were tuned on, 1.07 and 1.01 under the scaled counts over all five. The check that matters is the split kept out of the tuning. There the family scores 1.43 on the hardest half against its floor of 1.18, where throw_s18start untuned scores 1.39, and 1.04 on an ordinary half against 1.14, where throw_s18start scores 1.10. So the tuning's gain on the halves is mostly the tuning seeing them, and the tuned rows do not displace the two untuned rows at the floor. The row tuned with split 1 left out counts 2,968 cells by this report's convention and the row tuned on all five splits 2,944. An earlier form of the first, which read keyed rows of join tails, counted 3,718, over the bound. One caution belongs beside these figures. A search that tunes on four halves and reports the fifth can still drift toward the four. A row of the same family with the commonest words in a shared row scored 1.46 on the left-out half at its first full check and 1.72 at the end of its search. Its figure on the tuning halves improved over the same stretch. So the end of such a search is not by itself the honest row, and each figure above is labelled with the point of the search it was taken at. The clean protocol tunes on three halves, stops on a fourth and reports the fifth, which no choice in the search ever sees. The family was run that way once for each half, from its untuned start each time, under the scaled counts and the bound. Reported on split 1, tuned on splits 3, 4 and 5 and stopped on split 2, it scores 1.65 against that half's floor of 1.18, with 18 of the 27 statistics within two tolerances. Five statistics are beyond three tolerances there, the share of the ten commonest words, the share of the hundred commonest words, the share of once-used words, the exact repeats and the first-glyph gain. Through the search the stopping half went from 1.69 to 0.87 and the reported half from 1.73 to 1.65. At the end of the search the reported half stood at 1.65. Reported on split 2, tuned on splits 1, 3 and 4 and stopped on split 5, it scores 1.12 against that half's floor of 1.64, with 23 of the 27 statistics within two tolerances. Nothing is beyond three tolerances there. Through the search the stopping half went from 1.39 to 0.97 and the reported half from 1.69 to 1.12. At the end of the search the reported half stood at 1.16. Reported on split 3, tuned on splits 1, 2 and 5 and stopped on split 4, it scores 1.08 against that half's floor of 1.14, with 25 of the 27 statistics within two tolerances. One statistic is beyond three tolerances there, the spread of word lengths. Through the search the stopping half went from 1.38 to 1.06 and the reported half from 1.23 to 1.08. Reported on split 4, tuned on splits 1, 2 and 5 and stopped on split 3, it scores 0.95 against that half's floor of 1.33, with 23 of the 27 statistics within two tolerances. One statistic is beyond three tolerances there, the glyph entropy. Through the search the stopping half went from 1.23 to 0.86 and the reported half from 1.38 to 0.95. At the end of the search the reported half stood at 0.99. Reported on split 5, tuned on splits 1, 2 and 3 and stopped on split 4, it scores 0.98 against that half's floor of 1.18, with 26 of the 27 statistics within two tolerances. Nothing is beyond three tolerances there. Through the search the stopping half went from 1.38 to 1.07 and the reported half from 1.39 to 0.98. At the end of the search the reported half stood at 1.15. So the family writes splits 2, 3, 4 and 5, which it never saw, as well as the text's own halves match each other, and it does not write split 1. Its limit is split 1, whose figure is the one to quote for the family. The reported rows count 2,988, 2,950, 2,900, 3,000 and 2,872 cells by this report's convention. The page classifier tells their pages from the real ones 86, 86, 93, 91 and 88 per cent of the time, in the same order. The row for split 4 sits exactly at the bound of 3,000 cells, which the search admits. The split-3 row's tell is one the battery does not carry, since that row draws a word's beginning without regard to the word before.
A second grammar was run on the same halves under the same protocol. The automated search's parts_g1 writes each word by a program of rules over rows of word parts, with a working sheet and the pen habits at the line's edges. Its texts were scored here by the same 27 statistics, and its scorer and this report's agree on every statistic to the fourth decimal. Tuned on splits 1, 3 and 5 and stopped on split 4, it gives split 2 1.45 against the floor of 1.64, at 1,840 cells. Tuned on splits 2, 3 and 4 and stopped on split 1, no tuning step improved the stop half, so its clean point is the untuned start: 1.55 on split 5 against 1.18, at 2,670 cells. Tuned to the end it reaches 1.34 on split 5 while the stop half worsens from 2.01 to 2.43. Tested on split 1, with splits 2, 3 and 4 for tuning and split 5 as the stop, it gives 2.43 against the floor of 1.18 and the untuned 2.01, at 2,020 cells. That point is the same program as the end of the second rotation, seen from its stop half. With an objective that also weighs the head of the frequency curve and the neighbours at one edit it gives 2.23, at 2,304 cells. So tuning on the hand-2 halves (the halves written by the second of the five scribal hands Lisa Fagin Davis identified) makes hand 3's half worse under either objective. Beside the recipe family's 1.12, 0.98 and 1.65 on the same three halves, the second grammar matches the held-out halves less closely, and on the hardest half it fails the same way. Hand 3's ten commonest words take 0.230 and 0.196 of its text against the real 0.133, and its words run 4.75 and 4.81 glyphs against 5.21. The page classifier tells its pages from the real ones 87 to 94 per cent of the time.
A fair copy with a copyist's slips. Every row above is a draft written straight from its procedure, and the manuscript is a fair copy in a practised hand. A fair copy carries slips, and their rate in a book of made-up names has been measured. Across the copies of the Sworn Book of Honorius, six to ten names in a hundred are altered in each copy, mostly by one letter and mostly by a letter of like shape. So the final row's text was copied once more with slips at that rate. Each slip exchanges a glyph for one a copyist confuses with it, the pairs one stroke or one loop apart. The pairs are a and o, ch and sh, k and t, f and p, i and ii, ii and iii, m and n, and r and s. Nothing was tuned, the rate and the pairs were fixed before any scoring, and split 1 was never fitted. On split 1, the hardest half, the row goes from 1.43 to 1.17 at eight slips in a hundred against the floor of 1.18, with 21 of the 27 statistics within two tolerances against 17 before. Its four tuned halves stay under their floors, at 1.46, 0.87, 1.15 and 1.09 against 1.64, 1.14, 1.33 and 1.18. At six slips in a hundred split 1 reads 1.22 and at ten 1.14. The share of once-used words, the exact repeats, the share of the ten commonest words, the Zipf slope and the mutual information come inside two tolerances. The similarity of neighbouring words, the glyph entropy, the share of the hundred commonest words, the initial-glyph profile, the first-glyph gain and the word-length profile stay beyond two. Slips of the same number drawn without regard to shape do not do this. At six in a hundred the row reads 1.31 on split 1 against 1.22 with the look-alike pairs, its share of once-used words goes seven tolerances over, and all four of its tuned halves go over their floors. So the gain comes from the shape of the slips. A look-alike slip mostly turns a word into another the book already has, which is what the manuscript's neighbouring words look like. The row for split 1 under the clean protocol and under the bound moves from 1.65 to 1.30 at ten in a hundred, still over its floor. These are figures under the scaled counts. Under the fixed counts the final row's never-fitted half goes from 2.09 to 1.85 at eight in a hundred against its floor of 1.18, so there the pass takes about a quarter of the distance. On the whole text the copy at eight in a hundred lands within tolerance on 16 of the 23 statistics and within three on 21, against the draft's 12 and 23. The two statistics beyond three are the mutual information beyond the shuffle and the glyph-pair entropy. At six in a hundred it lands within tolerance on 16 and within three on 22. The page classifier still tells the pages apart, at 89 per cent before and after, on the same tells. Re-searching the draft's tables with the copy stage in place gives nothing back. The search moves the draft toward its tuning halves and hands the diversity to the slips, and the never-fitted half returns to where the row is without the stage. This is an account of the shape of the residue (the part of the miss that remains), not a match. The slips are a copyist's and not the procedure's, and a draft followed by a fair copy is a two-scribe or two-season job, costed in the kit's timing above at 65 six-hour days for both passes.
The slips inside the search. Two more tests put the slips where a searching scribe would have had them. In the first the same look-alike table is laid over each clean row's own text on the half it never saw, at six, eight and ten slips in a hundred, with nothing re-tuned. In the second the copying stage is written into the grammar and the search runs again under the clean protocol. The slip rate is one more setting, between 0.06 and 0.10, and the tuning, stopping and reported halves are as before. Every figure is the mean over five seeds on the half the search never saw, under the scaled counts.
| Half | Floor | Clean row without the stage | The stage laid over it, six / eight / ten in a hundred | Searched with the stage |
|---|---|---|---|---|
| split 1 | 1.18 | 1.65 (v1b_tail) | 1.42 / 1.34 / 1.30 | 1.51 at six in a hundred, 2,890 cells |
| split 2 | 1.64 | 1.12 (v2_tail) | 1.33 / 1.42 / 1.49 | 1.22 at six in a hundred, 2,840 cells |
| split 3 | 1.14 | 1.08 (v3_tail) | 1.07 / 1.11 / 1.17 | 0.83 at seven in a hundred, 2,916 cells |
| split 4 | 1.33 | 0.95 (v4_tail) | 1.05 / 1.12 / 1.19 | 1.12 at eight in a hundred, 2,710 cells |
| split 5 | 1.18 | 0.98 (v5b_tail) | 1.07 / 1.15 / 1.28 | 1.13 at six in a hundred, 2,702 cells |
Laid over the clean rows, the stage helps only the hardest half. Split 1 goes from 1.65 to 1.30 at ten in a hundred against its floor of 1.18. Every ordinary half gets worse, by 0.03 to 0.30 at eight in a hundred. Those halves are at their floors without the slips, and the slips move them off it. Searched with the stage in place, four of the five halves stay under their floors. The hardest does not. Split 1 reads 1.51 against its floor of 1.18, worse than the 1.30 of the stage laid over the clean row after the search. A search that knows about the slips hands them the diversity and tunes the draft toward its tuning halves, as the refit of the final row did. Against the clean rows the searched rows are better on split 1 and split 3 (1.51 against 1.65, 0.83 against 1.08) and worse on split 2, split 4 and split 5. The searched rows carry 2,702 to 2,916 cells, under the bound, and the page classifier tells their pages from the real ones 88 to 92 per cent of the time. So the slips do their best work on the hardest half when laid over a draft tuned without them. Inside the search they buy one half at another's cost.
A record can stand in for a table. One more step was tried on v1b_tail: a strip of notes beside each of the 300 front cells of the new-word row. When the scribe makes a joined new word, about one word in ten, the throws for its front and back go as before. Before writing, the scribe looks at the strip beside the front cell, where every word already made from that front stands. If the word is there, the scribe moves to the next back of that length without a throw, until reaching a word not on the strip or running out of backs. The word goes on the page and, if new, on the strip. The strips start blank, are never cleared, and serve both scribes. A twin check at chance 0.35 and the look-alike slips at four in a hundred go with the step. Held out on split 1 the row scores 1.27 under the scaled counts and 1.83 under the fixed ones, against 1.65 and 2.37 for v1b_tail alone and the floor of 1.18. The figure is not clean. The chance was tuned on lo1_final_tail's other four halves, and the slip rate was set after the copy rows on every half had been seen. The rows stay at 2,988 cells, since no lot is ever thrown on a strip. By the end of the book the strips hold 3,356 to 3,522 entries, about eleven a strip, looked at 3,476 to 3,643 times. Counted as apparatus they put the step at 6,344 to 6,510 entries in all and 5,616 to 5,782 for each scribe, over the bound. Counted as a working record they are larger than the tables they sit beside. Chance acts come to 3.53 a word, and 4.77 with the slips' own throws.
Where the hardest half misses. Split 1 holds out quire T, the twenty-two stars pages in hand 3 with 10,361 words, together with the Currier-B quires H and N and five herbal and pharmaceutical quires in hand 1. Its B tables are read from 8,371 words that are 91 per cent hand 2, 81 per cent of them from the biological quire M, and are written into 14,957 words that are 70 per cent hand 3. It is the only half that judges hand 3 with none of that scribe's pages in the tables. The other four read their B tables 77 to 84 per cent from hand 3 and write hand 2. A half made of real pages on some quires and generated pages on the rest places the miss, under the scaled counts with five seeds against the floor of 1.18. With the herbal and pharmaceutical quires generated and H, N and T real, lo1_final_tail scores 0.35 and v1b_tail 0.38, with all 27 statistics within two tolerances. With only H, N and T generated they score 1.18 and 1.36, the whole of the miss, against 1.43 and 1.65 with every quire generated. Quire T alone gives 0.89 and 1.00. What the generated stars pages get wrong is the scribe. The procedure writes one Currier-B profile whatever the split. Its ten commonest words take 17 to 22 per cent of the tokens on every row. Hand 2's real pages give them 18.5 per cent, hand 3's 13.1 and hand 5's 12.3. On split 1 the generated hand-3 pages give the ten commonest words 0.04 more of the text than the real ones and write words 0.24 glyphs shorter, 0.05 and 0.24 on v1b_tail. The generated B quires hold 3,168 distinct words against the real 3,695. Reading the tables from hand 3's own pages does not mend it. Quire T's later twelve pages written from tables read from hand 3's earlier pages give the ten commonest words 20 to 22 per cent against the real 14, and from hand 2's pages 18 to 21. The profile is set by the sizes of the tables and the fill rule, not by whose pages fill them. The other B scribes are missed the same way. Hand 5's pages, held out on splits 1 and 3, give the ten commonest words 0.08 to 0.09 more of their text than the real ones. Hand 2's pages on splits 2 to 5, where the tables are read from hand 3, give them up to 0.02 less, the other way and small. So one profile shaped on hand 2 is judged against hand 3's flatter head, broader stock and longer words on split 1, and against hand 2's own heavier head on the other four. The fair-copy pass reaches the floor without mending this. With only quires H, N and T copied and the rest real, lo1_final_tail reads 1.12 against 1.18 uncopied, and with only the herbal and pharmaceutical quires copied 0.46 against 0.35. After the pass the generated hand-3 pages still give the ten commonest words 0.03 more of the text than the real ones. The pass gets to the floor by broadening the vocabulary on every page, which repairs the pooled count of once-used words. Split 1's figure is one B scribe's procedure judged on the other B scribe's pages. A second table set of hand 3's own breadth is the change that would fit it, and fitting it needs hand 3's pages among the tuning material, which the quire halves do not give. Such a test splits quire T by page and lies outside the quire protocol. A first test of that kind gives hand 3's pages a full second set, read from the earlier eleven pages of quire T in that hand and sharing nothing with hand 2's set. Its breadth was chosen on a page split inside the quire. It takes the whole half from 1.65 to 0.91 on v1b_tail and from 1.43 to 0.94 on lo1_final_tail, under the floor of 1.18. The figure is not clean. The choice saw part of the half, and quire T's later twelve pages, never used in the choice, stay at 1.67 against their own floor of 1.12. What remains there is the spread of word lengths and the first and last glyph of the word. Hand 3's own set comes to 2,823 cells, under the bound for one scribe, and both scribes' sets to 5,811.
That set, with its rows filled by frequency and hand 3's own habits at the line's edges, was then run on the whole text, its rows read from all 29 pages in that hand, and on every held-out half. On the whole text it keeps the battery where the one-scribe rows have it, with 12 of the 23 statistics within a tolerance on lo1_final_tail and 11 on v1b_tail, against 12 and 9 without it. It loses the profile of where words recur. On lo1_final_tail two of the six distance classes stay within tolerance and on v1b_tail none, against three and five without it. The page classifier tells its pages from the real ones 88 and 82 per cent of the time. Held out on all five halves it takes lo1_final_tail from 1.01 to 0.88 and v1b_tail from 0.97 to 0.79. The whole of that gain is split 1's, from 1.43 to 0.71 and from 1.65 to 0.76, where the set and the habits were read off quire T's earlier eleven pages inside the held-out half. Splits 2 and 4 hold no page in hand 3 and do not move. Hand 3's set comes to 2,862 cells on lo1_final_tail and 2,888 on v1b_tail, under the bound for one scribe, and both scribes' sets to 5,830 and 5,876.
Hand 3's habits alone do more of the work. Given only settings of the scribe's own that need no table, on hand 2's own tables, with a working sheet of 18 cells added and no table of hand 3's own, the hardest half moves. Hand 3 takes it from 1.65 to 0.98 on v1b_tail, under the floor of 1.18, and from 1.43 to 1.16 on lo1_final_tail. The rows carry 3,006 cells on v1b_tail, six over the bound, and 2,992 on lo1_final_tail. The battery on the whole text and the page classifier stay where the one-scribe rows had them. The profile of where words recur keeps five of its six classes on v1b_tail and five on lo1_final_tail, against five and three for the one-scribe rows. The settings are how often the sheet is refilled from the page, no exact copying from the line above, the length drawn first, the reach along the rows and the pen habits at the line's edges. They were tuned on quire T's earlier eleven pages inside the half and tested on the later twelve, which stay at 1.87 and 2.23 against their floor of 1.12. So most of what a second set bought on the hardest half is hand 3's habits, and the tables it added cost the profile of where words recur.
Tuned instead on hand 3's six pages of quire Q on the training side, with nothing of the held-out half used, the habits take v1b_tail from 1.65 to 1.49 on split 1. The floor is 1.18, and fresh seeds give 1.45. On lo1_final_tail they move it from 1.43 to 1.39, within seed noise. That tuning keeps one of the six distance classes of the recurrence profile on v1b_tail against the one-scribe row's five.
A third tuning holds the rates of joined, copied, exactly copied, new and chained words at hand 2's. It tunes only the 17 settings the two earlier tunings agree on, among them the sheet, the length drawn first, the reach along the rows and the pen habits at the line's edges, again on the six Q pages. That gives 1.34 on v1b_tail and 1.33 on lo1_final_tail against the floor of 1.18. It has 23 of the 27 statistics within two tolerances. It keeps all six classes of the recurrence profile on v1b_tail, and one on lo1_final_tail against the one-scribe row's three. That is the two-scribe figure inside the quire protocol. Under the standard cut it is 2.02 against the one-scribe row's 2.37. Over all five halves it is 0.91 against 0.97 under the scaled cut and 1.27 against 1.35 under the standard cut. Quire T's later twelve pages, never used to choose, stay at 2.09 and 2.41 against their floor of 1.12. Hand 3's ten commonest words take 0.154 of the generated text against the real 0.135. The rows carry 3,006 cells on v1b_tail, six over the bound, and 2,992 on lo1_final_tail.
What no setting reaches is the spread of word lengths on the stars pages. Drawing each word's length first and its body from every body of that length in the source is a ceiling at 621,554 cells, not a row. It leaves quire T's test pages at 3.30 and 3.32 against their floor of 1.12 and thins the ten commonest words to 0.083 and 0.089 of the text. The body rows re-filled by length give 2.07 and 1.73. A second set of middle rows at hand 3's rate of joined words gives 1.28 and 1.29, inside seed noise of the own set's 1.44 and 1.41. Under every middle the three-letter words are over the real share by 0.02 to 0.08 and the five-letter words under by up to 0.055. The words built from the rows, 42 per cent of the words on the test pages, come out 4.82 glyphs long against the real 5.21, because the rows hold the commonest bodies and those are short. The words copied with a slip write three-letter words at 0.125 to 0.134 of their number against the real 0.067. The new words, 13 per cent of the words at 6.83 glyphs when joined, are the one long source.
The assembly of the parts is not the place either. On the two-scribe figure inside the quire protocol, re-run over three seeds, quire T's test pages stand at 2.13 and hand 3's words there run 4.98 glyphs against the real 5.23. Giving one word in ten, five or three a second middle part brings the lengths nearer, to 5.04, 5.14 and 5.28 glyphs. But the glyph pairs move the other way, the ten commonest words on those pages thin from 0.162 to 0.147, 0.131 and 0.116 of the text, and the test pages worsen to 2.31, 2.42 and 2.59. Joining words by their parts trades the same way more slowly. Dropping the ending of one word in ten or five, writing it twice, a rule tying the head of a word to its body and a join in the tail were tried too, alone and in combination. They leave the pages between 2.02 and 2.79, inside seed noise of the row or worse. A longer word made from hand 2's parts carries hand 2's glyph pairs.
On hand 3's own set the copying step is not the place. An oracle that lands every copied word on the length the page needs puts the copies' three-letter share at the real 0.066. But it thins the ten commonest words to 0.147 and 0.153 of the text, and the test pages stand at 1.67 and 1.91 on the first five seeds against the row's 1.44 and 1.41. No copying rule the scribe could follow beats the row by more than seed noise. Nor is the fill of the rows. Reweighting every body of the rows to the shape of hand 3's tune pages, an oracle fill, gives 2.62 and 2.41 on ten seeds against the row's 1.45 and 1.49. The shape is reached, but the ten commonest words thin to 0.088 and 0.099 of the text and the glyph pairs go far off. Reweighting only the 30 commonest bodies gives 1.27 to 1.48 on ten seeds, keeping the head and the once-only count. The same fills stood at the floor on the first five seeds, 1.12 to 1.14, and at 1.41 to 1.58 on the next five. One fill's luck moves a page group by 0.3. Nor is the draw at the row. A marker walked down the row one cell a word, with no lot, and a marker moved by a die a word both land within 0.05 of the lot's figure on both rows over ten seeds. The whole text, the profile and the classifier are unchanged, and the marker makes no chance act at the draw. That is an option for the account of production: the rows once written carry the chance, and the lot at the row is not needed. On both footings the lengths are the grammar's own word shape, not a table's size, a copying rule, a fill or an assembly rule.
The page classifier is the other residue, and it is not hand 3's alone. Run page by page on v1b_tail over three seeds, it calls the real and the generated pages right 95 per cent of the time on hand 1's 112 pages. On hand 2's 46 pages it is right 76 per cent of the time and on hand 3's 31 pages 83 per cent. On the 128 pages of the herbal section it is right 96 per cent of the time. On the 19 pages of the biological section it calls nearly every page generated, real or not, so it is right 53 per cent of the time, which is chance. With hand 1's pages left out, it is still right 81 per cent of the time on the rest, so hand 1's pages, three fifths of the pages, weigh most in the whole figure but do not carry it. No one family of statistics carries it either. On a ranking score where 1 is a perfect sort and 0.5 is chance, dropping any one family leaves it between 0.89 and 0.91 against 0.91 with all of them.
What it reads is that a real page is more its own than a generated page. Across hand 1's pages, the spread of the glyph shares from page to page in the generated text is 0.61 of the real spread, and of the glyph pairs 0.66. Every generated page is a fresh draw from the same tables and habits, and the real pages differ from each other more than that. The two-scribe figure inside the quire protocol leaves this where it was, at 94 per cent on hand 1's pages and 81 per cent on hand 3's. A search over the 31 settings that need no table, with the battery held, found none that moves it beyond seed noise. Its end point, the best of 409 candidates, scores 0.90 against the row's 0.91 over three seeds on the ranking score, and hand 1's pages stay at 0.99.
On hand 1's 95 herbal pages, what differs from page to page is the words. The mix of word families, the aiin family against the e family and the ol words against the ch words, gives the main axes of the differences, and half the variation takes five axes against six for the generated pages. No pair of look-alike spellings, ch against sh or o against a, sits on opposite sides of any axis. The generated pages vary along the same families, since the families are the vocabulary's, but by less: on the ten family glyphs the real pages' spread from page to page is 1.8 to 2.8 times the generated pages'. Consecutive pages are no more alike than in the generated text, 0.87 against 0.88 of the distance between any two pages, so the differences are page by page and not a drift. Three constructs that would give a page a draw of its own were tried too, one setting of each re-run here on the whole text. A region of the tables read by each page, the rows sorted by their first glyph, raises the ranking score to 0.97 and hand 1's to 1.00. But the burst and repeat statistics go 10 and 8 tolerances off, since a page reading one region writes words that begin alike. The scribe's habits and rates drifting from page to page leave the score at 0.92, and a page's leaning toward some word beginnings and endings at 0.89, with hand 1's pages at 0.98 and 0.98. Under both the spread of the glyph shares from page to page is still 0.65 and 0.69 of the real, and the burst statistic is 6 and 7 tolerances off. None of them gives a page its own words. So the text leaves two residues. Hand 3's word shape on the stars pages is the grammar's floor. The classifier's residue is on every hand and every section but the biological, cleanest on hand 1's pages. It is each real page favouring its own words, and with them its glyphs and pairs, in a way a fresh draw from a fixed store does not.
The closest match at any size is R9Y, on the first line. It is the procedure above with a die on the ruler and three redraw rules. The die sends the copy to the word directly above half the time. The q is struck on half of the varied q-words. The paragraph base is drawn again while it begins with q, d, y, s or l, and the last word of a line is drawn again while it ends in o. It is the first row to match the whole text in full by the score file's rule: 23 of 23 within three tolerances, six of six recurrence classes and four of four gap measures within the tolerance. Held out by quire it scores 1.29, at the floor, ranked first of nine generators in every split, tuned on no split. But it needs 21,305 cells, 6,231 entries and 71 page sides, ten times the bound. It is still told apart page by page, at 0.800 against a null of 0.61, on the breadth of a page's glyphs and the word above. Its joint cloud test fails, with a distance of 183.8 where the largest among its own seeds is 54.6, and its lines' words carry no information on their length. By the whole rule, the quire halves, the page classifier and the gap statistics together, it is not a full match either. Beside the recipe read off the page it buys nothing held out and 0.025 on the classifier, for seven times the apparatus.
The two walls. Every row on every line failed held out on the same two statistics, the spread of word lengths and the share of words used once, and the reason for each is understood. The first wall was in part a fault of the check. Under the scaled counts described above, the spread of word lengths comes within three tolerances on every split for throw_s18start, and the miss halves for refined_s18 and for the automated search's part rows. The rest of it, and the whole of the second wall, belong to the procedures. A generator asked to write pages its tables were not built on has to invent about one word in six that it never made before, with the right length and shape. The mechanisms tried make the new words in two ways. Adding a glyph to a word in sight gets the count right and the shape wrong. The manuscript's once-only words are one glyph longer than a word already on their page 10 per cent of the time, exactly the rate of the same page with its words shuffled. The candidates that add glyphs show 20 to 56 per cent. Building the new word from parts drawn independently gets the shape too wide, and capping the length collapses the count. Raising the splice rate to fix the count makes the length spread the dominant miss, so the two only trade. The automated search stated the wall in its own terms. A table set small enough to own cannot make the rare words look right and make the right number of them at the same time, in any grammar it could reach. The derivation of the text from its own pages, described in the next subsection, reached the same conclusion by another road. It also says what the missing step must look like: a new word whose length is settled first and whose glyphs are then spelled in the text's own way, from nothing in sight. Two late rows of the automated search test that step. Drawing a word's length first and filling its parts to it fixes the shape of the once-only words on the held-out halves with no other change. The length-spread statistic goes from about +8 to about -3 on all five splits while the once-only count holds. Every earlier candidate drew its parts independently of one another, and that alone was the length-spread failure. The best held-out row with the length drawn first, fit_g1, scores 1.74 against the floor of 1.29, at 2,742 cells and 9.1 sides, with the once-only share still short by 6.6 tolerances. On the whole text it reaches 6 statistics within tolerance and 17 within three. The derivation's new-word step is a front and a back cut from the once-only words, with the length drawn first and no glyph added. Applied to 8 to 10 per cent of inner words, it brings the held-out once-only share from -7 to about 0 while keeping the length spread at -3, at 2,346 cells and 7.8 sides, under the cell bound. Held out it scores 2.05 overall. It is the first candidate to get the once-only count and shape right together. But it overshoots novelty, and it is not a feasible row. On the whole text it scores 4 within tolerance and 13 within three tolerances, with a misfit of 62.8. Both steps as written reject a draw and try again until the length matches. So the length-first row costs 9.34 draws a word and the new-word row 10.3, over the bound of about four random acts a word. The new-word row's one-glyph-neighbour statistics are far off too, the adjacent one-edit pairs by 5.6 tolerances and the neighbour statistic by 6.9, and its recurrence profile matches none of the six classes. Those two rows were scored under the fixed counts, and the automated search reports that the length-first row, tuned against them, overshoots the spread once the counts are scaled. The reason is that its new words are all distinct front-plus-back combinations. About three quarters of the real once-only words are one exchange from a word found elsewhere in the book, which the derivation assigns to the spelling habit. So the two walls are understood. Parts drawn to independent lengths spread the words too widely, and new words made without the neighbour structure of the vocabulary stand too far from the rest. The recipe read off the page in the next subsection carries the length-first step in a form that draws once. The length is thrown first and the word is then spelled to it from the glyph rows, with no retry, at one act for the length and one a glyph. Built into the copying frame with no ruler and one lot a word, the recipe reaches 12 statistics within tolerance and 22 within three at 2,983 cells and 4.75 acts a word, with the recurrence profile six of six. The step brings the spread of word lengths within tolerance on the whole text. It lifts the share of once-only words among paragraph-first words from 0.17 to 0.41, against the text's 0.48. The one miss beyond three tolerances is the line-end statistic, the information a word's last glyphs carry about its standing at the end of a line, because a new word written there carries no line-end habit. The share of once-used word types stays at 0.65 against the text's 0.68 whatever the step's rate, because the rows' plain draws still make most of the once-only words, as near-copies of the page. Held out by quire it scores 1.25 against the floor of 1.29 with nothing tuned, and the share of once-used words is its one statistic beyond three tolerances, on four of the five splits, at three to six short. The second line built the same two steps into its one-lot-a-word tables and tuned the rates on the quire splits in rotation. On the splits used in the tuning both walls come down, the spread of word lengths within two tolerances and the once-used share within three, at 1.01 to 1.07 under the scaled counts over all five. Under the clean protocol the once-used share comes within two tolerances on splits 2 and 3 and is short by five tolerances on split 1, while the spread of word lengths is beyond three on split 3. The step that repairs both together on pages the tables were not built on is not yet in hand, and the search for it is open. The two walls are therefore a limit of the steps tried, not yet an established limit of hand procedures. The first was in part a limit of the check, and the second is the once-only words.
The labels. One limit is shared by every candidate. The manuscript has 1,075 label words on 57 pages, single words beside plants, stars, nymphs and jars, and they come from the same glyph pool as the running text. Measured on ten statistics, the labels are longer than the running words, 5.21 units against 4.33 for a random running word. Of the labels 40 per cent occur nowhere else in the book, 57 per cent begin with o against 21 per cent of running words, and almost none begins with a gallows or with q. Four candidates were made to write labels from their own tables in every mode they allow: the procedure above, R8c, the shrunk copying frame and the part rows. None writes them. Their best modes put none or one of the ten measures inside the real interval. Their labels are too short, at 4.0 to 4.6 units, and too seldom new. Their paragraph-first words begin with a gallows 62 to 72 per cent of the time, where the labels almost never do. A glossary of plant names is no better a source. The headwords of the medieval synonym lists Alphita and Sinonoma Bartholomei run to seven and a quarter letters with a wide spread, and only two in five are as short as the label words. A full production account therefore still needs a label step, and none of the candidates has one.
What a full match would and would not establish. Suppose a candidate passed the whole text, the quire halves, the page classifier and the gap statistics under the bound. That would show that a scribe of the period, with a few sides of tables and a die, could have written a text with every measured property of this one. It would be a production account. It would not show that the manuscript was made that way, since other procedures might leave the same traces. It would not show what any word means, because a procedure of this kind carries no message. It would not fix the identity of the rare words, which no candidate reproduces, and it would not explain the section vocabularies, which were not tested. The tables of every candidate were filled from the manuscript's own words. So the search shows that such tables suffice, not how a scribe composed them, and no procedure was carried out by hand. What the search establishes as it stands is narrower and firm. The whole-text score of the procedure above can be had at a tenth of its apparatus. The four habits and the join are properties of the text that any account must carry. And the once-only words are the open problem.
Sources of error. The misses reported here could in principle be faults of the programs and not of the procedures, so the checks that bear on that are set out. The held-out miss on the once-used words appears in every candidate of three lines whose generators and scorers were written separately, and in the derivation, which is a fourth method. Two rows of the second line were run again with this line's scorer and gave every statistic to the last digit. The derivation code, run unchanged on seven candidates, reproduced the manuscript's own rows exactly, and the frame of the recipe reproduced the hybrid J3 with no difference on any statistic. Where a family's results file and its summary row disagreed, the file was taken. The four disagreements that remain are counts of apparatus and not scores: the H3 cell count with and without the sheet and the rule table, 18,776 against 17,976, and three roundings of page sides. Three errors were found and corrected during the search, and one of them touched the held-out scores. A cell counter on the second line omitted the rows of the middle swap, which moved interim_s11 from 2,825 to 3,725 cells. The same convention, applied to the rest of that line's rows, moved refined_s18 from 2,386 to 3,402 cells, over the bound, and throw_s18start from 2,070 to 2,100. A count of random acts for the rows with the length drawn first had been carried over unmeasured, at 5.5 where the measured figure is 9.34. And the fixed counts by which the second line's tables keep a word part, set on the whole text, starved the tables built from half the book. Scaled to the half, they move that line's held-out figures by 0.1 to 0.2, as set out above, and the once-used miss stays. The scores on the whole text were not touched by it. The transcription is a further source. Transcribers disagree on about 2 per cent of characters, and the verdicts on the generators were re-derived on a second transliteration, where the misses at issue are four to eleven tolerances wide on the held-out halves. The tolerances come from half-samples of the text's own pages. So whatever slips the scribe made and whatever the transcribers misread are already part of the spread they measure, while the two misses fall on the same measures in every quire split. So the walls are properties of the procedures tried, and not of the programs.
Earlier work. Rugg (2004) and Rugg and Taylor (2017) proposed a table read through a Cardan grille. Timm and Schinner (2020) proposed a self-citation method in which each word is copied with changes from the page. Both are hand methods, and each reproduces some properties of the text. Zandbergen (2021) described wheels of word fragments and gave no statistics for them. Greshko (2025) published the Naibbe cipher, in which a card drawn from a shuffled deck chooses which of six tables spells each piece of the plaintext. It is the one published key of the period's kind whose sign variants are chosen by chance, and the lot-book and dice-table families here read their tables the same way, though they carry no plaintext. Bolte (1903) surveyed the lot books and Wozniak (2025) the wax tablets on which the size bound and the working sheet rest. New here are the bound on the apparatus as part of feasibility and the count of cells, entries and sides beside every score. New too are the search over some forty families of smaller procedure, the kinds of smaller procedure only, on four independent lines, the four habits and the join share read off the text, and the held-out rule with the tuned splits named. So are the two rows at the floor under the bound, the closest at any size, the two walls with the part of one that was the check, the simplest device and its ladder, and the label test.