Beinecke MS 408 · independent examination
Voynich Text Examination

What the text of the Voynich manuscript is like, measured on the page scans and on full transcriptions, and compared with real writing in thirty-one languages.

One bundle of sheets

Layfield and Davis found that the two halves of a folded sheet use words more alike than neighbouring pages do. They conclude the book was meant to be read sheet by sheet. Colin Layfield and Lisa Fagin Davis published that in July 2026. It is the one measured fact on these pages about how the book was made. Their paper is ‘Singulion Structure and the Voynich Manuscript’, in Digital Medievalist 19. It went online on 10 July. It is at journals.openedition.org/digitalmedievalist/2331.

This search came to the same result on its own, without having read their paper. That is because it was asked to work from the manuscript alone. The words ‘this search’ and the ‘we’ of this page mean computer programs. They took the measurements and wrote these pages. No person has checked these numbers by hand. The checking that exists is programs re-running each other's work from the manuscript's own text. This page's main result was re-run from the raw words. Every figure came out identical. This page gives their result and the result found here. It gives every check made on it, the results retired on the way, and what is still being tested.

Every figure names the list of words it was scored on. The size of the effect depends on that list. Every figure on this page is also listed with its source in the file figures.tsv beside this page.

One question stood over every figure on this page that says how far beyond chance a result is. The question is whether the sets of page pairs drawn to stand in for chance understate what chance can do. It has been tested against eighteen stretches of real prose. The section on the deepest limit gives the answer in two halves. The spread is how much a figure moves about by chance when the pairs are shuffled. The spread comes out right.

The other half is the bad one. Ordinary prose came up by itself to the level once called five in ten thousand. It did so once in eighteen tries. So every such figure on this page is given as a range where one has been computed. The ranks against real prose come first, because they need no set drawn to stand in for chance. The chance figure behind the headline result has since been tested against real prose too. The section on the deepest limit gives that test.

What a sheet is

A book of this age is made of sheets of parchment folded in half. Fold one sheet and you get two leaves, four pages. Stack several folded sheets inside each other, sew them along the fold, and you have a bundle.

In a bundle the two halves of one sheet end up far apart. The outer sheet gives the first leaf and the last leaf of the bundle. The innermost sheet gives the two leaves in the middle. Those two face each other. The manuscript has about twenty bundles. This page is about the last complete bundle of ordinary sheets. Its leaves are numbered 103 to 116. This page names a sheet by its two leaves, such as the sheet 104 and 115.

The published result

Layfield and Davis measured how alike two pages are in the words they use. Their method is called latent semantic analysis. It drops the words used only once. For each sheet they took one pair of pages. That is the back of the first leaf against the front of the second. Across the book those pairs score 0.433 on their scale. Pages that face each other in the binding score 0.298. In the last bundle the same-sheet pairs score 0.600. The facing pages there score 0.438.

Their comparison makes no allowance for how far apart two pages are in the book. It has no shuffle test. A shuffle test shuffles which pages count as the same sheet, many times over. It asks how often chance alone gives as large a figure as the real folding does.

They judge whether their result is beyond chance by a test between two groups of pairs. One group is about 38 same-sheet pairs. The other is about 67 facing pairs. All scribes and all sections are thrown in together. The middle sheet is kept in. Their check is two real bound manuscripts, the Aberdeen Bestiary and a Speculum Humanae Salvationis. Each is scored in its own arrangement. In both, the same-sheet pairs come out lowest.

Their conclusion is that the book was a stack of single-sheet bundles written one sheet at a time. They also propose an order for the sheets of the last bundle. The section on their order below takes it up. The text their paper analyses leaves out the last page of the bundle. So do the tests here.

Why their paper mattered had been said by someone else four years before it appeared. Nick Pelling wrote in December 2022:

I'm unaware if there are any studies that have specifically used page-to-page word list difference metrics (of the kind Julian Bunn did) to exhaustively evaluate all the different nesting permutations, i.e. to suggest what the original nesting order was

That is the gap their paper filled in 2026.

The result found here

In this one bundle, take the pages that were the two halves of one folded sheet. They share more of the book's ordinary middling-common words than any other way of folding those same leaves would give. The effect is not found on the bundle's own words. Those are the words common inside it and rare in the 182 pages outside it. And it is not found in the other bundle we tested. And no text our own programs write comes near it.

An earlier wording of this page said the bundle's own words were the carriers. That was wrong. The section on the words says how it was found to be wrong. The bundle's own words are one part of the carrier. They are the smaller part.

The deciding test comes first, because nothing in it is drawn at random. The bundle's leaves can be paired off with each other in 120 ways. The test ranks the real folding against all of them on the same measure. Here is how the 120 are built.

The bundle has fourteen leaves. They are numbered 103 to 116. Two of them are lost from the manuscript, 109 and 110. Two more are set aside. Leaf 116 has only one surviving side. Leaf 103 is its sheet-mate. Their sheet can therefore never be measured. Ten leaves are left. They make five sheets. Five sit in the front half. Those are leaves 104 to 108. Five sit in the back half. Those are leaves 111 to 115.

In a bundle of nested sheets, every sheet joins one front leaf to one back leaf. So there are 120 ways to fold them. The first front leaf can pair with any of five back leaves. The next can pair with any of the four left. The next has three to choose from. Then two, and then one. Multiplied together, those give the 120. The true folding is the nested one. Its outermost sheet is 104 and 115. Next comes 105 and 114. Then 106 and 113. Then 107 and 112. The middle sheet is 108 and 111.

The narrower run leaves out the middle sheet, 108 and 111. Its two leaves sit closest together. No comparable pair at that distance exists to measure it against. Four leaves remain on each side. They can be paired in 24 ways.

We did not establish the folding ourselves. We took it from the published description of the manuscript. That description agrees with pairing by position on 148 of 148 pairs. Then we asked whether the manuscript's words prefer it over the other 119 ways. The count of 120 assumes the bundle is nested front leaf to back leaf. If a sheet could join any two leaves at all, there would be 945 ways instead.

The number 120 means two different things on this page. There are 120 ways of folding. Separately, there are 120 pairs of pages that get compared. The two have nothing to do with each other. Their being equal is a coincidence. Here is where the 120 pairs come from. There are six front leaves. Each is set against each of five back leaves. Each pair of leaves gives four page pairs, one for each way their sides meet.

Twelve leaves survive. They have 24 sides. One side is missing, which leaves 23 pages. There are four measurable sheets. Each gives four same-sheet page pairs, one for each way its sides meet. That makes the sixteen same-sheet pairs. The thirty leaf-pair figures are one for each front leaf against each back leaf. Each is the sum of its four side pairs.

In the deciding test no set of pairs is drawn at random. So nothing can go wrong the way the matched comparison sets did. On the words where the effect lives, the real folding comes first of 120. On the narrower run it comes first of 24. One in 120 is the strongest thing this test can say. The test says it. One in 24 is the strongest it can say on the narrower run. It says that too.

On the bundle's own list of words, chosen a different way, the real folding comes eleventh of 120. On the narrower run it comes fourth of 24. So the list never supported the fold on its own by this test. The finding lives on one choice of words and dies on the other. The 120 pairings were then counted a second time by a separate program, from the same files. The ranks came out the same.

The rank says 120 rivals and none higher. One more figure belongs beside it. Take the average over all 120 pairings. It is +0.51 on the test's scale. On that scale a larger figure means the paired leaves share more of those words. The real folding's figure is +2.81.

The rank rests on the outer three sheets. Above all it rests on the sheet 106 and 113. That sheet gives the largest of the thirty leaf-pair figures the test is built from. The two innermost sheets are a tie.

Only two of the 120 pairings carry those outer sheets. The real folding and the runner-up differ only in how the two innermost leaves are paired. The gap between them is about a twentieth of how much that figure moves about by chance on that swap. On the narrower run the gap to second place is about six tenths of that movement. So the outer three sheets are settled. The two inner ones are not.

Three caveats belong here and not below. The first is how close the runner-up is. The arrangements that come closest are the real one with one or two inner sheets swapped. They come in at +2.738 on the same scale. The real folding is at +2.812. So the second place is our own answer lightly disturbed.

The second caveat is that the comparison runs across different page distances. With each leaf pair's usual figure for its distance apart subtracted, the ranking was run again. The real folding is then second of 120 on the common words. So the distances do not make the rank on their own.

The third is that this run was not blind. The per-sheet figures were in hand. The prediction was written from them. Once the measure is fixed, the ranks are fixed too. That is not thought to touch the result. What it could touch is the choice of test and of pairings to rank against. Leaving the first and last leaves of the bundle out of the pairings does not help the real folding either. The first leaf's figures are poor, so pairings that use it rank low.

One more thing belongs here. It counts for the finding. The earlier wrong-fold result fell apart because of a matching rule. Pairs were compared only against other pairs at the same page distance. At one distance that rule left a single comparison. So one number appeared twice with its sign flipped.

The deciding test does no matching at all. Every sheet pair is compared against one and the same set. That set is every pair of leaves that is not a sheet. So it cannot fail in that way. The test was built without matching because of what matching had just done. Nothing more is claimed from that than that the one fault found cannot be present here.

The evidence that needs no reckoning of chance comes next. Three real books were written out into the manuscript's own page and line lengths. It is as if a scribe had copied them into this book. That was done from a different starting point each time. Each start is one placement. A text written out that way is called poured on this page.

There are twenty placements each of the King James Bible, Don Quixote and Montaigne. That is sixty placements in all. This page calls that check the sixty placements. Twenty of them are of Montaigne. That text is struck. Its file held the original and a modern rendering one after the other. So forty placements stand. But the counts given here include all sixty.

On each placement the same thing was measured as on the manuscript. That is how many more words the two halves of a sheet share than their comparison pairs do. This page calls that extra sharing the excess.

The manuscript's excess stands above all sixty placements. That holds under all five ways of choosing the words. It does so on both of its comparison levels and under both of its ways of scoring. Not one of the sixty comes up to it under any reading. That is the one statement about these placements that survives every reading with no exception. It is the one this page leads with.

A comparison level is the figure the comparison pairs sit at. The pooled level takes all of them together. The matched level takes only those matched to each sheet pair.

The first three lists are drawn from the front, the back and the whole of the bundle. Two more came later. On the pooled comparison level the front list gives the manuscript +1.47 on that scale. The back list gives +2.10. The whole-bundle list gives +1.55. On the matched comparison level the front list gives +2.49. The back list gives +2.33. The whole-bundle list gives +1.85. On that scale a larger figure means the two halves of a sheet share more.

A rank against real prose needs no comparison set. So nothing in the question stated below can flatter it. Every such count can move by about one placement in twenty. That happens when the text is written out starting a little earlier or later. The section on real texts says how that was measured.

The lowest value a count over the standing placements can report is not one in forty-one. Twenty placements of one text are that many different stretches. The words that carry the figures overlap by only about an eighth from one stretch to the next. But two texts are still only two texts.

On the excess the standing texts sit at about one level, as far as two texts can show. So the lowest value that count can report lies between one in twenty-one and one in forty-one. This page prints that range and never one in forty-one alone. How many separate tries forty placements of two texts amount to was not worked out. This page claims no number for it.

One more limit rides with every count over the sixty. The King James Bible uses far fewer different words on those pages than the other two texts. It has 1,163 distinct words a placement. The other two have 2,739 and 2,461. The manuscript has 3,043. And King James placements come closest to the manuscript on one of the five measures below.

Four more measures of that kind were made on the same placements. Each count below names the list of words it belongs to. A reading here is one list of words under one way of scoring. So the original lists give six readings.

Two of the measures below are ones where a low figure counts for the finding. On the pages that face each other the manuscript's figure is a low one. So is its figure on pairs at odd distances apart. A placement comes up to it there by sitting at or below it. On the gap between sheet mates and facing pages the manuscript's figure is a high one. There a placement comes up to it by sitting at or above it.

On the pages that face each other in the binding, the three texts differ in level. So the statement is per text. There the manuscript sits below all twenty placements of each text. That holds under four of five lists. Under the fifth it sits below all twenty placements of two of the texts. Of Montaigne's it sits below nineteen of twenty.

Under the front list the manuscript's figure for those pages is −0.14 on that scale. Under the back list it is −0.05. Under the whole-bundle list it is −0.42. Every prose placement runs positive there, and many run above +3.

The gap between the manuscript's same-sheet pairs and its facing pairs is the next measure. Under the two further lists it stands above all sixty placements. Under the original lists it stands above 59 of 60 on every reading but one. That is five of six readings. On the one left, it stands above 58 of 60. The placements that come up to it there are King James placements. One does under every reading, and a second under one.

The manuscript's figure for pairs at odd distances apart is the next measure. Under three readings it sits below all placements but one. That is 59 of 60. Under two readings it sits below 58 of 60. Under one it sits below 57 of 60. One King James placement comes up to it under every reading. Beside it, three placements of Don Quixote do so under one reading each.

On the matched comparison level under the first list, the three texts differ in level for the excess too. There the statement is per text. The manuscript stands above all twenty placements of each text.

Those two rankings are what this page stands behind. One is the real folding at one in 120. The other is the manuscript above all sixty placements. The finding this page leads with is not a probability at all. It is a ranking of the real folding against all 120 ways of pairing those same leaves. All 120 are counted out one by one. Nothing in it is shuffled and nothing is drawn. That is why this page can carry a probability beside it without the finding depending on it.

A drawn comparison set is a set of page pairs drawn at random to stand in for chance. Every probability built on one comes after the two rankings, with its calibration attached. Calibration here means checking the comparison set against real prose, where no sheet can matter. That shows how far chance alone can go.

On the book's own middling-common words, take the sixteen same-sheet pairs. They share +5.29 more words a pair than their comparison pairs do. On the shuffle test, one shuffle in 200,000 matched it. On a second run with a different shuffle, none did. So the figure lies somewhere below about one in a hundred thousand. And 200,000 shuffles cannot put it more finely than that. The question in the section on the deepest limit is attached to it.

That figure has since been tested against real prose. The section on the deepest limit gives the test. Real prose poured into the same pages and scored the same way gets about halfway to the manuscript's figure. It gets no further. The test flatters itself mildly, by about a fifth. It keeps its raw value only if the wrong-fold test is ruled exempt. Whether it is exempt is still being decided. The calibration in that section has not been applied to it.

Folded one leaf early, the figure is −2.03 shared words a pair. Folded one leaf late, it is −0.58. What that collapse is made of is in the section on the deepest limit. It is less than it looks. Twenty artificial texts were written into the same pages. The manuscript sits seven and a half spreads above their average. None comes up to it. That says only that our own programs do not produce the effect. It does not calibrate the shuffle figure.

There are two fair ways to compare the manuscript with a real text poured into its shape. This page prints both together and never one alone. The first is the size of the effect, which is what the finding is about. On it the manuscript stands above every poured placement. That holds on every test that has made the comparison on the same kind of words.

Three tests have made that comparison. None of the eighteen stretches on the comparison matched on distance comes up to it. None of the sixty placements does. And none of the sixty on the test of the headline's chance figure does. Each of those sixties has twenty placements of the struck Montaigne file among them. So forty stand in each. The manuscript is above all forty in both.

Those two sixties are two different tests run on overlapping texts. Both included Montaigne and the King James Bible. One adds Don Quixote where the other adds Caesar. So they are not two clean shots at the same target.

On each text's own middling-common words, the answer is the same. Five books were written out into the pages. That made a hundred placements. Twenty of them were of a text we later struck. So eighty stand. The manuscript is above all eighty under both readings. The nearest comes in at +3.230 on that scale. The manuscript's figure is +5.328.

On the other list of words, the bundle's own, the counts run the other way. Livy was written out into the pages four different ways. Its placements come up to the manuscript anywhere from none in twenty to five in 45. One placement in twenty of a fourth text does too. Two lists, two answers, and both are on this page with their lists named.

The raw size of the effect depends on how many different words a text uses. The poured texts differ in that. So the fairer scale measures each text against its own noise and not against the others. On that scale the manuscript stands above every placement on two of those three tests. On the third, the comparison matched on distance, one stretch of Manzoni comes to a near tie. It stands at 3.03 spreads. The manuscript stands at 3.22.

Neither of those comparisons leads this page. The ranking of the 120 pairings leads, because it compares the manuscript with itself. It uses the same words and the same pages, with nothing poured in and nothing shuffled. So the question of scaling one text against another never arises in it. The headline result does not depend on any comparison with other texts. The comparisons with other texts answer a different question.

One limit on the sixty placements must be said plainly. Their figures are as one program reported them. No second program has read them from its files. The eighteen stretches were read from files by a second program. So were the sixty on the test of the headline's chance figure. So one of the supporting counts on this page rests on a single report.

The book's middling-common words were drawn up two ways, by different routes and by different rules. One list has three hundred and ten words. The other has a hundred and ten. A hundred of the smaller list's words sit inside the larger one. So the smaller list is almost wholly contained in the larger. The larger carries two hundred and ten words the smaller leaves out. The two counts of the manuscript's pages agree to the word. Both count the same 10,820 words. Both count the same 23 pages.

Both lists got the same answer against poured prose. On one list, five books were written out into the pages. That made a hundred placements. Twenty of them were of a text we later struck. So eighty stand. The manuscript is above all eighty under both readings. The other list was used against sixty placements. On it the manuscript is above all forty that stand. That is agreement between two different tests built on overlapping material. It is not two routes arriving at the same list. Neither list was made from the other.

Among the matched comparisons, the strongest figure is from the comparison matched on distance. It compares each pair of pages from one sheet only with pairs the same distance apart in the book. It also requires the same fronts and backs. Pairs made of the front and back of one leaf are a separate class in it. So it is about sheets and not about leaves.

On its list of words, the same-sheet pairs share 7.39 more distinct words than their matched comparison pairs do. They stand above all twenty samples of every kind of artificial text written into the same pages. That holds in all six of its comparison groups.

A shuffle of which pages count as the same sheet had put the figure at about five in ten thousand. Calibrated against eighteen stretches of real prose, that becomes a range and not a number. The section on the deepest limit gives how. By rank against those texts it is about one in eighteen. The range on that runs from roughly one in a thousand to one in four. That rank is the widest statement of it.

The part of the comparison's error that can be accounted for gives a narrower range. That is about one in four hundred to one in seven hundred and fifty. The part that cannot be accounted for could take it further. Those are not collapsed into one number. Under our own artificial texts' spread it would have been seven in a hundred thousand. That is the number this page would have published if it had only ever checked itself against our own machinery.

The three kinds of pair below were scored on the list of words used by the comparison matched on distance. The explanation of those kinds was withdrawn. That withdrawal is a different comparison. So it does not weaken this figure. But the figure comes from a part of the work that did not survive untouched.

The comparison matched on distance also ranked the manuscript's raw figure against all 43 of its texts, real and artificial. The raw figure is the one before any shuffle is subtracted. The manuscript comes first in all four of its comparison groups. Its figure is +7.39 more shared distinct words a pair. The best stretch of real prose, one stretch of Manzoni, gives +4.12. The best artificial text gives +1.34. The gap between the manuscript and the best real stretch is 1.9 times the manuscript's own spread.

Five of that comparison's texts are independent of each other. Against those, the lowest value that ranking can report is one in six. Against all eighteen stretches of prose it is a range. The range is between one in six and one in nineteen. Against the artificial texts it is one in twenty-six.

That test adds a correction of its own. The average of the manuscript's own shuffles is +1.93 and not zero. So its excess over that average is 5.46. That is why one Manzoni stretch comes up to the manuscript's shuffle probability. No text comes up to its raw figure.

One test shuffles which pages are called the same sheet. This page calls it the shuffle test. It scores the 122 words common inside the bundle and rare outside it. Which words those are depends on which outside pages are used to decide it. The run tried several ways of cutting the same 182 pages.

The manuscript's figure was matched by eleven in 200,000 shuffles. A second run with a different shuffle matched it seven times. That is a raw shuffle figure. The calibration in the section on the deepest limit has not yet been applied to it. Under every way of cutting the outside pages, the chance was between two and twenty in ten thousand.

The same shuffle run on the book's own middling-common words comes out stronger. There the figure came up once in 200,000 shuffles. On a second run with a different shuffle it did not come up at all. And that is on fewer occurrences of the words. So again the figure lies somewhere below about one in a hundred thousand. That is the measurement that changed what this page says the carrier is.

The same-sheet pairs were also scored under five lists of words. This page calls that check the five lists. The lists were drawn from five places. They were the front of the bundle, the back of it and the whole of it. The fourth was the words common to front and back. The fifth was the rest of the book. The effect is positive under all five. It runs between +1.47 and +3.82 on its scale. On that scale a larger figure means more sharing. Its size depends on the list. Its existence does not.

The deepest limit

One limit stands over everything on this page. It belongs here and not at the back. It is about how a result was judged to be more than chance. It was found on a comparison that rested on eighteen pairs of pages. This page calls that the eighteen-pair comparison. Here is the limit, stated in full.

To judge whether the pages that were folded together as one sheet share more words than chance would give, we compared the sheet-mates against many sets of page pairs drawn at random from the same gathering, and the sheet-mates came out above almost all of them. That comparison rests on an assumption: that a set of page pairs drawn at random is a fair stand-in for what chance produces. We tested the assumption by pouring ordinary printed books into the same page layout, where no sheet can matter, and asking the same question of them. The books' sheet-mates did not sit in the middle of the random sets, as they should have if the assumption held; they wandered about twice as far from the middle, in both directions, as the random sets said was possible. So for a text that flows like ordinary prose, the random sets understate how much a group of page pairs can differ by accident, by roughly a factor of two. The real pages' distance from the random sets is therefore large if the real pages behave like our generated texts, whose sheet-mates do sit in the middle, and unremarkable if they behave like ordinary prose. Which of those the real pages do cannot be told from these pages, and until it can be told, the sheet-mates' standing is a number on the record and not a finding.

Why a drawn comparison can mislead is understood. Here is the reason. The real group of pairs is clustered. It is five blocks of pairs. Each block is the four pairs between the same two leaves. A drawn group is scattered across many leaf pairs. Prose has effects that run across two leaves. A subject runs for pages and comes back. So the real group feels such an effect four times over. A drawn group feels it once.

That is why the drawn comparison understates how far real prose drifts. It understates it by something between about a third again and three times.

Our own artificial texts carry such block effects too. Their block share is 0.148. The poured texts' block share is 0.221. The rule that makes a leaf's two pages alike also makes their pairings with other leaves move together. So the clean results on artificial texts did not come from their lacking the structure. The block effect accounts for a factor of about 1.3 to 1.4 in either kind of text. That is not the whole of the gap between them. The rest of the gap is unexplained.

Take the comparison matched on distance. There the factor is applied to a figure near 3.9 spreads from the comparison's middle. Apply the factor of 1.3 to 1.4 there. The figure becomes about one in four hundred to one in seven hundred and fifty. Suppose the factor were 2 instead. Then the figure would have been one in forty.

The part that cannot be accounted for could take it further. The rank against the eighteen poured stretches of prose remains the widest statement. It is about one in eighteen, with a broad range. Those three are not collapsed into one number.

The same doubt was then taken apart under four more comparisons. They were built on the same pairs as the eighteen-pair comparison. Each keeps a different part of how the pages are arranged. One of the four keeps both the balance of pages and the blocks of leaves. Under it the poured texts' figures spread 1.31 times as far as the comparison says. That figure is uncertain by 0.54 either way. The artificial texts' figures spread 0.98 times as far. That figure is uncertain by 0.16. On that comparison the manuscript sits 2.00 spreads above the middle.

A re-pairing test was also run on ten leaves, with the two centre leaves left out. It put the arrangement under test fifth of 120. The same-sheet claim from the eighteen-pair comparison stays withdrawn. The deciding test ranks the figures as measured. This comparison takes the page balance and the blocks out first. In one sentence:

the factor stands as measured, 1.3 to 3 on four values; the page effect and the blocks account for the generated side of the gap; what is left is inside the four-pour error and unexplained.

Generated texts there are the artificial texts of this page. The four-pour error is the uncertainty on the poured texts' factor.

The wrong-fold test is the clearest example on this page of a result that looks decisive. It is partly a property of the comparison. In this bundle, at one page spacing, two sheets are each other's only comparison. They are the true sheet and the sheet made by folding one leaf early. So the true sheet scored against its comparison gives +38.5 shared words. The early fold's sheet scored against its comparison gives −38.5. That is the same number with its sign reversed.

That one comparison is 46 percent of the true fold's total on the book's common words. On the bundle's own list it is 56 percent. It is more than the whole of the early fold's negative.

Take that comparison out. Three sheets remain. On the book's common words they give +3.8 shared words per pair for the true fold. The early fold gives +0.5. On the bundle's own list the true fold gives +1.9. The early fold gives +0.4. About half the collapse under a wrong fold is the true result's largest single term, seen from the other side. It was always going to look like a collapse.

Its independent content is three sheets on each side. The true fold is ahead there, carried mostly by one sheet. The figure at a shuffle probability of 0.049, given in the section on the wrong fold, has the same shape. It is one comparison group against one, at a different spacing.

Every probability attached to those figures comes from the drawn-comparison machinery. That machinery's calibration is in question. The twenty artificial texts do not calibrate it. They say only that our own programs do not produce the effect. The test that replaces it ranks the real folding against all 120 ways of pairing those leaves with each other. On the narrower run it is all 24 ways. Both use the same measure. Every one of those pairings is the same kind of object as the real one. So nothing is drawn. The doubt cannot touch it.

The section on the result found here gives that test. On the words where the effect lives the real folding comes first of 120. On the bundle's own list it comes eleventh.

The calibration is this. Eighteen stretches of real prose were poured into the bundle's own page lengths and page order. The shuffle test was run on each exactly as on the manuscript. Twenty-five artificial texts went through the same. Each had a hundred thousand shuffles. They were run in four comparison groups.

The good half first. The feared widening of the spread is not there. The poured texts' results spread almost exactly as the comparison says they should. The factor is 1.03, uncertain by about a sixth. That is not the 1.5 or 2 that was proposed. By the letter of the rule set beforehand, the test reads as exempt in practice in all four groups. Take the five main poured texts. None sits in the lowest five in a hundred of the comparison values. None sits in the highest five in a hundred either.

The bad half is worse than the factor would have been. One of the eighteen is a stretch of Manzoni with no folded sheets in it at all. It came out at a shuffle probability of 0.00062. That is about six in ten thousand. The manuscript's own level is 0.00074. That is about seven in ten thousand. This page had called that level five in ten thousand. Ordinary prose came up to it by itself, once in eighteen tries.

There is also a tilt the spread does not capture. Twelve of eighteen poured stretches have their same-sheet group raised. The artificial texts lean the other way.

So the figure is given as a range, with the accounts that produce it, and not as a single number. Several figures follow for one result. The one this page stands on is the first, the rank. The rest show how the same result looks under each account of the comparison's error.

By rank, one of eighteen poured stretches is at or below the manuscript. That is about one in eighteen. Its range runs from roughly one in a thousand to one in four. A single event in eighteen fixes very little. By the measured factor alone, it is about one in a thousand. With the poured texts' tilt taken out as well, it is about one in two hundred. Under the artificial texts' own factor, it is seven in a hundred thousand.

Holding page distance and side fixed does not make a comparison immune to this doubt. That had been argued for the comparison matched on distance. The argument is withdrawn, on a count. The shuffle there draws groups of 21 pairs. Such a group covers about sixteen distinct blocks of leaves. The real pairs sit in six. So the clustering doubt applies to that comparison too. The other shuffle figures on this page have not had the same calibration applied. Each says so where it stands.

What has been calibrated and what has not is stated here in one place. How far one test is flattered in judging chance, by the way its comparison groups are built, has been measured. The same test was run on sixty placements of real texts. Forty of them stand. The struck text's twenty are out of the count. The answer is mildly.

The shuffled comparison's spread is understated by a factor of about 1.1 to 1.25 over the standing texts. The struck text's own figure is 0.90. It is out of the count along with its placements.

Take the level of one in twenty. Four of forty placements came out at it, where two would be expected by chance. One came out at the level of one in a hundred. And one placement passes the line the test itself sets, without coming up to the manuscript. A second figure came with it. Within one text, the comparison understates how much the number moves from one placement to another. It does so by about a third.

That measurement covers the test that works on each placement's own uncommon words. That is where the figures on the sixteen same-sheet pairs come from. It does not cover the common-word result, the one this page rests on. That result was tested separately. Its test comes next. So the uncommon-word test's judging of chance has been measured. It is mildly flattered.

And in that run of the uncommon-word test, the manuscript does not stand out. Its own figure comes out at about one in twelve, which is nothing. Five of sixty poured placements of real texts come up to it or beat it. The shuffle test on the bundle's own words itself gave eleven in 200,000 shuffles. A second run gave seven. That one in twelve and the shuffle test's figures have not yet been set side by side.

The chance figure behind the headline result is the shuffle count on the book's common words. That is one match in 200,000 shuffles, and none on a second run. It has been tested against real prose in the same way. Three real texts were each written out into the manuscript's own page and line lengths. Each was written out at twenty different starting points. That is sixty placements in all. Twenty of sixty are of the Montaigne file later struck. So forty stand.

The figures over all sixty are kept here. The standing counts follow. For each placement the same kind of word list the finding uses was drawn. That is the middling-common words of that stretch of that text. Then the same test was run on each. Not one of the sixty came close. The largest figure among them was +2.762 shared words per pair. The manuscript's is +5.29. So the largest is about half of it.

None came up to the manuscript's standing against its own shuffles either. The largest stood 2.98 spreads above the middle of its own shuffles. The manuscript's figure is 4.55. The best chance figure any of them produced was one in seven hundred and sixty. The manuscript's is one match in 200,000.

The test does flatter itself, mildly. Its shuffled comparison moves about a fifth less than it should. Take the level of one in twenty. Four of forty standing placements came out at it, where two would be expected. One came out at the level of one in a hundred. And one placement passes the line the test itself sets, without coming up to the manuscript. The flattering is carried by one of the two standing texts, the King James Bible. It is also the text that comes closest on every other measure.

The lowest value that test can report is this. The manuscript is first of twenty-one within each standing text. Over all the standing placements together it is first of forty-one. The forty are not so many separate tries. So the statement is between one in twenty-one and one in forty-one. It is not narrowed. That comes from the same counting as the sixty placements' range in the section on the result found here. It is not a second measurement agreeing with it.

Two texts at twenty placements each give those two ranks by construction, whatever the texts are. So the deepest limit on this page is not that the headline's chance figure is unmeasured. It is that the figure is mildly flattered, by about a fifth. And real prose poured into the same shape gets about halfway to the manuscript and no further.

The last sentence of the limit quoted above is about its own comparison and about nothing else. Does the same limit cover every comparison built the same way? Every shuffle figure on this page is such a comparison. The eleven in 200,000 on the bundle's own words is among them. So is the one in 200,000 on the common words. What has been measured of that is given above in this section.

Whether the wrong-fold test is exempt is an open question. It compares the same pairs under a different folding and not a drawn set.

The deciding test re-pairs the same 23 pages instead of drawing other pages. It pairs them in all 120 possible ways. Poured prose reads as ordinary on it. No matched comparison in this search has managed that. The section on the result found here gives its figure: the real folding first of 120.

Sheet by sheet

Two of our own tests took the result apart by sheet. They disagree. They used different lists of words and different comparisons. The shuffle test, on the bundle's own frequent words, scored sixteen pairs of pages. Those pairs cover four sheets. It found no sheet more than 1.4 spreads of its own noise away from an even effect. An even effect means the same effect on every sheet.

The comparison matched on distance, on its own list of words, found the opposite. As first built, under one way of matching, one sheet sits nearly four spreads above an even effect. Another sits nearly five below it. Under the other way of matching, one figure is 2.81 above. The other is 2.57 below. The unevenness across the four sheets comes out as far from even as that test can report.

So it cannot be said which of the sheets carry the effect. The reason is not that one test found nothing. It is that two of our own tests disagree. No third test has been run to break the tie. Any third test would be chosen knowing what answer was wanted. Two tests on two lists of words can differ. This page's own argument about noise says why. The disagreement is a finding about how hard the question is.

The result of the comparison matched on distance comes with its own limits, all of them against it. It is sensitive to page length. With page length taken out, it survives under one way of matching. It vanishes under the other, where every sheet is ordinary.

Two descriptions of single sheets made on the way are gone. One sheet had been described as also carrying the effect. On the comparison matched on distance it is not supported at all. It sits higher than 95 in a hundred of its comparison's values under one way of matching. Under the other it sits higher than 72 in a hundred. Another sheet had been called flat. Its flatness is the comparison's own unevenness.

On the book's own common words, all four sheets are positive. That includes the one that was negative on the bundle's own list. That is one more reason not to say the effect belongs to particular sheets.

What can be said is this. The sixteen pairs together carry the effect on the shuffle test. On the bundle's own words that is eleven in 200,000 shuffles. On the book's common words it is one match in 200,000, and none on a second run. It is spread across the pairs of at least three sheets. And on both lists of words it falls under a wrong fold. The section on the deepest limit takes that fall apart.

The shuffle test's figures by sheet are in the table below. They are a description of its run and not evidence of anything. Beside each sheet's observed excess is what an even effect would put on it. That allows for its page sizes and its share of the words. The four sheets differ widely in the first column. All four are within one and a half of their own noise of that even effect. The amount of writing on the four is nearly even. So it is not simply that one sheet has more text on it.

The differences are in how many of the bundle's own words sit on each sheet's second leaf. They are also in how many comparison pairs each pair of pages has. The rule set before the run was this. To call the unevenness real, some sheet had to sit more than two spreads from the even effect. None came near it on that test. The comparison matched on distance, on its list, does not agree. The first paragraph of this section says so.

SheetObserved excessExpected from its sharing, and the deviation in its own spreadExpected from its word-list occurrences, and the deviationIts own spreadWords on itWord-list occurrencesComparison pairs per pair of pages
104 with 115+29.59+16.64 (+1.29)+15.29 (+1.42)10.051,7862134
105 with 114+5.19+13.80 (−1.08)+13.50 (−1.04)7.991,60218812
106 with 113+17.00+13.03 (+0.52)+12.85 (+0.54)7.651,89817916
107 with 112+0.93+9.23 (−1.05)+11.06 (−1.29)7.881,7631548

An earlier wording of this page said something else. It put the effect in two sheets and part of a third, not in the bundle. That claim is withdrawn. The published paper takes every bundle together and one pair per sheet. So it cannot see the sheets separately.

Their paper does print a table by bundle, with three scores for each. One score is for the two halves of a sheet. They call it conjoint. One is for the two sides of a leaf. They call it confoliate. One is for pages that face each other. Their reading of that table is one sentence, their paragraph 39. For this manuscript the facing score tends to be the lowest of the three. The other two are higher.

Their only other remark by bundle is about something else. It is whether the centre sheet of a bundle scores highest among its own sheet pairs. They say that happens only inconsistently. Their table covers nine bundles. It happens in three of them. No sentence of theirs says the relation between the three scores differs from one bundle to another. No bundle is named as an exception. The exceptions their notes name are in the two Latin comparison books.

What is added here is our own subtraction, done on their printed columns. Subtract their facing score from their sheet score, bundle by bundle. The difference runs from −0.095 to +0.282. Its middle value is 0.143. One bundle is reversed outright.

They printed no differences, no numbers of pairs, no spreads and no probability by bundle. A reader can do the subtraction in ninety seconds with their table in front of them. The reversal is measured against the facing score, and here is why. Against the leaf score that bundle's difference is only −0.067. That is not secure. Against the facing score it is −0.095. That survives the obvious doubt.

The doubt is that the reversed bundle is missing a leaf. Their own sentence on exclusions says which bundles they dropped. Those were foldouts, bundles missing single leaves, and bundles too short to be meaningful. So a reader will ask why this one is in. It is in their table and in their figure for the whole book. Taking their values by bundle together, with it in, reproduces their printed 0.433. Dropping it gives 0.452.

The one weakness of the reversal is for this page to state. That bundle is the only one whose sheet average contains no centre pair. Its centre sheet survives as a single leaf. They leave that leaf out. A missing centre pair would have to score 0.586 to lift its sheet average to its facing average. That is above every plant-bundle average in their table. Their highest plant-bundle average is 0.448.

So the reversal holds. But it could have been questioned, for the reason just given. We are pointing at a pattern in their own table that they did not draw attention to. We are not correcting them. We do not know what they made of it.

Folding the bundle in the wrong place

One check is to pair the pages about the wrong fold. Re-assign the sheets as if the bundle had been folded one leaf earlier. Then do it again as if it had been folded one leaf later. The pairs are still mirrored about a fold, but about the wrong one. If the effect survived that, it would not be about sheets at all. It is the only test in the search that compares the same pairs under a different folding. The others compare against a set of other pairs.

On the earlier list of words, the bundle's own frequent words, it does not survive. The effect is +3.29 shared words per pair under the true fold. On this test's own shuffle that is eleven in 200,000 shuffles. A second run with a different shuffle gave seven. The question in the section on the deepest limit is attached to it.

Under one wrong fold the figure is −1.57. Chance matches that in 94 of a hundred shuffles. Under the other wrong fold it is −0.87. Chance matches that in 88 of a hundred shuffles. That is nothing. Take the true pairs out of the comparison pairs. The first wrong fold then gives +0.90. Chance matches that in 11 of a hundred shuffles. The second gives −0.24. Chance matches that in 62 of a hundred shuffles. Still nothing.

The same test was built a second time, independently, on the eighteen-pair comparison. There the figure falls from +1.64 to +0.02 and +0.31 on its own scale. The eighteen-pair comparison's same-sheet figures are withdrawn, for the reason the section on the deepest limit gives. Its wrong-fold figures stand. Two constructions, the same collapse, on that list of words.

The deciding test then ranked the real folding eleventh of 120 on that list. So the list never supported the fold on its own by that test. The collapse published on it goes with the identity given in the section on the deepest limit. That identity is the same number with its sign reversed.

The test was then run on the words that carry most of the effect. Those are the book's own common words at that level of commonness. It had only ever been run on the earlier list. On those words the sixteen pairs give more. It is +5.29 more shared words per pair than their comparison pairs. On the shuffle that is one match in 200,000 shuffles, and none on a second run with a different shuffle. The same question is attached to it.

Folded one leaf early it gives −2.03. Its shuffle probability is 0.932. Chance alone beats it in more than nine of ten shuffles. Folded one leaf late it gives −0.58. Its shuffle probability is 0.677. Chance beats it about two times in three.

The collapse is the same shape as on the earlier list. It starts from a higher figure. The comparison against artificial texts was run on words of this kind too. None of twenty comes up to the manuscript. The manuscript sits seven and a half spreads above their average. The other bundle shows nothing on either list of words.

What this collapse is made of is set out in the section on the deepest limit. It is less than it looks. About half of it is the largest term of the true result, seen from the other side.

One result went the wrong way for the finding. It goes here and not in a note. Remove the sheet pairs from the comparison sets. The early fold then reads a shuffle probability of 0.049. That is about one in twenty. That is at the line set beforehand.

The explanation offered is this. Once those pairs are removed, one sheet pair has no comparison set left and drops out. The pair that drops is the early fold's most strongly negative. So what is left is three mild comparison groups. That explanation is convenient for this search. It has not been tested. Until it is, this figure rests on an explanation and not on the number.

What the words are, and what removing them does

An earlier wording of this page said the effect was carried by the bundle's own words. Those are the words common inside it and rare in the 182 pages outside. That was wrong. Here is how it was found to be wrong. Every way the test was built here defined its list as words rare under whichever cut was in use. So the book's own common words sat outside the test by construction. They were then measured.

Measured, the book's own common words at the same commonness fire the same comparison harder. On the shuffle of which pages count as one sheet, they give one match in 200,000 shuffles. They do so on fewer occurrences. The bundle's own list gives eleven in 200,000. Compared one for one against words of the same commonness, the listed words do not exceed them. The probability there is 0.064. That is about one time in sixteen, which is no evidence either way.

So the finding is this. The two halves of a folded sheet share more middling-common words than matched pairs elsewhere in the bundle do. That holds whatever the origin of those words. The bundle's own words are one part of the carrier. They are the smaller part.

The four-text check found the same reversal by a different test on a different text. The section on real texts gives that check. On the book's common words the manuscript stands above all 45 windows of Livy. It stands above all 39 of Livy's other run too. That holds under both of its comparisons. On the bundle's own list it fails. Two independent tests giving the same reversal are worth more than either alone. The section on real texts gives the figures.

The figures that supported the earlier wording stay, as a record of what they showed and did not. The measure adds up, for each pair of pages, the words both pages carry. So a word used on many pages enters many more pairs.

Removing the bundle's frequent layer takes away 21.9 percent of the word occurrences. It takes away 17.8 percent of the distinct words. But it takes away 61.3 percent of the pair mass. The result falls by 72 percent.

The fall tracks the loss of pair mass close to one for one. The sign is the same under all four ways of removing that layer. Per unit of evidence the rarer words show about seven tenths as much. What those figures show is that the bundle's own words carry much of the evidence in that list. They could not show what the list left out.

The ten words below are examples of the kind of word involved, and not the carriers of anything. They are the ten that carried most of the excess in the shuffle test's run on its list. They carry 48 percent of that run's gross positive total. They carry 78 percent of the net. They have 62 shared pairs in all. Of those, 54 rest on a word that appears exactly once on at least one of the two pages. And 22 rest on a single occurrence on both pages.

That is the definition used for the 54. A slightly different definition counts 53. So the ranking among the ten is not stable. A few single occurrences moving would reorder it. The whole list has 1,053 occurrences inside the bundle. The bundle has 23 pages. The list has 676 occurrences in the rest of the book.

WordContributionOccurrences in the bundle / in the rest of the bookPages in the bundle / in the whole bookOf the sixteen pairs: on either page / shared on both / shared where a page carries it once / where both do
cheeo+0.34418 / 114 / 1514 / 10 / 10 / 5
qokchedy+0.33329 / 1113 / 2214 / 10 / 6 / 0
kchedy+0.32313 / 910 / 1813 / 7 / 7 / 5
qopchedy+0.27621 / 1014 / 2115 / 9 / 8 / 3
qopaiin+0.2505 / 25 / 76 / 4 / 4 / 4
otair+0.24518 / 49 / 1213 / 5 / 4 / 1
qoeedy+0.21911 / 88 / 148 / 4 / 4 / 0
aiir+0.20814 / 88 / 1612 / 4 / 2 / 1
qotchdy+0.20811 / 108 / 179 / 5 / 5 / 3
chdaiin+0.19810 / 77 / 1210 / 4 / 4 / 0

One thing was checked about these words and came out empty. The list has 122 members. For 74 of them, the book has other words of the same commonness to compare against. Those comparison words number 588. On those 74, the listed words do not share across the fold any more than their comparison words do. The comparison shows nothing. Its probability is 0.344. That is about one time in three.

For the other 48, no equally common comparison word exists anywhere in the book. A word on seven or more of these pages is almost always one of the bundle's commonest. That is what put it on the list. Whether the two groups differ from each other beyond noise was then tested. They do not.

The division inside the list cannot be told apart from one even effect. So no claim rests on it. In the middle range of commonness, sharing across the fold looks like a property of words of that commonness. It does not look like a property of this list. That is the finding itself, and not a qualification of it.

Scored in another bundle, the same words do almost nothing. There are 24 occurrences of them there. In the last bundle there are 150. Six of ten leading words are present at all. The figure is +0.02 shared words per pair. Its probability is 0.44. That is an absence of power, not a demonstration of absence. The words are the last bundle's own. So the test elsewhere has little to work with. That is one reason the finding is one bundle's. Another bundle tested by the comparison matched on distance does not show the effect either.

The two pages in the middle

The two pages either side of the middle of the bundle face each other. On the way they were treated as a strong case of the sheet effect. Their height was the largest single number in the whole search. That support is lost. What follows says how. Ordinary prose poured across the two centre pages comes up to the pair's height in a minority of placements.

On one test it does so about one time in eight. That is five of forty placements of Caesar and the King James Bible. The lowest that test could report is one in forty-one. On the other test it does so about one time in six. That is fourteen of eighty placements of Livy, the King James Bible, Jókai and Don Quixote. The lowest that one could report within a text is one in twenty-one.

So the pair's height is within what prose on neighbouring pages can do. A first run of this measurement used one placement per text. It said that no prose came up to the pair. The sixty-placement run reversed it. One draw is not a control. Every such count can move by about one placement in twenty from where the text is started. The section on real texts says so.

Their size is almost entirely the word list's. They were scored under the five lists of words. Under a list drawn from the front of the bundle they score +14.83 on that scale. Under one drawn from the back they score −2.57. Under the whole bundle they score +7.65. Under the words common to front and back they score +2.59. Under the rest of the book they score +12.86.

Both pages are back pages. So under a front-drawn list they share the back's common words far beyond an ordinary back-to-back pair. Under a list that treats front and back alike, the pair sits at the level of the other same-sheet pairs. It is not a special pair. It is an ordinary same-sheet pair that looks special under one choice of words.

What the two pages share is not the words that carry the sheet effect. On the bundle's own frequent words they show no excess under one way of scoring. They sit +0.22 shared words above the average. That average is 11.78. Under the other way of scoring they show a weak positive, +3.04 shared words. Their sharing is on rare words. The rare words move as much under a wrong fold as under the right one. That is the same statement as there being no fact on the rare words. Nothing on this page rests on the rare words.

How the comparison was made

Two doubts stand against any result of this kind. Pages nearer each other in a book naturally share more words. And a page's front and back may differ in what they carry. The comparison matched on distance removes both before any figure is worked out. Each pair of pages from one sheet is compared only with pairs the same distance apart in the book. And it is compared only with pairs that have the same combination of fronts and backs.

One limit on the result on the bundle's own frequent words is this. On that list the effect is carried by the bundle's own frequent words. Removing those words takes the figure down sharply under two ways of making the cut. But that fall could come from having less data on rare words. So it does not show that the effect is about frequent words and not rare ones. Measured per unit of evidence, the rarer words show about seven tenths as much. The section on the words gives that figure.

That question has since been overtaken. The book's own common words, measured afterwards, carry the larger part of the effect. The section on the words says so.

The two tests were not looking at the same pairs. Their two largest pairs both sit on the sheet 104 and 115. On the shuffle test's frequent words, the largest single pair is the first leaf's front with the second's back. On its rare words the largest is the first leaf's back with the second's front. That is the pair one other figure rested on. Two lists, two largest pairs. No claim is built on either.

The largest of the sixteen pairs is settled as ordinary. Take the set's average out. It then sits higher than 87 in a hundred of its comparison pairs under one way of matching. Under the other it sits higher than 81 in a hundred. That is where the largest of sixteen usually sits. It is not a special pair. No claim is made that it stands above every artificial text.

Three kinds of pair, and an explanation that died

A sheet has four pages. They make three kinds of pair. The two sides of one leaf are one kind. The front of one leaf with the front of the other is a second. The back of one with the back of the other is a third. Measured on the sharing of words, the three kinds came out in a fixed order. We had an explanation we liked. Pages that sat closer together in the folded stack should share more words simply because of that closeness. The order would just be that closeness showing through.

Two separate checks killed that explanation. One built the table of distances from the way the bundle is folded. The three kinds sit in the same order under both readings of the distances. The pattern is clean sheet by sheet. The other came to the same sums by its own route. Both arrived at the same thing. A closeness effect predicts the opposite of the order that was measured.

So the explanation is dead. What is left is an order in the measurements that nothing so simple explains. Any difference between the two physical surfaces of a sheet was also tested for. None was found. A reader should know that we tried to explain our own finding away and could not. What stands is a measured pattern with no cause yet found for it. The sixty placements add a third line against the explanation. Prose's own falling-off with distance does not reproduce the order either.

Real texts written into the same pages

The other main check takes real texts and writes them into the manuscript's own page shapes. Then it measures the same thing on them. A text used that way is a control. It is a text of known origin whose answer is known. If real writing does what the manuscript does, the finding is just something written text does. No published study has written a real text into this manuscript's page layout. The nearest is a 2026 paper by Rozanova and Temerev. It wraps control texts to the manuscript's line lengths to study its letters.

The whole count of what was poured is one sentence:

We poured fifteen real texts into the manuscript's own page and line layout, in 645 separate placements across seven independent checks; one text at twenty offsets counts as one text and twenty placements; the headline test rests on four of those texts at twenty placements each, with a floor of one in twenty-one within a text.

Two words in that sentence need saying. An offset is a starting point. A floor there is the lowest value a count of twenty placements can report. That floor is one in twenty-one. The count says texts and not books. One text can be many books. Livy alone is twenty books of one work. The King James Bible is sixty-six books. Jókai is there as two novels. Every pour is an extract. No text was poured whole. The longest stretch used was about two thirds of one text. The full inventory of what was poured is in POURS.md beside this page.

Five books were written out into the bundle's pages. Each went in at twenty placements. The bundle has 23 pages. That made a hundred placements in all. Twenty of them were of a text we later struck. So eighty stand. They are Livy, the King James Bible, Jókai and Don Quixote. This page calls that check the four-text control. On that grid the manuscript stood above every placement of three of four texts. It stood above nineteen of twenty placements of the fourth. Which one is the fourth depends on where the windows fall.

The placements were then re-scored with front matter and title lines cut out. This page calls that re-scoring the cut. Under the cut the manuscript stands above twenty of twenty placements of Jókai. It stands above nineteen of twenty of Don Quixote. That count has since weakened. It is not what this page rests on.

Livy was then run at 45 windows. The earlier grid had twenty. The windows ran one after another through 35 of its books. A window is one stretch of the text long enough to fill the bundle's pages. Those counts are on the bundle's own list of words. So are the twenty-placement counts above.

Livy was written out into the pages four different ways. How often it comes up to the manuscript's excess runs from none in twenty to five in 45 placements. The none in twenty is the grid where each window starts at the head of one of Livy's books. The five in 45 is the continuous grid in reading order. The other two ways are in the as-made order. One gives two of nineteen. The other gives two of 39. Both fall between those two rates.

That is the whole range there is. The width of that range is itself the finding. The rate cannot be pinned down with the text poured.

The 45 windows were then re-scored with the title lines cut. That count becomes four. At 45 placements the excess is refuted under the reading order. That is on the reading that takes all comparison pairs together. Under the as-made order it is undecided. Two of thirty-nine placements fall in the middle of the rule's range. There the rule decides nothing.

On each text's own middling-common words the answer is the opposite. A hundred placements were written out into the pages. They were of five books. Twenty of them were of a text we later struck. So eighty stand. The manuscript is above all eighty under both readings. The nearest placement sits at +3.230 on that scale. The manuscript sits at +5.328.

The manuscript still stands above all of them after the text is shifted by a page. But its lead over the nearest of them roughly halves when it is. Under one reading the lead goes from 2.1 to 1.2. Under the other it goes from 3.4 to 1.7. The rank survives the shift. The size of the lead does not.

A second measurement makes every count softer. Shifting a window by some tens or hundreds of words moves a placement's excess. The move is up to 1.7 on its scale. Before the run, the predicted move was under 0.5. So every count of so many placements in so many can move by about one in twenty. The movement comes from where the page boundaries happen to fall. This page calls that movement the jitter.

The jitter differs by list of words. Shift by one page on the texts' own middling-common lists. A single placement moves by up to 2.8 or 3.2 on that scale. On the bundle's own list it moves one by up to 1.7. Two verdicts on the four-text control move with the shift, each by a single placement. One goes from missed to held. Another goes from refuted to missed.

The rule the counts are judged by has a pass range and a fail range. Whether anything lies between them depends on how many placements were laid. With only twenty placements, a single placement moving decides the verdict. One placement coming up to the manuscript is a pass. Two is a fail. That is why the same test came out two different ways under a shift of less than a page. At nineteen only a clean zero passes.

The 45-window result is a real failure and not a product of the counting. At 45 placements there is room between passing and failing. That middle is three or four placements. At 39 placements there is a middle too. It is two or three placements wide.

That jitter of one in twenty holds for the last bundle, under shifts of less than a page. Its pages carry about 470 words each. The other bundle's pages carry about 340 words each. There a shift of a whole page moved one count from 10 to 6 of 20. It moved another from 9 to 7. So every count on the other bundle's pages carries a jitter of up to four in twenty.

Twice a real text has come up to the manuscript by a single placement. One gap was +0.108 on the same scale. The other was +0.004. Both gaps are inside the jitter.

What the four-text check leaves standing is not a count. It is a difference in how a real text gets there. It is a reading of the figures and not yet a fact. A real text that comes up to the manuscript does it in one of two ways. Either its excess is spread out across its pairs. Or its excess sits close to its comparison level. The comparison level is the figure the comparison pairs sit at.

The manuscript's own comparison level is −0.321 on that scale. That sits above the averages of all four texts. Their averages run from −0.42 to −0.56. And it is ordinary inside each text's own range. Each text has twenty placements. Six to nine of them sit at or above it. So the manuscript's result is made by what the two halves of its sheets share. Its comparison level is unremarkable.

Both real-text placements that come up to the manuscript are scenes of a novel. Their words recur at the sheet distances. One is the Jókai placement. Take its three highest pairs. Between them they share 45 distinct words. The words are those of a runaway-mill scene, a dialogue and a health inspection. Not one is a name. The other is the Don Quixote placement that comes up to the manuscript under the cut.

Both placements fall below the manuscript on the matched reading. They also fall below it on the reading with the same-sheet pairs taken out. The manuscript itself does not fall on those readings.

On the other bundle the manuscript's comparison level itself sits above every placement of all four texts. So on those pages the whole block of far-apart pairs is raised. It is not the sheet pairs alone. The sheet excess sits inside the range of all four texts. On the book's common words the picture is different. There the manuscript stands above all 45 windows of Livy. It stands above all 39 as well, under both comparisons. The section on the words puts that beside the same reversal on the shuffle test.

The sixty-placement pour is the one the evidence on this page leads with. It poured the King James Bible, Don Quixote and Montaigne into the same page layout. Each went in at twenty placements. That is sixty in all. The same-sheet excess was measured on each. Five lists of words were used. The manuscript stands above all sixty under every one of them. That holds on both comparison levels and under both ways of scoring.

The section on the result found here gives the figures. It gives the four further measures with their counts. And it gives the limits that ride with them. A rank against real prose needs no drawn set of comparison pairs. So the question in the section on the deepest limit cannot make it look stronger than it is. The jitter above applies to it as to every count.

The same pour gives three negatives. Each strengthens this page by taking something away. The first concerns the region effect. That was an early result saying the effect covered the front half of the bundle against the back. It was retired. The record below says how. All three poured texts reproduce the region effect. So it was a by-product of the layout. Prose makes it too. It stays retired.

The two pages in the middle behave the same way on prose. That confirms their withdrawal. And prose's own falling-off with distance does not reproduce the order of the three kinds of pair. That is a third independent line against the explanation we liked and abandoned. The section on the three kinds of pair tells that story.

On every twenty-placement grid the real texts fell short of the manuscript's level. The paragraphs above say what happened at 45. One limit rides with the sixty placements. In that pour the twenty placements of one text overlap heavily. They step about two thousand words at a time. Each window is about eleven and a half thousand words. So they are not twenty independent draws.

The cut above is a re-scoring of the four-text control. It cuts front matter from two placements and title lines from a third text. That front matter and those title lines were in every figure before the cut.

Every real book came out a little below its comparison pairs, not level with them. There is a reason. In these tests the same-sheet pairs sit on shorter pages than their comparison pairs. So any text at all written into those slots reads a little low there. The comparison is tilted against the same-sheet pairs. When the tilt is subtracted the manuscript still stands well clear. But a subtraction is not a controlled test. The comparison has not yet been re-run under control. That is left open.

The worry that page length drives the result was the largest open question here. It has been tested directly. The eighteen-pair comparison removed the effect of page length from its test of the sixteen same-sheet pairs. By leaf number the pairs still sit higher than 98.0 per hundred of the comparison's values. In writing order they sit higher than 97.2 per hundred.

The 31 poured texts were run through the same control. Their average comes out at +0.4 on its scale. By leaf number the manuscript sits 1.67 spreads above that average. In writing order it sits 1.61 spreads above it. The comparison's own scale is extra shared words in a pair, beyond the comparison pairs. On that scale the manuscript is 1.901 above the average by leaf number. In writing order it is 1.946 above. Two of the 31 poured texts come up to or pass the manuscript under both matchings. So the worry was tested and not merely acknowledged.

The block share moves the eighteen-pair result under one matching. The block share is the part of the comparison's error that the blocks of leaves account for. The section on the deepest limit explains those blocks. With it applied, the result goes from one in three hundred and twenty-two to one in a hundred and sixty-one. Under the other matching it moves from one in ninety-three to one in fifty-one.

Only the block part is applied. The unexplained part is not. Neither its cause nor the manuscript's own kind is known. The two matchings, and why there is no third, are stated in one sentence:

the two weightings differ only in the weight of one pairing of two pairs, whose mean has twice the variance of the others; one holds the poured texts and leaves the artificial ones off-centre, the other centres the artificial ones and loses the poured ones; no third weighting is proposed, because a weighting chosen after seeing those two would be chosen to pass, not built to test.

That comparison's own limits ride with it. Take the bundle's own lists of words, without the control. There the same comparison sits higher than 85 and 94 per hundred of the values. It does not sit higher than 99. Two things account for the fall. One is the mass of words counted. The other is how the comparison pairs are composed. Apply the word list and the control together. Then the figure sits somewhere between 71 and 85 per hundred. Under both together the result is not claimed.

On its own comparisons the shuffle test sits higher than between 85 and 95 per hundred. It never crosses the line. That is a consistent excess that never crosses the line on that particular test. It does not move with the list of words either. One list puts it higher than 92.0 per hundred. Another puts it higher than 92.2. The third puts it higher than 89.1. After the control, the gap over the best of our own imitations is one to two spreads and not three.

The same question across the whole book

Our own test of the same question covers the whole manuscript. It is rougher. It takes the two leaves of one sheet and finds how many of their distinct words they share. It uses all their words and not a chosen list. It compares that with pairs of leaves the same distance apart in the same bundle by the same scribe.

Over 24 sheets the two halves share more than the comparison pairs do. The extra is about one part in a hundred. That is 0.011 on the test's scale. On that scale 0.15 is the usual sharing between two leaves. Shuffling which leaves count as one sheet matches that figure about one time in forty.

Latin prose written into the same pages shares 0.009 more. That is as much. So on this whole-book test the manuscript is not unusual. The figure also does not survive the correction applied for having run many tests. Nine of 24 sheets go the other way. The sheets written by the first scribe give 0.006. The second scribe's give 0.016. The third scribe's give 0.018.

The largest differences are on other sheets than the last bundle's. One is the sheet 50 and 55. It gives 0.066. Another is the sheet 51 and 54. It gives 0.065. Each of those is against two comparison pairs. Next is the sheet 78 and 81. It gives 0.037. That is against six comparison pairs. The middle sheet of the last bundle is the sheet 108 and 111. It gives 0.036. That is against eight comparison pairs, the largest number of any sheet.

The comparison pairs in this test share less than pairs in general do. That holds in the manuscript. It holds just the same in every real or artificial text written into its pages. It is what makes the choice of comparison matter. It is not a quirk of the manuscript.

Measure instead against all pairs in the same bundle by the same scribe. Then the manuscript's excess falls to 0.001. The Latin's falls to 0.002. Measured against every comparison pair with equal weight, it is nil. So at the scale of the whole book, the choice of comparison decides the size of the effect. By one of the three natural choices it is nil. It is weak by one way of measuring and absent by another. That is why the claim is one bundle's.

A second whole-book test asks a different question. It asks how predictable the text of each page is. Then it asks whether the four pages of one sheet are alike in that. They are. The resemblance is large. But two real texts written into the same pages show a resemblance of the same size. Score the same pages by bundle instead of by sheet. That shows the resemblance does not come from position in the book.

A sheet's four pages are two adjacent pairs, the front and back of each leaf. A continuous text scores on this test through the leaf. So this test separates the sheet neither from the place in the book nor from the leaf. The comparison matched on distance, in the last bundle, can. There leaf pairs are a separate class. That is why the one figure comes from the bundle test and not from this one.

What is added to the published result

Layfield and Davis found it first. A reader should go to their paper. Their result rests on a comparison with no allowance for distance and no shuffle test. The result here puts the same conclusion on ground that can be attacked. It survives the attack. What is added is this list.

One thing on this page is not claimed as our own idea. Stripping out the common words was Nick Pelling's suggestion in 2022. He proposed a return to an earlier measure with the common words left out. Nobody carried it out. What is added here is having done it and measured what happens. Leaving the common words out is also why the larger part of the effect went unseen here. It stayed unseen until it was measured. The section on the words says so.

Where this page disagrees with the paper

Their central result is not disputed. Pages joined as one sheet are more alike than pages that face each other in the binding. Their Table 8 gives it bundle by bundle. It covers nine of the book's bundles. The book has twenty or so. The result shows in eight of nine. The measurements here agree with it. The disagreement is about one further step. It is about that step and not about their work.

They call a bundle made of a single sheet a singulion. Their claim is that the book was written as a stack of single sheets. The four pages of each sheet were written in sequence. The sheets were gathered into bundles later.

The step that gets there is a test. It asks whether the two halves of a sheet differ from the two sides of a single leaf. The test gives a probability of 0.168. Their stated rule accepts no difference when the probability is above 0.05, one in twenty. The two sides of a leaf are by definition read in sequence. So, they conclude, the two halves of a sheet must be too. That step is what turns ‘these pages are alike’ into ‘these four pages are read as a sequence’.

The reply, in the plainest terms this page can manage, is this. A test that fails to find a difference has not shown there is none. On their own numbers the sheet halves sit above the leaf sides. Over the book the sheet halves sit at 0.433. The leaf sides sit at 0.383. In the last bundle the sheet halves sit at 0.600. The leaf sides sit at 0.568. In another bundle the sheet halves sit at 0.603. The leaf sides sit at 0.530.

A difference of that size is beyond what their test could detect. It pools 38 measurements on one side. It pools 76 on the other. They come from nine bundles that differ a great deal from each other. On the test used here, a leaf's two sides behave differently from a sheet's two halves.

That is not a new doubt. In 2010 Nick Pelling commented on Julian Bunn's post. He rated how alike the two sides of each leaf in this bundle are on Bunn's measure. He reprinted the rating in December 2022:

103 good, 104 very bad, 105 very good, 106 bad, 107 excellent, 108 excellent, 111 excellent, 112 good, 113 excellent, 114 excellent, 115 very bad

Against that, he found the two halves of the sheet 104 and 115 alike across the sheet. Each leaf's own two sides are ‘very bad’. That is a precedent from 2010. Already then the two sides of a leaf and the two halves of a sheet were seen behaving differently. That is exactly the point where this page disagrees with the paper. So the disagreement puts a measurement on something noticed sixteen years ago.

The limit of that reply comes next. Their measure is a similarity score over whole pages. It drops the rarest words and weights the frequent ones down. The measure here is the sharing of middling-common words against comparison pairs the same distance apart. Two things can be indistinguishable on their measure and distinct on this one. Neither measure need be wrong for that. So the disagreement is about the inference and not about the number. A shared occasion of writing accounts for words shared across a sheet. It makes no claim about the order the pages were read in.

Their proposed order, and a trap

This section is not a result. It is a demonstration of a trap. It is kept because we walked into it.

Layfield and Davis propose an order for the six sheets of the last bundle. They chose it by their own measure of how alike the pages at each join are. We rebuilt their method from their published description. We used our transcription and not theirs. Then we scored every arrangement of the six sheets. There are 720.

Their proposed order comes out first among the fully scored arrangements under the rebuild. The gap between it and the rest runs the same way as the one they report. It is also of about the same size. Redone here, their computation gives the same answer.

We then had a number that looked like independent confirmation of their order. Our own measure of sharing put their order in its top five percent of all the arrangements. There are 720 of them. Their order scored 0.035 on our scale. We measured how much of that was guaranteed by how they had chosen the order. They had chosen it on a similar measure. The answer was nearly all of it.

Across the 720 arrangements, their measure and our own rise and fall together. They agree at a correlation of 0.87. That is high enough that their order was always going to score well on our measure. They chose it by making as large as possible something close to what we measure.

That agreement alone predicts 0.036 for their order on our measure. The observed figure is 0.035. The part our measure adds beyond that agreement sits in the middle of all 720. We also ran a simulation at that level of agreement. It puts the arrangement best on their measure inside our top five percent. The probability of that is 0.976.

So our measure adds no preference of its own for their order. The two measures agree because they measure the same thing on these pages. That is the common words shared at the joins. Our measure does not confirm theirs independently.

One claim made on the way is withdrawn. It was that the book as bound sits below the middle of the 720 on our measure. Scored their way, as a sequence and not a nesting, the book as bound sits at 243 of 720. That is the middle of the pack. Its own joins score at the level the comparison expects. There is no finding here about the book as bound looking bad on a measure built from its own words.

One thing survives as our own. On this book our measure can tell arrangements apart at all. Across the 720 arrangements its scores have a spread. That spread is 2.70 on its own scale. That is above every one of five artificial texts. They sit between 1.05 and 1.31. A text with no structure has no preferred arrangement. This book does. That is worth one clause and no more.

Two notes go with the rebuild. Our transcription is a later one than theirs. So any difference between the rebuild and their figures could be on our side and not theirs. And the rebuild is our reading of their description, not their code. We get 935 chunks of text. Their count was 606. We get 2,345 distinct terms. Theirs was 2,161.

The good news did not survive its own check.

Earlier work, and why it was not read

This search did not read earlier work on the manuscript until the end. It had been asked to work from the manuscript alone. So the finding on this page was made before anyone here knew it was in print.

The idea is older than the paper. In 1976 Prescott Currier found that the manuscript's two styles of writing change by folio, meaning by leaf. Its hands, the different handwritings in it, change by folio too. He spoke of folios and pages, not of sheets. His own account of how the book was made is leaf by leaf:

some sixty-five folios were prepared ahead of time with drawings on them. They were placed on a table so. The first twenty-five folios were taken, one at a time, off the top and filled in with writing by one individual.

That is the opposite of a sheet at a time. So he is no precedent for the finding on this page. It was René Zandbergen who read Currier's finding at the level of the sheet. He did so on his page on the origin of the manuscript. Zandbergen claims that within each sheet the hand holds steady and so do the counts made on the text. He claims no more than that. On the same page he writes:

It seems impossible to say whether all four drawings on a bifolio were done before all four text sections, or it was done page by page.

So he claims nothing about the order the pages of a sheet were written in. In 2013 Nick Pelling made the same reading. Any one sheet, on that reading, carries one style of writing only.

In 1997 Zandbergen grouped some leaves of this bundle by the words they use. He took six of its leaves. They were 103, 107, 108, 111, 112 and 116. He grouped them as three sheets alike in their words.

In 2010 Julian Bunn published a table of how alike each pair of pages is in its words. Pelling, in a comment on that post, flagged two pairings. Both were on the sheet 104 and 115. One was the back of 104 with the front of the sheet's other leaf. The other was the back of 115 with the front of the sheet's other leaf. He also flagged the sheet 108 and 111. He called the sheets 105 and 114, 106 and 113, and 107 and 112 unconnected. His rating of each leaf's two sides, from the same comment, is quoted in the section on the disagreement.

In 2022 Pelling wrote of the sheet as a working unit in this bundle. He suggested going back to Bunn's measure with the common words left out. He did not compute it. Nobody else did. In 2019 Vladimir Dulov measured what share of one reference text's words each page carries. He averaged that by sheet. He concluded that each sheet was written before the bundle was sewn.

Works referred to on this page

Each entry ends with one of three words. Read means the work itself was read in full at the link given. Described means only an abstract or another author's account was available. Not read means the work could not be opened. The entry says why. The full list is in sheet_finding_bibliography.md beside this page. It says what each work claims. It says which lines are our own sums.

What is not claimed

It is not claimed that the finding is new. It is published. The first paragraph of this page says where. What is claimed as new is the list of checks above.

It is not claimed that some sheets carry the effect and others do not. Two of our own tests disagree on it. The section sheet by sheet says so.

It is not claimed that the words carrying the effect are the bundle's own. The book's own common words carry the larger part. The section on the words says how that was missed.

It is not claimed that the collapse under a wrong fold is independent evidence in full. About half of it is the true result seen from the other side. The section on the deepest limit says so.

It is not claimed that the effect is found on the bundle's own words. Those are the words common inside the bundle and rare elsewhere in the book. On that list the deciding test ranks the real folding eleventh of 120.

It is not claimed that any shuffle figure on this page is settled. One question stands over all of them. The section on the deepest limit states it.

It is not claimed that real texts written into the same pages never come up to the manuscript. One text was poured at 45 placements. It comes up to the manuscript about one time in nine. The section on real texts says what is left.

It is not claimed that the sixty placements are that many separate tries. They are placements of three texts. The section on the result found here gives the range that count can report and what limits it.

It is not claimed that our measure confirms the order Layfield and Davis propose for the sheets. The section on their order says why.

It is not claimed that the two pages in the middle of the bundle are part of the sheet effect. They are an ordinary same-sheet pair that looks special under one choice of words. And what they share is on separate words.

It is not claimed that several separate results confirm each other. The figures once treated as agreeing are several labels on mostly the same hundred comparison pairs in one book. What this page rests on is one book and one set of page pairs, measured several ways. It is checked against a wrong fold and against real texts written into the same page layout.

The record of the search, dated 17 September 2026

This section keeps the results that were found and then retired during the search. Beside each is what retired it. It also keeps the mistakes found in our own work.

A region effect, retired four times. An early result said the sheet effect covered a region of the book. That region was the front half of the bundle against the back. Four separate tests retired it. Real prose written into the same pages reproduces it. A second run used five real texts. On all five the manuscript sat inside the range of prose. A third test found that the region effect it saw came from its word list. That list had been drawn from one half of the bundle. The effect had stood above every one of twenty artificial texts. Under a list drawn from the other half the difference reverses. The manuscript's value becomes an ordinary one. The fourth killing is the wrong-fold test above. The region reading cannot pass it.

The pour of sixty placements then reproduced the region effect too. All three of its texts made it. So prose makes it.

A leaf effect, retired. An early result said the front and back of one leaf share unusually much. Real prose written into the pages matches the manuscript's figure. The result was restated. The unusual thing is that the first page of the bundle shares little. It is not that the leaf shares much.

One sheet as the carrier, retired. An early result said one sheet in the last bundle carried the whole effect. It also said a second test had found its two strongest pairs on the same sheet. The second test had not. The two tests share one pair of pages. The second test's other strongest pairs are on the neighbouring sheet. The sheet-by-sheet breakdown above replaced the claim.

Two sheets as the carriers, withdrawn. A later wording of this page said the effect belonged to two sheets and part of a third. It said the effect was not a property of the bundle. The uneven look across the four sheets was then put to the shuffle test. Could it be one even effect seen through uneven noise? It can. No sheet sits more than 1.4 spreads of its own noise away from one even effect. The rule set beforehand needed a sheet more than two spreads from the even effect. The claim came out. The comparison matched on distance, on another list of words, then found the sheets far from even. The section sheet by sheet carries the disagreement.

The order of the three kinds of pair. The section on the three kinds of pair tells it in full. An explanation was offered. Two separate checks killed it. A later reading said the order rested on two page pairs on the outermost sheets. That reading has not been compared with what the noise of those pairs alone would produce. So it is kept off the page.

The two pages in the middle, shrunk three times. An early result said this. No ordinary prose written into the pages came up to the height of the two middle pages. That rested on one low placement per text. It was withdrawn. Later runs, with the rule fixed beforehand, put the pair within what adjacent prose can do. On one test prose comes up to it about one placement in eight. On the other it is about one in six. A figure in between, one in twelve, was computed with a text since struck. So it is not used.

Another test then found the pair's sharing is on separate words from the sheet effect. Under five lists of words, its size is almost entirely the word list's. Under a list treating front and back alike it is an ordinary same-sheet pair. On the sixty placements the pair behaves the same way on prose. It does so on all three poured texts. That confirms the withdrawal.

A figure withdrawn. On the eighteen-pair comparison the same-sheet pairs sat higher than 98.9 in a hundred of the drawn values. There are eighteen same-sheet pairs in it. The figure was withdrawn. The number had not moved. The control built to defend it had failed its own test on real texts. That control was meant to remove differences in page length. The corrected figure turned out to be the original figure under another name. The poured texts did not come back to zero under it. The control brought none of the four real texts poured into the pages back to zero. So the figure stays withdrawn. Its same-sheet figures go with it.

The same comparison's re-pairing test was run on ten leaves, with the two centre leaves left out. It put the arrangement it was testing fifth of 120. The reason why is the deepest limit on the whole search. The section on the deepest limit carries it. A better test was then built. The section on real texts gives what it found.

A defect in the comparison matched on distance. In that comparison, one comparison group was empty at one distance. That left half its combinations undefined. The test was run again. Two smaller blemishes remain. An earlier sentence had said the pairs furthest apart were the strongest. It was corrected. It had read a difference measured against comparison pairs as if it were a raw value.

A largest difference named wrongly. The middle sheet of the last bundle had been named as giving the largest difference in the whole-book test. It is fourth. The two largest are on sheets in another bundle. The page above names them, with their figures and the number of comparison pairs beside each.

A claim of blindness, and its limit. The shuffle test's runs and their verdict sentences were fixed before the result of another test arrived. The code and the logs were in place before that result arrived. The verdict wording is fixed in the code as literal text. The figures choose between the options. The written reading that followed carries a time two minutes after that result arrived. So it cannot be checked. So this page says the runs and their conclusions were fixed beforehand. It does not say the same of the discussion written around them.

A poured text, struck. One pour used five real texts. One of them was struck after everything was computed. The Montaigne file held the sixteenth-century original and a modern French rendering of the same text, one after the other. So several placements lay in the translation and carried the same content as earlier placements of the original. The text was struck. Every clause was then checked on the four texts that remain. Three things improved when the bad text went. A clause that had failed only because of it holds without it. The prediction made beforehand about the kinds of page holds in all four. And on the other bundle the manuscript sits above all twenty placements of every remaining text. That holds under both ways of scoring. That count carries the larger jitter of the other bundle's pages. There the jitter is up to four in twenty. The section on real texts describes it.

Placements that are not separate from each other make a count look bigger than it is. That happened twice in the controls here. The Montaigne file was the first time. The second was in another pour, where the twenty placements of one text overlap heavily. They step about two thousand words at a time. Each window is about eleven and a half thousand words long. So they are not twenty separate draws.

A control that could not fail. One of our tests asks whether the number of words on a line depends on the words on it. A writer fitting a line to the edge of the page would produce that. We had compared the manuscript with Latin written into the same lines. The Latin showed no dependence. It could not have. That pour gave each line the manuscript's own number of words, whatever the words were. We replaced it with a control that could fail. That control is real texts written to fill each line to the manuscript's own width. Real texts written that way show a dependence two to three times the manuscript's. So the manuscript's fit to the line is a partial fit.

A stored figure that does not reproduce. We tried to reproduce that Latin comparison. We found where the Latin figures stored on our pages came from. They came from a copy of the reference file that is gone from the disk. The file was rewritten on 14 September 2026. The rewrite was at 20:17. The stored result had been saved at 16:22. That is four hours earlier. The file as it stands gives 0.8 percent. The stored result gives 0.6. Both are small. The record carries both.

A predictability figure, cut. These pages had put a figure on how much of a page's predictability the sheet explains. Every text written into the same pages gives a figure swollen in the same way. So the number added nothing. What remains is the comparison inside groups of pages by the same scribe, against pages shuffled among sheets. A shuffle figure in that test moved between two runs. That happened because the shuffles started from a different random point each time.

The carrier, misdescribed. Every earlier wording of this page said the bundle's own words carried the effect. Those are the words common inside the bundle and rare outside it. Every test built here had defined its list as words rare under the cut in use. So the book's own common words were never measured. When they were measured, they fired the same comparison harder. They gave one match in 200,000 shuffles. The bundle's own words had given eleven. The wrong-fold test was then run on them. The collapse reproduced, from a higher figure. The four-text control found the same reversal on its own test. The section on the words carries all of it.

The pour control, weakened. One count had been called the strongest control on this page. It was this. The manuscript stood above every placement of three texts. Each text had twenty placements. And it stood above nineteen of twenty of a fourth text. That count was a property of how many windows were taken. Livy was then taken at 45 windows. There ordinary Latin prose comes up to the manuscript one time in nine. A second measurement showed that every such count carries a jitter of about one placement in twenty. What is left is a difference in how a real text gets there. The section on real texts says what that is.

Two descriptions of single sheets, gone. One sheet had been described as also carrying the effect. Another had been described as flat. On the comparison matched on distance the first is not supported at all. The second's flatness is the comparison's own unevenness. The shuffle test and the comparison matched on distance disagree about the sheets. No third test has been run to break the tie. Any third test would be chosen knowing what answer was wanted.

A largest pair, settled as ordinary. The largest of the sixteen same-sheet pairs had been treated as standing above every artificial text. Once the set's average is taken out, it sits where the largest of sixteen usually sits. The claim is off the page for good.

A number that looked like confirmation. There are 720 arrangements of the sheets. Our measure put the order Layfield and Davis propose in its top five percent of them. How much of that was guaranteed by the way they chose the order was then measured. The answer was nearly all of it. The section on their order tells it as a demonstration of a trap.

A defence of a strong number, withdrawn. The comparison matched on distance had been defended as immune to the clustering doubt. The defence was that it holds page distance and side fixed. Counting the blocks of leaves shows why the defence fails. The shuffle draws groups of 21 pairs. A drawn group covers about sixteen distinct blocks of leaves. The real pairs sit in six blocks. So the doubt applies to the way that comparison is built.

Its figure was then checked against eighteen stretches of real prose. One of them was at the manuscript's level. The figure that had been called five in ten thousand is on this page as a range. It runs from about one in eighteen to about one in seven hundred and fifty.

Numbers larger than the evidence supports. Earlier wordings of this page led with a shuffle figure of about one in twenty thousand. For a time it led with one under five in a million. Both come from drawn comparison sets. Those sets were then checked against real prose. Both figures are on this page as supporting figures, with ranges. Both are printed as counts on this page. One is eleven in 200,000 shuffles. The other is one match in 200,000. Each has its second run beside it. What this page stands behind is two ranks. One is a rank of one in 120. The other is a rank above sixty poured placements.

The bundle's own list, and what the deciding test says of it. The whole search was built on a list of words common in this bundle and rare outside it. On that list the real folding ranks eleventh of 120 ways of pairing the leaves. On the narrower run it ranks fourth of 24. That list, then, never supported the fold on its own by the test that decides it. The collapse published on that list goes with what the next item takes apart.

The wrong-fold collapse, taken apart. The collapse under a wrong fold had been called the central piece of evidence on this page. At one page spacing the true sheet and the early fold's sheet are each other's only comparison. So half the collapse is the true result with its sign reversed. What is left is three sheets on each side, with the true fold ahead. The section on the deepest limit carries the figures.

A claim about the book as bound, withdrawn. On our measure, the book as bound had been said to sit below the middle of the 720 arrangements. Scored their way, as a sequence, it sits at 243 of 720. Its own joins score at the level the comparison expects.

A re-run of the deciding test, completed. The deciding test was run a second time, from the manuscript's words, on separate copies of the programs. It went through the counts and the thousand shuffles. Then it went through the chance figure and the thirty sheet-by-sheet figures to the final ranking. Every single number came out identical. The re-run did not start from our numbers. It ran everything again from the text.

What was still open

Several questions were still open when this page was finished. None of them is answered on it. Is the wrong-fold test exempt from the question in the section on the deepest limit? That decides whether the one match in 200,000 keeps its raw value. How many separate tries do the sixty placements amount to? They are placements of three texts. Does the explanation for the one clause the wrong-fold test missed hold?

Is the difference in how a real text gets there a measured fact, or only a reading of the figures? What does the tilted comparison give when re-run under control? What does the rest of the four-text control's re-scoring give, with front matter and title lines cut out? And two runs of the test on the bundle's own words had given different answers. One gave about one in twelve. The other gave eleven in 200,000 shuffles. The two had not been set side by side.