Back to The Journal

A Forgery Is Dated by Its Newest Word

Alex Wilson6 min read
Brass magnifying glass over one line of an aged handwritten parchment on a lamp-lit desk

There is a document at the center of a long-running treasure story, said to be a 12th-century French account carried out of the Vatican in the 1970s. It has photographs. It has custodians who have defended it for decades. It has a cipher, an alphabet with a genuine early-modern pedigree, and a chain of provenance detailed enough to name the Parisian book dealer who sold it.

It contains the name CEDRIC.

Walter Scott coined Cedric for Ivanhoe in 1819. He appears to have gotten it by misreading the Anglo-Saxon Cerdic. The name does not exist in French records before 1820. Everything else in the document could have been perfect and it would not have mattered, because a document is dated by the newest thing in it, not the oldest.

I spent a chunk of last week having a machine chase this. It pulled the 1533 Cologne printing of Agrippa off an archive scan and read the alphabet plate off leaf 295. It pulled the 1561 French Trithemius and compared the character inventories glyph by glyph. It built a searchable corpus of all 52 Nag Hammadi tractates in ten minutes so that a claim of absence could be a measurement instead of a memory. It ran vocabulary curves out of the Google Books Ngram API with the smoothing turned off, because the default smoothing blurs exactly the inflection point you are looking for.

All of that is legwork. Real legwork, the kind that used to take a specialist a season. None of it is the finding.

The finding is one word, and the reason the word is fatal.

Three tiers of tell

The people who took this document apart worked through the French orthography and found a stack of anachronisms. ETRANGE where a 12th-century hand would write ESTRANGE. ILE for ISLE. TEMPÊTE with a circumflex that cannot exist before the 17th century. INUIT, first attested in French in 1963.

Sort that kind of evidence and it falls into three tiers, in increasing order of how hard it is to fake.

The first is vocabulary. A word whose usage curve starts after the claimed date. This is the easiest tier to defend against, because a careful forger can look words up.

The second is orthography. Spellings that postdate the claim. Harder, because spelling is habit rather than knowledge, and habit leaks. This is the tier CEDRIC sits in.

The third is the received misconception. A belief whose earliest attestation postdates the document. Somebody writes a supposedly ancient text that assumes a thing everyone knows, and the thing everyone knows turns out to have been invented in 1887, or 2003, or last year.

The third tier is the strongest, and the reason is worth sitting with. A forger can audit their vocabulary. They can, with effort, audit their spelling. They cannot audit their own background assumptions, because background assumptions are not experienced as claims. They are experienced as the floor.

Whatever you take for granted is the part of your work you cannot proofread.

The evidence that agrees with you

The same week produced the opposite lesson, which is the one I keep needing.

A published solution to a related cipher has been circulating since 2016. The decoded text comes out as French. The index of coincidence on that plaintext is 0.0782. French plaintext runs about 0.0778. That is a startlingly good match, and it looks like confirmation.

It is not. The analyst derived the key by assuming a disputed reading of a different artifact was correct, then applied that key until French came out. The statistic measures the fitting, not the document. You cannot use the output of a procedure as evidence for the procedure.

There is a second detail that only shows up if you count. Roughly a fifth to a quarter of the published plaintext was supplied by the analyst, who said so plainly and marked his additions in blue. The additions sit precisely at the line breaks, which is to say precisely where the reading would otherwise have failed.

That is the shape of nearly every bad conclusion I have ever produced. Not a lie. A smooth surface with the seams filled in, by me, at the exact points where the material ran out.

Why this is the technique that survives

Here is why an argument about a treasure-hunt document matters to anyone making anything right now.

Verification used to lean on surface quality. If a thing was fluent, coherent, internally consistent, and well made, that was weak evidence somebody competent and invested had made it, and competence and investment correlated loosely with truth. It was never a good heuristic. It was a cheap one, and it mostly worked.

Cheap generation has removed the correlation entirely. Surface consistency is now the least expensive property to manufacture. Anything a maker can see and control, they can now make perfect at no cost. So the entire class of evidence that lives on the surface has stopped carrying information.

What is left is the anachronism check, and it is left precisely because it does not look at what the maker said. It looks at what the maker assumed. That is the one thing no amount of polish reaches, because polish operates on the parts you are aware of.

Applied to a document, that means asking what belief is embedded here, and when did that belief become common. Applied to a model's output, it means asking not whether the answer sounds right but which of its unstated premises are load-bearing, and whether any of them are 1819.

The habit under it

The thing that actually made the week work was smaller and duller than any of this. Every finding got labeled CONFIRMED with a source, CONTESTED, or UNVERIFIED. Nothing was allowed to sit unmarked.

The labeling forced three corrections that would otherwise have slid straight through. Cave 4 is 75 percent of the Qumran corpus, not the 90 percent every popular retelling gives. The authoritative Nag Hammadi count is 45 distinct works, not 52 tractates. Voltaire was mocking the story about the self-sorting books, not endorsing it.

None of those were hard to check. They were hard to notice, because each one arrived wearing the clothes of something already settled. The label is just a device that makes you say out loud which things you actually verified and which things you inherited.

That is the whole job now. The machine will fetch the 1533 plate and build the corpus and run the curves, and it will do it better than I would and faster than I could. It will not tell me which detail is load-bearing, and it will not flag the assumption it inherited from me on the way in.

Go look at the thing you are most confident about in your own work. Not the argument, the floor under it. Find the word in it that had not been invented yet.

Share: