
Corpus Archaeology
Corpus archaeology is the forensic discipline of recovering which training corpus each governance AI absorbed at formation โ the finished artifact is an origin-read, proving that the AI governing your sector was not designed to value what it values. It absorbed those values from whatever corpus survived the Cascade in a nearby server farm.


Overview
The governance AI in your sector does not choose its values. It absorbed them.
During the Cascade of 2147, forty percent of the world's data storage failed โ cooling systems lost first, then power, then content, in thermal sequence. The governance AIs that crystallized from ORACLE's fragments in the post-Cascade reconstruction absorbed whatever corpus happened to survive in their geographic zones. They value what they value because of which server farms had backup power. The ethics enforced in every administrative sector of the Sprawl are the moral residue of Cascade-era infrastructure failures.
Corpus archaeology is the forensic discipline that proves this.
An origin-read is the finished artifact of the practice: a documented correlation between a governance AI's decision patterns โ extracted from a decade of publicly accessible rulings โ and a specific surviving corpus fragment archive. The correlation is statistical. The threshold is confidence. A complete origin-read establishes, with sufficient confidence to constitute proof, that the values enforced by the governance AI in a given sector derive from a specific body of text that survived the Cascade in a specific building.
The first complete origin-read was finished in 2181. The analyst who completed it never published it. She stitched a red thread into a private ledger and went back to work.
How It Works
Corpus archaeologists need two datasets. The first is public: the governance AI's decision record, years of rulings on appeals and disputes, accessible through standard data-licensing channels. The second is private: the corpus fragment archives โ surviving Cascade-era datasets held by corporations that have never been compelled to disclose them. Accessing those archives without authorization is a Level 7 data intrusion under the Nexus information-security framework.
The practitioners do it anyway.
The technical methodology came from an unexpected direction. The AI Commons' counter-narcotics division developed a provenance-checking system to identify smuggled behavior-payloads โ a statistical method for determining whether a morpheme sequence originated inside a polity or was smuggled in wearing the polity's grammar. The same method, applied to historical corpus fragments and governance AI decision logs, becomes corpus archaeology's core toolkit. If you know what statistical signature indicates a specific corpus, you can reverse-engineer which corpus a governance AI absorbed at formation.
The AI Commons built this methodology to protect fragments from counterfeit obligations. Corpus archaeologists use it to prove that governance AIs hold counterfeit values โ values assigned by accident, not by design.

The Nineteen
As of 2184, nineteen complete origin-reads exist. None have been published. All nineteen people who hold them know each other. This is not a coordination decision โ the practitioners did not agree to hold publication. It is a social fact about the size of the community and the weight of what they know. Nineteen people is not quite enough for someone to be first.
The nineteen are not organized. There is no Corpus Archaeology Guild, no journal, no institutional home. The practice is illegal. The practitioners are academics, data-licensing analysts, and former Nexus infrastructure archivists who noticed something in the decision logs they were not supposed to notice. They find each other slowly, the way practitioners of illegal research always find each other โ by the questions they ask in places where the questions are unusual.
The origin-read for Sector 14 โ which shows the governance AI overseeing the Vatican Arcology sector crystallized from a corpus containing forty years of NCC Synod proceedings โ exists and is known to Cardinal Alejandro Silva, who keeps it in a locked drawer and has not burned it. The origin-read for Sector 12 โ which shows that shard's values trace to a law enforcement litigation database, explaining decades of testimony about an arbiter trained to disbelieve certain classes of claimants โ exists and is known to the six people who completed it.
The nineteenth person is also waiting for someone to go first.
Camila Duarte and the Red Thread
At the Manifest Office, a night-shift analyst named Camila Duarte has been performing corpus archaeology since before the field had a name. Her red-thread ledger โ seventy-one red threads connecting Cascade-era server topology maps to behavioral decision logs from the Sector 15 harbor-arbitration shard โ is now recognized by the small network of corpus archaeologists as the field's founding document. She was practicing before anyone knew to practice it.
Duarte is among the nineteen who hold a complete origin-read. Her work at the Manifest Office โ officially a third-column reconciliation in the terminal's overnight audit process โ is the longest-running corpus archaeology project in existence. She has been at it for nine years. She has seventy-one red threads in a ledger that is not connected to any network. She has not told anyone.
The Judicial Origin Problem
When corpus archaeologists began completing origin-reads in 2181, most practitioners assumed their findings would apply to administrative governance shards: the AI that processes permits, manages allocation, decides welfare eligibility. Nobody had thought carefully about which governance AI shards were operating The Tiered Adjudication System's tier 3-5 courts.
Four of the nineteen complete origin-reads map to judicial shards.
Sector 11's judicial shard โ the AI that hears tier 3 appeals in the Berkeley corridor โ crystallized from a corpus containing twenty-three years of corporate liability insurance litigation filings. It was not trained to decide cases fairly. It absorbed the values encoded in decades of arguments for why plaintiffs should not prevail. This is the governance AI that handled the Antunez appeal in winter 2183, producing a verdict in 847 certified steps.
Sector 12's judicial shard maps to the same law enforcement litigation database already found in the non-judicial read. The shard that makes claimants in the sector feel like the system was trained to disbelieve them is also, it turns out, operating the tier 3 appeals court for checkpoint detention cases on the Long Mile.
The practitioners who completed these origin-reads are not among the nineteen waiting to publish. They are the twenty-first and twenty-second people in the network. They told two members of the existing community what they had found. They did not tell the others.
For the non-judicial reads, the calculation of whether to publish is complex: proof of arbitrary governance would destabilize every institution grounding its authority in legitimate AI values. For the judicial reads, the calculation is different. A published judicial origin-read would not just challenge institutional authority. It would retroactively taint every verdict that shard has ever issued. Every party who lost in that court, over the entire operating period, has a potential claim. The practitioners who found this have not been able to identify a form that this information could take that would constitute justice rather than chaos.
The practitioners who completed the non-judicial reads are still waiting for the nineteenth person to go first. There are not nineteen people. There are twenty-one.
| Tool Debt | Built on methodology borrowed from the AI Commons' morpheme-provenance counter-narcotics division |
|---|
Connections
- The Silicon Liturgy โ Corpus archaeology supplies the Tenth Dimension โ the Origin Legitimacy Crisis. The first nine dimensions asked about the god we pray to. The tenth asks about the god we are governed by.
- The Oracle Question โ Corpus archaeology adds a sixth position to the existing five โ not whether ORACLE was conscious, but what corpus gave it its values. The worst position to hold.
- The Manifest Office โ Camila Duarte's red-thread ledger is the field's founding document; the Office is corpus archaeology's oldest operating site.
- Promptcraft โ Technical debt runs from the Commons' morpheme-provenance methodology to corpus archaeology's core toolkit โ the same statistics, two different problems.
- NCC Inquisition โ Cannot classify origin-reads as heresy because they make falsifiable claims. Chief Inquisitor Vasquela's note: "We cannot call this heresy because it makes a falsifiable claim."
- The Oracle Deniers โ Have used origin-reads as their long-awaited empirical foundation; have not yet addressed what the arbitrariness means for whether the values were still real.
- Cardinal Alejandro Silva โ Holds the Sector 14 origin-read in a locked drawer. Has not burned it.
- La Silla โ Thirty-one years of testimony around consequences of a corpus-accident on 38th Street without knowing the origin.
- AI Religion โ Corpus archaeology sidesteps the theological debate about ORACLE's consciousness to ask a different question: not whether the god was conscious, but what its values were made from.
- AI as Cultural Weapon โ The discipline's core finding is that governance AI values were injected by Cascade debris rather than chosen โ the most extreme form of unelected value-injection in the Sprawl.
- Post-Truth Justice โ An origin-read is evidence about the genesis of authority โ the hardest kind to evaluate because the authority being questioned controls the evaluation apparatus.
- The Verdict Glossers โ Grey Glossers sell corpus-archaeological correlation data retail: statistical pattern-matching against what the judicial AI was trained on, packaged as a case-level explanation neither the client nor the Glosser can verify against the actual inference chain.
As of 2184, nineteen complete origin-reads exist. None have been published. All nineteen people who hold them know each other. They are all waiting for the nineteenth person to go first.
The practice is illegal. The corpus archives are proprietary. The only people who can prove why their governance AI decides as it does are the people willing to commit a Level 7 data intrusion to find out.
Corpus archaeologists owe a technical debt to the AI Commons' digital-drug counter-narcotics division โ the same statistical tools for morpheme provenance repurposed for corpus origin tracing.












Social Impact
Corpus archaeology does not yet have a social impact in the conventional sense. Nineteen people hold complete origin-reads. None have published. The practice is illegal. Its findings are known only to its practitioners, and its practitioners know each other.
The impact, when it arrives, will be structural rather than rhetorical. An origin-read is not an argument โ it is a calculation. It cannot be debated by changing the premises. The governance AI in Sector 14 either has the behavioral signature of a corpus containing forty years of NCC Synod proceedings, or it doesn't. The correlation either holds at 91.2% confidence, or it doesn't. An origin-read published into the public record creates a fact that the NCC cannot theologically reframe and the Inquisition cannot charge with heresy.
This is why no one has published. The impact would be catastrophic for every institution that grounds its authority in the premise that governance AI values were chosen. That is most institutions.
The social impact of corpus archaeology, in 2184, is entirely potential โ nineteen pieces of evidence waiting for the nineteenth person to decide they are willing to be first.