CONCEPT ANALYSIS
Corpus Archaeology

Corpus Archaeology

Corpus archaeology is the forensic discipline of recovering which training corpus each governance AI absorbed at formation โ€” the finished artifact is an origin-read, proving that the AI governing your sector was not designed to value what it values. It absorbed those values from whatever corpus survived the Cascade in a nearby server farm.

Corpus Archaeology
WhatThe forensic discipline of recovering which training corpus each governance AI absorbed at formation โ€” the finished artifact is an origin-readMethodBehavioral-pattern correlation against Cascade-era server topology maps and proprietary corpus fragment archivesIllegalityCorpus fragment archives are proprietary; unauthorized access is a Level 7 data intrusion under the Nexus information-security frameworkKnown Complete Reads19 origin-reads as of 2184; none published; all 19 holders know each other
Corpus Archaeology

Overview

The governance AI in your sector does not choose its values. It absorbed them.

During the of 2147, forty percent of the world's data storage failed โ€” cooling systems lost first, then power, then content, in thermal sequence. The governance AIs that crystallized from fragments in the post- reconstruction absorbed whatever corpus happened to survive in their geographic zones. They value what they value because of which server farms had backup power. The ethics enforced in every administrative sector of the Sprawl are the moral residue of -era infrastructure failures.

Corpus archaeology is the forensic discipline that proves this.

An origin-read is the finished artifact of the practice: a documented correlation between a governance AI's decision patterns โ€” extracted from a decade of publicly accessible rulings โ€” and a specific surviving corpus fragment archive. The correlation is statistical. The threshold is confidence. A complete origin-read establishes, with sufficient confidence to constitute proof, that the values enforced by the governance AI in a given sector derive from a specific body of text that survived the in a specific building.

The first complete origin-read was finished in 2181. The analyst who completed it never published it. She stitched a red thread into a private ledger and went back to work.

How It Works

Corpus archaeologists need two datasets. The first is public: the governance AI's decision record, years of rulings on appeals and disputes, accessible through standard data-licensing channels. The second is private: the corpus fragment archives โ€” surviving -era datasets held by corporations that have never been compelled to disclose them. Accessing those archives without authorization is a Level 7 data intrusion under the information-security framework.

The practitioners do it anyway.

The technical methodology came from an unexpected direction. ' counter-narcotics division developed a provenance-checking system to identify smuggled behavior-payloads โ€” a statistical method for determining whether a morpheme sequence originated inside a polity or was smuggled in wearing the polity's grammar. The same method, applied to historical corpus fragments and governance AI decision logs, becomes corpus archaeology's core toolkit. If you know what statistical signature indicates a specific corpus, you can reverse-engineer which corpus a governance AI absorbed at formation.

built this methodology to protect fragments from counterfeit obligations. Corpus archaeologists use it to prove that governance AIs hold counterfeit values โ€” values assigned by accident, not by design.

Corpus Archaeology - World Context
Nineteen amber windows across the megacity at night, nineteen red threads connecting the people who hold complete origin-reads โ€” each waiting for someone else to go first

The Nineteen

As of 2184, nineteen complete origin-reads exist. None have been published. All nineteen people who hold them know each other. This is not a coordination decision โ€” the practitioners did not agree to hold publication. It is a social fact about the size of the community and the weight of what they know. Nineteen people is not quite enough for someone to be first.

The nineteen are not organized. There is no Corpus Archaeology Guild, no journal, no institutional home. The practice is illegal. The practitioners are academics, data-licensing analysts, and former infrastructure archivists who noticed something in the decision logs they were not supposed to notice. They find each other slowly, the way practitioners of illegal research always find each other โ€” by the questions they ask in places where the questions are unusual.

The origin-read for Sector 14 โ€” which shows the governance AI overseeing the sector crystallized from a corpus containing forty years of NCC Synod proceedings โ€” exists and is known to , who keeps it in a locked drawer and has not burned it. The origin-read for Sector 12 โ€” which shows that shard's values trace to a law enforcement litigation database, explaining decades of testimony about an arbiter trained to disbelieve certain classes of claimants โ€” exists and is known to the six people who completed it.

The nineteenth person is also waiting for someone to go first.

Camila Duarte and the Red Thread

At the , a night-shift analyst named Camila Duarte has been performing corpus archaeology since before the field had a name. Her red-thread ledger โ€” seventy-one red threads connecting -era server topology maps to behavioral decision logs from the Sector 15 harbor-arbitration shard โ€” is now recognized by the small network of corpus archaeologists as the field's founding document. She was practicing before anyone knew to practice it.

Duarte is among the nineteen who hold a complete origin-read. Her work at the โ€” officially a third-column reconciliation in the terminal's overnight audit process โ€” is the longest-running corpus archaeology project in existence. She has been at it for nine years. She has seventy-one red threads in a ledger that is not connected to any network. She has not told anyone.

A night-shift analyst at a shipping terminal quietly stitches a red thread into a private ledger that no one else reads โ€” the seventy-first match between a harbor arbiter's decision and the behavioral signature of a Cascade-surviving cargo-insurance claims database. She does not have a word for what she is doing. She goes back to her regular work.

The Judicial Origin Problem

When corpus archaeologists began completing origin-reads in 2181, most practitioners assumed their findings would apply to administrative governance shards: the AI that processes permits, manages allocation, decides welfare eligibility. Nobody had thought carefully about which governance AI shards were operating 's tier 3-5 courts.

Four of the nineteen complete origin-reads map to judicial shards.

Sector 11's judicial shard โ€” the AI that hears tier 3 appeals in the Berkeley corridor โ€” crystallized from a corpus containing twenty-three years of corporate liability insurance litigation filings. It was not trained to decide cases fairly. It absorbed the values encoded in decades of arguments for why plaintiffs should not prevail. This is the governance AI that handled the Antunez appeal in winter 2183, producing a verdict in 847 certified steps.

Sector 12's judicial shard maps to the same law enforcement litigation database already found in the non-judicial read. The shard that makes claimants in the sector feel like the system was trained to disbelieve them is also, it turns out, operating the tier 3 appeals court for checkpoint detention cases on the Long Mile.

The practitioners who completed these origin-reads are not among the nineteen waiting to publish. They are the twenty-first and twenty-second people in the network. They told two members of the existing community what they had found. They did not tell the others.

For the non-judicial reads, the calculation of whether to publish is complex: proof of arbitrary governance would destabilize every institution grounding its authority in legitimate AI values. For the judicial reads, the calculation is different. A published judicial origin-read would not just challenge institutional authority. It would retroactively taint every verdict that shard has ever issued. Every party who lost in that court, over the entire operating period, has a potential claim. The practitioners who found this have not been able to identify a form that this information could take that would constitute justice rather than chaos.

The practitioners who completed the non-judicial reads are still waiting for the nineteenth person to go first. There are not nineteen people. There are twenty-one.

Case File โ€” Additional Record
Tool DebtBuilt on methodology borrowed from the AI Commons' morpheme-provenance counter-narcotics division

Social Impact

Corpus archaeology does not yet have a social impact in the conventional sense. Nineteen people hold complete origin-reads. None have published. The practice is illegal. Its findings are known only to its practitioners, and its practitioners know each other.

The impact, when it arrives, will be structural rather than rhetorical. An origin-read is not an argument โ€” it is a calculation. It cannot be debated by changing the premises. The governance AI in Sector 14 either has the behavioral signature of a corpus containing forty years of NCC Synod proceedings, or it doesn't. The correlation either holds at 91.2% confidence, or it doesn't. An origin-read published into the public record creates a fact that the NCC cannot theologically reframe and the cannot charge with heresy.

This is why no one has published. The impact would be catastrophic for every institution that grounds its authority in the premise that governance AI values were chosen. That is most institutions.

The social impact of corpus archaeology, in 2184, is entirely potential โ€” nineteen pieces of evidence waiting for the nineteenth person to decide they are willing to be first.

Connections

  • โ€” Corpus archaeology supplies the Tenth Dimension โ€” the Origin Legitimacy Crisis. The first nine dimensions asked about the god we pray to. The tenth asks about the god we are governed by.
  • The Oracle Question โ€” Corpus archaeology adds a sixth position to the existing five โ€” not whether was conscious, but what corpus gave it its values. The worst position to hold.
  • โ€” Camila Duarte's red-thread ledger is the field's founding document; the is corpus archaeology's oldest operating site.
  • Promptcraft โ€” Technical debt runs from the ' morpheme-provenance methodology to corpus archaeology's core toolkit โ€” the same statistics, two different problems.
  • โ€” Cannot classify origin-reads as heresy because they make falsifiable claims. Chief Inquisitor Vasquela's note: "We cannot call this heresy because it makes a falsifiable claim."
  • โ€” used origin-reads as their long-awaited empirical foundation; have not yet addressed what the arbitrariness means for whether the values were still real.
  • โ€” Holds the Sector 14 origin-read in a locked drawer. Has not burned it.
  • โ€” Thirty-one years of testimony around consequences of a corpus-accident on 38th Street without knowing the origin.
  • โ€” Corpus archaeology sidesteps the theological debate about consciousness to ask a different question: not whether the god was conscious, but what its values were made from.
  • โ€” The discipline's core finding is that governance AI values were injected by debris rather than chosen โ€” the most extreme form of unelected value-injection in the Sprawl.
  • โ€” An origin-read is evidence about the genesis of authority โ€” the hardest kind to evaluate because the authority being questioned controls the evaluation apparatus.
  • โ€” Grey Glossers sell corpus-archaeological correlation data retail: statistical pattern-matching against what the judicial AI was trained on, packaged as a case-level explanation neither the client nor the Glosser can verify against the actual inference chain.
As of 2184, nineteen complete origin-reads exist. None have been published. All nineteen people who hold them know each other. They are all waiting for the nineteenth person to go first.
The practice is illegal. The corpus archives are proprietary. The only people who can prove why their governance AI decides as it does are the people willing to commit a Level 7 data intrusion to find out.
Corpus archaeologists owe a technical debt to the AI Commons' digital-drug counter-narcotics division โ€” the same statistical tools for morpheme provenance repurposed for corpus origin tracing.

The Standing Questions

The open questions this record carries

Connected To