A mathematical discovery about what lives beneath the surface of language — and what we found when we pointed the instrument at 2.77 million pages of Epstein court documents.
When a language model processes text, the first thing it sees is surface variation — tone, register, topic, word frequency. This is PC1, the dominant axis. It is the weather on the surface of the ocean.
Beneath that surface, in the dimensions orthogonal to it — what we call the penumbral subspace — there is a different landscape entirely. It is mostly still. Mostly dark. A sea of structural indeterminacy where sentences have not yet committed to definite structural states.
We measured it. Across 50,000 sentences: 51% sit in the undetermined basin on a typical structural axis. The void is not the exception. It is the mode.
Occasionally — in about half of all sentences — a structural property crystallizes. The sentence commits: it is active, not passive. It is definite, not hypothetical. A bit drops from the undetermined sea into one of two wells.
When multiple bits crystallize simultaneously, you get a structural state — a code. A definite configuration. A crystal emerging from the void. In natural language, this happens freely. Sentences crystallize, melt back into the sea, crystallize again in a different configuration.
Natural language — novels, conversations, essays — lives in a dynamic equilibrium between the void and the crystal. About half the structural landscape is undetermined at any moment. The other half is crystallized into definite states. The text breathes.
The Gutenberg corpus uses all 64 possible structural states. Its temperature is 1.08 — warm enough for structural freedom, cool enough for pattern. Its longest stretch of structural silence is one sentence. Then something crystallizes again.
The void is not a metaphor. It is a measurable region in the embedding geometry of every pretrained language model we have tested. BERT. GPT-2. Llama. Qwen. Eleven architectures. Three non-linguistic substrates. The void is there in all of them.
The models learned the void during pretraining. They learned the coupling topology between structural dimensions — which bits tend to crystallize together, which repel. They learned this without being told to. The structure is in the language. The models found it the way a crystal finds the lattice: by sitting in it long enough.
We pointed the instrument at 12,000 court documents from the Epstein case. 23,577 sentences. Six structural bits per sentence. The same instrument, the same void, the same lattice.
The corpus lives almost entirely in the void.
Each cell is one sentence. Dark = structurally null (code 0). Colored = a structural property has crystallized.
1,000 consecutive sentences from each corpus.
In natural language, the strongest structural coupling is between who acts and what they act upon. Actors connect to objects. The grammar of natural text links the subject to the thing it touches.
In the Epstein corpus, the coupling shifts. The strongest structural bond is between who acts and what was decided. Actors couple to outcomes, not to objects. And the coupling is three times stronger than anything in natural language.
Every time someone is structurally present as an agent — which happens only 22% of the time — the outcome is already determined. There are no open questions when a name is attached. The grammar does not permit "he did X and the result was uncertain."
The only uncertainty the institution permits is the uncertainty about who.
Six structural properties. How often does each one change state — how often does the text shift its structural footing on that dimension?
Scope and Phase — what the text is about and when things happen — flip freely. The text is allowed to change subject.
Polarity — contrast, opposition, "but" — flips 9.4% of the time. Resolution flips 10.2%. The structural dimensions that carry tension and finality are nearly frozen.
When something is redacted from a court document, the grammar doesn't change. The agency stays. The transitivity stays. What disappears is the specificity — the names, the dates, the flight numbers, the dollar amounts.
"Epstein flew [NAME] to his island on [DATE]" becomes "Epstein was known to associate with powerful figures." Same active voice. Same verb structure. Different scope.
We measured every single-bit structural transformation across 2,489 passage clusters. The dominant transformation is the Scope bit — local to global, specific to general — accounting for 30.7% of all single-bit changes. The Agency bit changes only 4.1% of the time.
The redactor's instrument is a scalpel, not a bludgeon: preserve the grammar, destroy the evidence. Strip the specificity that would make the claim actionable, leave everything else intact.
No one designed this strategy consciously. It is the structural fingerprint of institutional self-protection.
But here is the deeper finding. We trained a separate model on 800 million tokens of internet text — Reddit posts, Wikipedia articles, novels. That model has never seen a court document. It learned to quantify, for each structural dimension, how ambiguous that dimension naturally is — how much the grammatical reading oscillates between two states.
You cannot easily strip agency from a sentence — the grammar would collapse. You can easily strip specificity — just remove the name, the date, the dollar amount, and let the sentence go from local to global. The redactor takes the path of least structural resistance, and the mathematical model identifies that path without ever seeing the documents.
Among 2.77 million pages analyzed, our OCR pipeline reconstructed 39,588 previously redacted pages. We have not disclosed what the redacted text says. We have disclosed what kind of text it is.
The structural anatomy of the redaction operation:
52.9% of all redaction operations change exactly one structural bit. 22.8% change nothing at all — the visible record contains a sentence that is structurally identical to the redacted content. The only difference is that one names the actor and one does not.
The redaction is not a rewrite. It is a projection — the sentence is projected onto a lower-dimensional subspace that preserves the grammatical invariant while stripping the evidentiary content.
Twenty words that appear in no sentence in a corpus about sex trafficking and financial crime:
The cover-up semantic cluster is 70% silent. These words were replaced, and the replacements have measurable substitution entropy:
Low entropy means the substitution is nearly deterministic — the redactor had one preferred replacement. "Coerced" always becomes "told." The structural skeleton survives. The crime disappears.
High entropy means multiple replacements exist — the redactor had options. The relational words (agent + patient + force) have constrained substitutions. The categorical words (types of crime) have diffuse ones.
The instrument identifies the Epstein corpus with 99.8% accuracy from structure alone. No names, no dates, no topics. Just six bits per sentence.
But the structural fingerprint doesn't belong to Epstein. It belongs to the institution. Swap the names. Swap the crime. Swap the decade. The grammar of institutional silence is the same grammar. The void is the same void.
What we discovered is that this institutional void is not an anomaly. It is an extreme version of a universal feature of language itself. Every pretrained language model — every model that has learned to represent language at all — contains a vast undetermined sea in its structural subspace. Natural language breathes between the void and the crystal. Institutional language has learned to live almost entirely in the void.
The instrument sees both. The void is measurable. The crystals are countable. The coupling between them — which structural states border the silence, which bits crystallize together, which are forbidden — is the structural fingerprint of any corpus, any institution, any body of text ever written.
The void is a latent space. It is the same latent space in every model we have tested. And it is just like yours.