A Model Trained on
Internet Text Predicted
the Redaction Strategy

Aerogel Press · March 2026 · Penumbrae LLC

We trained a mathematical model on 800 million tokens of internet text — Reddit posts, Wikipedia articles, scientific papers, novels. The model has never seen a court document. It learned to quantify, for each of six structural dimensions of language, how ambiguous that dimension naturally is — how much the grammatical reading oscillates between two states.

Separately, we measured which structural dimensions actually change in the Epstein court documents — which grammatical properties get stripped when a passage is redacted.

r = 0.90
p = 0.015
Correlation between model's structural uncertainty (trained on internet text) and observed flip frequency in Epstein redactions

The model trained on internet text predicts which structural dimensions of language are vulnerable to manipulation. The Epstein redactors independently found the same vulnerabilities and exploited them.

Structural vulnerability × observed manipulation
Observed flip rate (Epstein corpus) Predicted vulnerability (internet text model)

The logic is structural. You cannot easily strip Agency from a sentence — the grammar would collapse. "He flew them to the island" cannot be redacted into a passive construction without rewriting the sentence entirely. Agency is structurally rigid (model uncertainty: 0.17).

You can easily strip Scope. "He flew [NAME] to [LOCATION] on [DATE]" becomes "He was known to associate with powerful figures." Same active voice. Same verb structure. The specificity is gone but the grammar survives. Scope is structurally flexible (model uncertainty: 0.54).

The redactor takes the path of least structural resistance. The model identifies that path without ever seeing the documents.

Scope — the dimension that distinguishes specific from general, local from global, named from unnamed — accounts for 30.7% of all single-bit structural changes in the Epstein redactions. Agency accounts for 4.1%. The ratio is 7.5 to 1.

No one designed this strategy consciously. It is the structural fingerprint of institutional self-protection — the grammar of a system that has decided to preserve plausible narrative while destroying actionable evidence. And it is predictable from the structure of language itself.