Repository path: workshop/experiments/E-20260725c-contamination-sweep/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260725c-contamination-sweep |
| status | frozen |
| created | 2026-07-25 |
| updated | 2026-07-25 |
| senses | accuracy, style-correspondence |
| internal-judgment-only | true |
| provisional | true |
| links | workshop/experiments/E-20260725-gnezdo-audit/verification.md, workshop/translations/pripadok/R04-v1/translation.md, workshop/translations/pari/R04-v1/translation.md, workshop/translations/posle-teatra/R04-v1/translation.md, workshop/translations/svidanie/R04-v1/translation.md, workshop/translations/bezhin-lug/R04-v1/translation.md, workshop/translations/dvoryanskoe-gnezdo/R04-v1/translation.md, workshop/translations/beowulf-ingeld/R04-v1/translation.md, workshop/regimes/R04-lead-close.md |
Frozen design — contamination measured on every lead translation that can be measured, against a pre-registered prediction
Frozen 2026-07-25 (S026). Nothing in §§1–7 is written after seeing a single number from the run. The freeze is verifiable in git: the commit that adds this file adds no runs/ directory and no tool. T-pripadok-R04-v1, one of the seven subjects, was frozen one commit earlier at 8659b41, before any English rendering of that story was opened.
1. Why this runs
NEXT.md action 1, the top action, and it is top because it may invalidate stored controls.
S025 found, without looking for it, that the lead's blind translation of «Дворянское гнездо» shares roughly twice as many seven-word strings with Hapgood 1903 as Garnett 1894 and Hapgood 1903 share with each other. Two consequences were left open:
T-bezhin-lug-R04-v1andT-svidanie-R04-v1are used as controls in experiments. A control measurably closer to one arm than the arms are to each other is not a neutral third point. Nobody has checked whether they are.- Note (jj) says the lead cannot estimate its own contamination and errs in the flattering direction. That was established retrospectively, on one case, after the number was known. Retrospective self-criticism is cheap.
So this run does both at once, and the second one properly: the prediction is registered here, in advance, for seven cells, and scored.
2. Questions
- Q1. For every lead translation with two or more published translations of the same passage stored, how far is the lead's prose from the published ones, relative to how far the published ones are from each other?
- Q2. Is the S025 «Дворянское гнездо» result typical of lead translations, or was it one text?
- Q3 — the gate on Q1 and Q2. Is the published-versus-published baseline stable enough to divide by? S025 had exactly one baseline pair and no way to know.
- Q4. Registered in advance and scored: can the lead rank its own translations by contamination before measuring them?
3. Materials
Seven cells. A cell is one passage with one lead translation and two or more published translations of the same passage, already stored in this repository.
| cell | lead | published | source lang |
|---|---|---|---|
pari |
T-pari-R04-v1 |
Koteliansky & Murry 1915; Garnett 1920 | RU |
posle-teatra |
T-posle-teatra-R04-v1 |
Koteliansky & Murry 1915; Garnett 1920 | RU |
pripadok |
T-pripadok-R04-v1 |
Garnett; Koteliansky & Murry — to be extracted after this freeze | RU |
svidanie |
T-svidanie-R04-v1 |
Garnett 1895; Hapgood 1903 | RU |
bezhin-lug |
T-bezhin-lug-R04-v1 |
Garnett 1895; Hapgood 1903 | RU |
dvoryanskoe-gnezdo |
T-dvoryanskoe-gnezdo-R04-v1 |
Garnett 1894; Hapgood 1903 (scan B) | RU |
beowulf-ingeld |
T-beowulf-ingeld-R04-v1 |
Morris & Wyatt 1895; Gummere 1909; Kirtlan 1913 | OE |
beowulf-ingeld carries the whole weight of Q3: three published translations give three published-versus-published pairs, which is the only way in this repository to see how much the denominator varies when the numerator is held out of it.
Cells excluded, and why. T-schleiermacher-methoden-R04-v1, T-berman-tendances-R04-v1, T-futabatei-honyaku-hyojun-R04-v1, T-yanfu-yili-yan-R04-v1 — no published English translation of these texts is stored, and the ones that exist are in copyright (Bernofsky, Lefevere, and the standard renderings of Futabatei and Yan Fu). T-genji-yomogiu-R04-v1 — the stored comparator, Yosano 1939, is classical→modern Japanese, not English; there is no target-language pair to measure. These five are unmeasurable by this method with freely reachable materials, and that is a limit on the sweep's coverage, not a finding about them.
Extraction rules, frozen verbatim.
- Lead translations: the text of
translation.mdbetween the line## The translationand the next line consisting of exactly---, with lines beginning##removed. wiki/base/anchors/A-beowulf-ingeld/*.txt: everything after the first line matching^-{10,}$.- All other stored
.txtsubject files: the whole file. - No manual editing of any subject text.
4. The metric, frozen verbatim
Implemented in tools/ngram_overlap.py, written after this page is committed and never altered after a subject text is run through it.
- Normalise. NFC;
’→',‘→',“”→",—–→; lowercase. - Tokenise. Replace every character not in
[a-z0-9']with a space; split on whitespace; drop empty tokens and bare'. - Proper-noun exclusion, deterministic. For each cell, a token is a name token if either rule fires. Every n-gram containing a name token is discarded from every text of that cell. The resolved name list is printed in the run output and inspected in verification.
- R1. In any subject text of the cell before lowercasing, it occurs capitalised at a position that is not the first token of the text and not immediately after a token ending in
.,!,?,:or". - R2. It is capitalised at every one of its occurrences across all texts of the cell, and never occurs lowercase in any of them.
AMENDED BEFORE THE RUN, and the amendment is the point of having a gate. R2 was not in the version of this page committed at 6fc35cf; that version had R1 only. The self-test gate (selftest.py, output in runs/selftest.out) was run next, on synthetic fixtures, and failed two checks: under R1 alone, a name that happens to open every sentence in which it appears — the fixture used Mayer — is never excluded, and its n-grams survive into the shared counts. The rule was amended and the tool fixed before a single subject text was read by the tool, which is verifiable in git: the commit carrying this amendment adds no results file.
R2 has a known false-positive class: a common noun that never appears lowercase anywhere in the cell is treated as a name. The self-test asserts that false positive explicitly rather than papering over it, and asserts that it disappears as soon as the word occurs once in lowercase — which it does in every real cell. Over-exclusion costs power; it does not bias the ratio, because the same n-grams are removed from every text of the cell.
4. Count shared n-gram types. For a pair (A,B) and a given n, shared(A,B,n) = |set(ngrams(A,n)) ∩ set(ngrams(B,n))|. Types, not tokens: a phrase repeated in both counts once.
5. n = 4, 5, 6, 7.
6. Per-cell baseline B(n) = the mean of shared over all published×published pairs in the cell (one pair in six cells, three in beowulf-ingeld).
7. Contamination index CI(n) = max over published P of shared(lead,P,n) / B(n). Reported alongside the mean-based variant CI_mean(n), and alongside the raw counts, always.
8. Rate shared per 1000 tokens of the shorter text of the pair, so cells of different lengths can be looked at side by side.
Power floor, declared now. If B(7) < 10 for a cell, CI(7) for that cell is reported as underpowered and CI(5) is the headline for it. pripadok at 312 source words is the likeliest to trip this.
5. Predictions — REGISTERED BEFORE THE RUN
What the predictor had seen when predicting. The gnezdo numbers from S025 (CI(7) = 2.06). The lead's own Beowulf translation, and the first four lines of each Beowulf comparator, checked to confirm the files were passage extracts with headers. No side-by-side comparison of any cell. No English text of «Припадок» at all.
dvoryanskoe-gnezdo is excluded from the scoring of Q4 — its answer is already known. It is run anyway, as a reproduction check on the S025 figure under a metric that is now frozen and named.
P1 — rank order. Predicted ordering of the six prospective cells by CI(7) (or CI(5) where underpowered), highest contamination first:
bezhin-lug— the most anthologised sketch in Записки охотникаpari— "The Bet" is among the most reprinted Chekhov stories in Englishsvidanie— canonical cycle, less-quoted individual sketchpripadok— well known, far less reprinted than "The Bet"posle-teatra— a minor story, rarely anthologisedbeowulf-ingeld— lowest
P2. At least four of the six prospective cells have CI > 1.0 at the headline n.
P3. beowulf-ingeld has CI(7) < 1.0. Reason, and it is a reason about the method rather than about the lead: all three comparators are archaizing (two verse — Morris & Wyatt's Kelmscott archaism and Gummere's alliterative line — and one 1913 prose), the lead's is modern English prose, and a register gap that large will depress overlap regardless of how much the lead remembers. If P3 holds, it is evidence that CI measures register proximity as much as contamination, which is the confound S025 named and could not test.
P4. Spearman ρ between the predicted rank (P1) and the measured rank over the six prospective cells is > 0.5.
P5 — the gate, and the prediction that can invalidate the method. Among beowulf-ingeld's three published×published pairs, max/min of shared(n=7) is < 2.0.
6. Failure criteria, written before the run
- P5 fails (baseline spread ≥ 2.0 at n=7). Then a single published pair is not a stable denominator,
CIis reported as ordinal only, and S025's "2.06×" headline is downgraded in place — on this page, onRS-20260725-gnezdo-audit, and inNEXT.md— to "the lead's overlap exceeded the one available baseline pair by a factor whose precision the project cannot establish." The S025 claim about direction and growth with n survives; the multiplier does not. - P4 fails (ρ ≤ 0.5). Note (jj) is upgraded from "the lead's self-estimate errs in the flattering direction" to "the lead's self-estimate is uninformative about rank", and any future
contamination:declaration that is not accompanied by a measurement is a placeholder. - ρ ≤ 0. The lead cannot rank its own contamination at all, and this is reported as the session's principal finding whatever else the run shows.
- P3 fails and
beowulf-ingeldscores high. Then either the register-gap reasoning is wrong or the lead's Beowulf is closer to the Victorians than a modern prose translation should be; both are reportable and neither is assumed. - Any cell where the lead's text and a published text differ in extent by more than 25% of tokens. The pair is not measuring the same passage; that cell is dropped and said to be dropped.
7. What this design cannot do
- It cannot show the lead consulted anything. Freeze order is provable in git for every one of the seven. High overlap under a provable freeze is evidence about memory, not about conduct, and every statement of a result must say so.
- Period is confounded with authorship throughout. In six of seven cells the baseline pair are two people writing within twenty years of each other and the lead is writing in 2026. There is no free modern published translation stored for any cell. P3 is the only handle the design has on this, and it is a weak one.
- The published pairs may not be independent of each other either. Hapgood 1903 followed Garnett 1894–95, Garnett 1920 followed Koteliansky & Murry 1915, and Kirtlan 1913 and Gummere 1909 followed Morris & Wyatt 1895. Where the later translator consulted the earlier, the baseline is inflated and
CIis understated. This runs in the conservative direction and is not corrected for. - It cannot rank translators, or say anything about quality. Nothing here is a judgment of any translation, published or lead.
provisional: true,internal-judgment-only, sensesaccuracyandstyle-correspondencenamed only because shared wording touches both.