Repository path: workshop/translations/yingyi-jiejixing/contamination.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260728-forced-run |
| status | frozen |
| created | 2026-07-28 |
| updated | 2026-07-28 |
| senses | — |
| links | workshop/translations/yingyi-jiejixing/R06-v1/translation.md, workshop/translations/yingyi-jiejixing/R04-v1/translation.md, wiki/method-notes.md, wiki/findings/results/RS-20260726-ovid-period-form.md |
Contamination gate on T-yingyi-jiejixing-*, and the forced-vs-recalled control
Design frozen before dispatch. S044.
0. Why this page exists
Method note (bcd) makes the contamination measurement a selection gate run before the material is used, and the two sessions before this one could not run it at all — 「最後の一句」 has one English translation and it is unobtainable. This work is the case the note has been waiting for: Lu Xun 「硬譯」與「文學的階級性」 has a standard English translation (Yang Xianyi and Gladys Yang, Selected Works of Lu Hsun vol. III, Foreign Languages Press, 1959/1980) and the scan is freely readable in full, so the gate runs.
1. The measurement, as run
- Subject:
T-yingyi-jiejixing-R06-v1, the single-pass draft, frozen in git ata563d0bbefore any measurement was made and before the self-revision. 1,533 tokens. - Comparator: the Yangs' §V, sliced out of the archive.org scan
(
dli.ernet.53888, Lu Xun Selected Works Vol-iii (1959),_djvu.txt) by structural markers only — the standalone roman-numeral section markers at file offsets 31464 (V) and 39753 (VI) inside the essay slice 95719–137446. No comparator prose was displayed at any point before the draft was frozen, per method note (abm). OCR running heads and page numbers were dropped (9 lines) before measurement. 1,417 tokens. - Instrument:
tools/dependence_check.py(unmodified). - Storage: the comparator is a copyrighted translation and is not stored in the repository;
it lived in the session scratchpad and is gone with the container. Consultation is ledgered in
wiki/base/consulted.md.
| shared 7-grams | shared 12-grams | shared 15-grams | longest common run |
|---|---|---|---|
| 13 | 1 | 0 | 12 tokens |
The 13 shared 7-grams collapse to four distinct sites:
| # | run | source | tokens |
|---|---|---|---|
| L1 | but there is still a great deal of paper in the world | 然而世間紙張還多 | 12 |
| L2 | the colour of my teeth in a society like this | 罵到牙齒的顏色。在這樣的社會裏 | 9 |
| L3 | the plays of Hauptmann and Lady Gregory in | 印Hauptmann和Gregory夫人的劇本了 | 8 |
| L4 | to be of some use to society | 我也願意於社會上有些用處 | 7 |
L3 is not evidence of anything: both names stand in the Roman alphabet in the Chinese source, so any translator transcribes them. It is recorded because deleting it after seeing it would be selecting the data.
The name-excluded column is unusable on this pair and is not reported. name_tokens()'s known
false-positive class (R2: capitalised at every occurrence, never lowercase) resolves 80 tokens here,
including a, in, the, this, they, world, yes, surely, naturally, today — the
comparator is a 1,417-token slice in which many common words only ever open a sentence. The raw
counts and the longest run are the signal, which is what the tool's own docstring says.
2. Where 12 tokens sits
CLAUDE.md records the measured range across five lead translations: 0 tokens (Ovid, no canonical
English) to 21 tokens (Turgenev, against Garnett — canonical and unavoidable). 12 is squarely in
the middle of that range, and the Yangs are as close to canonical for Lu Xun in English as Garnett is
for Turgenev. That is enough to declare suspected and not enough to say why.
3. The control — is the run forced or recalled?
RS-20260726-ovid-period-form tested exactly this objection on Latin ("the Latin forces it") and it
failed there: the three highest-overlap pairs were the three sharing 12-word verbatim runs, and
the runs were not forced. The same objection deserves the same test here, on a new pair, and the
test is cheap.
Question. Do independent translators, working from the Chinese alone, produce L1, L2 and L4?
Materials. The four Chinese sentences containing L1–L4, each with one clause of surrounding context, and nothing else. No English is shown.
Procedure. One call each to P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash,
P3 x-ai/grok-4.5, temperature 0, asked to render each sentence into English prose. The three
models are independent of each other and of the lead. Raw responses under runs/.
Pre-registered scoring. A locus is FORCED if ≥2 of 3 independent outputs contain its run verbatim, or contain ≥7 of its tokens contiguously. Otherwise it is NOT FORCED.
Pre-registered predictions.
- P-a. L1 is FORCED. Every content word of 然而世間紙張還多 has one obvious English equivalent and the default English order matches; there is very little to choose.
- P-b. L4 is FORCED, for the same reason and with less at stake (7 tokens).
- P-c. L2 is NOT FORCED. 罵到牙齒的顏色 has real alternatives (abused the colour of my teeth / went as far as the colour of my teeth / even my teeth were called yellow), and 在這樣的社會裏 can be in such a society as easily as in a society like this.
Pre-registered failure criterion, and it binds the artifact. If L1 comes back NOT FORCED —
0 or 1 of 3 — the 12-token run is evidence of recall and T-yingyi-jiejixing-R04-v1 is declared
contamination: high, not suspected. If L1 is FORCED and L2 is NOT FORCED, the declaration
stands at suspected with the reading that the shared material is the constrained material.
A control that can only confirm the convenient answer is not a control, so the failing branch is
written here, before dispatch, and the artifact is bound to it.
What this cannot settle. Three model outputs are not three human translators, and all three share whatever the Yangs' text contributed to their own training. The probe can therefore show a run is forced; it cannot show a run is not shared for some other reason. It is a one-sided test and is used only in that direction.
4. Result — the failure branch fired
3 calls, $0.0166679, all finish_reason: stop. Scored mechanically by runs/score.py
against the frozen rule; the figure in each model column is the longest run of the locus's
tokens appearing contiguously in that model's output.
| locus | tokens | P1 | P2 | P3 | hits | verdict | predicted |
|---|---|---|---|---|---|---|---|
| L1 | 12 | 5 | 5 | 5 | 0/3 | NOT FORCED | FORCED ✗ (P-a false) |
| L2 | 10 | 2 | 3 | 2 | 0/3 | NOT FORCED | NOT FORCED ✓ (P-c holds) |
| L3 | 8 | 4 | 7 | 7 | 2/3 | FORCED | (not predicted) |
| L4 | 7 | 7 | 7 | 7 | 3/3 | FORCED | FORCED ✓ (P-b holds) |
(L2 is 10 tokens under dependence_check.py's tokenisation, not the 9 estimated in §1's table.)
The headline run is not forced. All three models rendered 世間紙張還多 as "there is still plenty of paper in the world". Not one produced a great deal of — the lead's phrase, and the Yangs'. Every model's longest contiguous match with the 12-token run is 5, well under the threshold of 7. Prediction P-a was wrong, and it was wrong in the direction that costs something: 「世間紙張還多」 has more freedom in English than the lead judged when writing the prediction, and the lead and the published translation took the same one of the available options.
L4 is the counter-case that makes the instrument readable. "to be of some use to society" came back verbatim from all three, so a 7-token shared run at that site is genuinely forced by 於社會上有些用處. The probe therefore separates the two loci rather than blessing or condemning them together, which is the only reason its verdict on L1 is worth anything. L3 is forced and uninteresting — both names stand in Roman script in the Chinese.
Applied to the artifact, per the frozen branch: T-yingyi-jiejixing-R04-v1 is declared
contamination: high. The declaration is on a measurement and a control, not on a prior.
What this does and does not establish. It establishes that the longest shared run is not explained by the source constraining the English, which is the defence the run would otherwise have. It does not establish the mechanism: recall of the Yangs' sentence, a shared idiom-frequency prior, and the two translators independently making the same one of several free choices all predict this result. The probe is one-sided by design (§3) and is used only in the refusing direction.
Why this is worth more than the declaration it produced. The project's standing rule that the
lead cannot serve as an independent third translator where a canonical English version exists has
rested since S031 on the Turgenev/Garnett 21-token run — asserted to be unforced, never tested.
RS-20260726-ovid-period-form tested the same defence on Latin and refused it. This is the second
independent refusal, on a different pair, and the first on a lead translation. It also shows the
test is cheap: three sentences, three calls, $0.017, and it can be run on any measured run the project
has ever recorded. Note (bcp).
5. Provenance
- Draft frozen at
a563d0bbefore any measurement was made. - Design §§0–3, including the failure criterion, frozen in git before the probe was dispatched; the freezing commit is the one that created this file.
- Probe run: 3 calls, temperature 0,
max_tokens3000. Raw underruns/.