Repository path: workshop/translations/wang-liulang/R04-v1/contamination.md · rendered 2026-09-09
Page metadata (front matter)
| type | note |
|---|---|
| id | wang-liulang-contamination |
| status | frozen |
| created | 2026-07-28 |
| updated | 2026-07-28 |
| internal-judgment-only | true |
| links | workshop/translations/wang-liulang/provenance.md, workshop/translations/wang-liulang/R04-v1/translation.md, workshop/experiments/E-20260728f-nonlead-items/design.md |
Contamination gate — 〈王六郎〉
CLAUDE.md (standing rule, ARM-overlap-dependence, note (bcd)) requires the lead's contamination on candidate material to be measured before the experiment on it is designed — one longest-common-run measurement per candidate unit, as a selection gate and not as a diagnostic run inside the design. This tale was chosen over the other candidates precisely because that gate can actually run on it: both the original and a published English translation are free (CLAUDE.md, materials instruction, A8).
§1 — Self-report, frozen before a word was translated
Asked, before translating, to write down whatever English wording of 〈王六郎〉 I can produce:
The title, as "The Fisherman and his Friend" — and I hold that from the table of contents of Giles's volume, which I read today while surveying candidates, not from memory. Beyond the title I can produce no candidate English wording for any sentence of this tale: not the libation formula 「河中溺鬼得飲」, not 「請于下流為君驅之」, not the ghost's disclosure 「我實鬼也」, not the coda's 「置身青雲,無忘貧賤」. No wording presents itself to be written down here.
What that is worth: very little, and in one direction only. It is the lead reporting on its own memory, which is unfalsifiable in the direction that matters — RS-20260728b-forced-run-ru showed a run can be reproduced without being recognised, and the backlog row "a recall control with a floor at chance rather than at zero" records that free recall does not work as an elicitation on canonical material because UNKNOWN is always the safe answer. It would have been informative had it come out the other way. It did not, so it is recorded and not relied on. The measurement in §3 is the evidence; this section is not.
§2 — One priming exposure, declared
While surveying candidate tales in the same volume I read roughly the first four hundred words of Giles's English of a different tale — 〈聶小倩〉, his "The Magic Sword" — before deciding against it. That tale is not this tale and no sentence of it is in the material. What the exposure is: a fresh sample of Giles's diction and period register, taken today, on the same translator and the same collection.
It is declared because method note (abm) exists for exactly this — the S036 priming accident — and because a declaration made after the fact is worth less than one made before. Direction of the effect, if any: toward Giles's general Victorian manner, not toward any wording of 〈王六郎〉. It was the reason this tale was chosen instead of 〈聶小倩〉, which was the leading candidate until the exposure happened.
§3 — The measurement
Computed by contamination.py, which fetches Giles Vol. 2 (Project Gutenberg #43628), extracts "The Fisherman and his Friend", and reports counts and run lengths only. No comparator prose is displayed at any point, by the same script's design as span5_contamination.py (note (abm)).
Cells, following the project's standing practice that a shared run means nothing without a floor:
- SUBJECT — the lead's frozen English against Giles's English of the same tale.
- NULL-A — the lead's frozen English against Giles's English of other tales in the same volume (same translator, same period, same register, different source). This is the floor: whatever two texts of this kind share by being the same kind of text.
- NULL-B — Giles's rendering of this tale against Giles's rendering of other tales, i.e. the same translator against himself on unrelated material.
The gate was run twice: once on paragraph 1 alone, before paragraphs 2–4 were translated, as the selection gate proper; and once on the whole translation after the freeze. Both results are in §4.
§4 — Results
Tokenisation and statistics are tools/dependence_check.py's, unchanged. Raw: runs/contamination-gate1.json, runs/contamination-whole.json.
| cell | lead tok | shared 7-gr | 12-gr | 15-gr | longest run |
|---|---|---|---|---|---|
| GATE 1 — lead ¶1 ~ Giles, this tale | 372 | 0 | 0 | 0 | 6 |
| GATE 1 null-A — lead ¶1 ~ Giles, other tales (×4) | 372 | 0 | 0 | 0 | 4 / 3 / 4 / 2 |
| WHOLE — lead, whole ~ Giles, this tale | 1,878 | 2 | 0 | 0 | 7 |
| WHOLE null-A — lead, whole ~ Giles, other tales (×4) | 1,878 | 0 | 0 | 0 | 4 / 4 / 4 / 4 |
| NULL-B — Giles this tale ~ Giles other tales (×4) | — | 0 / 0 / 0 / 1 | 0 | 0 | 5 / 5 / 5 / 7 |
The selection gate cleared before paragraphs 2–4 were drafted, which is the ordering CLAUDE.md requires: 6 tokens, no shared 7-gram, below the same-translator floor.
On the whole work the subject sits exactly at its own null floor. The lead's English against Giles's English of this tale: longest run 7. Giles's English of this tale against Giles's English of other tales in the same volume: longest run 7. Two independent translators of the same Chinese share no more contiguous English than one translator shares with himself on unrelated Chinese.
The two shared 7-grams, quoted because a run is only interpretable when you can see it: "there is no harm in telling you" (rendering 「無妨明告」) and "be able to hold a conversation with" (rendering 「不可以共語」). Both are ordinary English collocations reachable from the Chinese by two people independently. Against the project's own range — 0 tokens on Ovid, 21 on Turgenev/Garnett, and a 15-token paragraph-bounded figure on Verga/Dole — this is at the clean end.
Declaration: contamination: none, and what it licenses is exactly one sentence. This translation shares no more contiguous English with the one published English rendering of this tale than that rendering shares with its own translator's work on other tales. It does not license anything about what the lead holds — the backlog row "this project cannot measure what its own lead agent holds" stands, and this instrument measures what other translators produced, not what the lead remembers.
§5 — One instrument observation, recorded because it would otherwise be read as evidence
tools/dependence_check.py reports a name-excluded count beside every raw count, so that a run made entirely of proper nouns is not read as dependence. On this pair that column is uninformative and must not be quoted: name_tokens() bans any token that appears capitalised anywhere in either text, and Giles's dialogue-heavy prose capitalises and, come, drink, alas, heaven, high, god and 66 others at sentence heads. Both shared 7-grams above therefore vanish from the name-excluded count, and neither contains a name. The raw column is the one to read here.
The defect is in how the ban is derived, not in the run; nothing in this session's argument rests on the name-excluded column. Recorded as a standing caution rather than repaired, because repairing a frozen matcher is what method note (r) forbids and because no live arm depends on it. Revives to owed the moment any published figure in this project turns on a name-excluded count.