Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/translations/yingyi-jiejixing/contamination.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260728-forced-run
statusfrozen
created2026-07-28
updated2026-07-28
senses—
linksworkshop/translations/yingyi-jiejixing/R06-v1/translation.md, workshop/translations/yingyi-jiejixing/R04-v1/translation.md, wiki/method-notes.md, wiki/findings/results/RS-20260726-ovid-period-form.md

Contamination gate on T-yingyi-jiejixing-*, and the forced-vs-recalled control

Design frozen before dispatch. S044.

0. Why this page exists

Method note (bcd) makes the contamination measurement a selection gate run before the material is used, and the two sessions before this one could not run it at all — 「最後の一句」 has one English translation and it is unobtainable. This work is the case the note has been waiting for: Lu Xun 「硬譯」與「文學的階級性」 has a standard English translation (Yang Xianyi and Gladys Yang, Selected Works of Lu Hsun vol. III, Foreign Languages Press, 1959/1980) and the scan is freely readable in full, so the gate runs.

1. The measurement, as run

shared 7-grams shared 12-grams shared 15-grams longest common run
13 1 0 12 tokens

The 13 shared 7-grams collapse to four distinct sites:

# run source tokens
L1 but there is still a great deal of paper in the world 然而世間紙張還多 12
L2 the colour of my teeth in a society like this 罵到牙齒的顏色。在這樣的社會裏 9
L3 the plays of Hauptmann and Lady Gregory in 印Hauptmann和Gregory夫人的劇本了 8
L4 to be of some use to society 我也願意於社會上有些用處 7

L3 is not evidence of anything: both names stand in the Roman alphabet in the Chinese source, so any translator transcribes them. It is recorded because deleting it after seeing it would be selecting the data.

The name-excluded column is unusable on this pair and is not reported. name_tokens()'s known false-positive class (R2: capitalised at every occurrence, never lowercase) resolves 80 tokens here, including a, in, the, this, they, world, yes, surely, naturally, today — the comparator is a 1,417-token slice in which many common words only ever open a sentence. The raw counts and the longest run are the signal, which is what the tool's own docstring says.

2. Where 12 tokens sits

CLAUDE.md records the measured range across five lead translations: 0 tokens (Ovid, no canonical English) to 21 tokens (Turgenev, against Garnett — canonical and unavoidable). 12 is squarely in the middle of that range, and the Yangs are as close to canonical for Lu Xun in English as Garnett is for Turgenev. That is enough to declare suspected and not enough to say why.

3. The control — is the run forced or recalled?

RS-20260726-ovid-period-form tested exactly this objection on Latin ("the Latin forces it") and it failed there: the three highest-overlap pairs were the three sharing 12-word verbatim runs, and the runs were not forced. The same objection deserves the same test here, on a new pair, and the test is cheap.

Question. Do independent translators, working from the Chinese alone, produce L1, L2 and L4?

Materials. The four Chinese sentences containing L1–L4, each with one clause of surrounding context, and nothing else. No English is shown.

Procedure. One call each to P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5, temperature 0, asked to render each sentence into English prose. The three models are independent of each other and of the lead. Raw responses under runs/.

Pre-registered scoring. A locus is FORCED if ≥2 of 3 independent outputs contain its run verbatim, or contain ≥7 of its tokens contiguously. Otherwise it is NOT FORCED.

Pre-registered predictions.

Pre-registered failure criterion, and it binds the artifact. If L1 comes back NOT FORCED — 0 or 1 of 3 — the 12-token run is evidence of recall and T-yingyi-jiejixing-R04-v1 is declared contamination: high, not suspected. If L1 is FORCED and L2 is NOT FORCED, the declaration stands at suspected with the reading that the shared material is the constrained material. A control that can only confirm the convenient answer is not a control, so the failing branch is written here, before dispatch, and the artifact is bound to it.

What this cannot settle. Three model outputs are not three human translators, and all three share whatever the Yangs' text contributed to their own training. The probe can therefore show a run is forced; it cannot show a run is not shared for some other reason. It is a one-sided test and is used only in that direction.

4. Result — the failure branch fired

3 calls, $0.0166679, all finish_reason: stop. Scored mechanically by runs/score.py against the frozen rule; the figure in each model column is the longest run of the locus's tokens appearing contiguously in that model's output.

locus tokens P1 P2 P3 hits verdict predicted
L1 12 5 5 5 0/3 NOT FORCED FORCED ✗ (P-a false)
L2 10 2 3 2 0/3 NOT FORCED NOT FORCED ✓ (P-c holds)
L3 8 4 7 7 2/3 FORCED (not predicted)
L4 7 7 7 7 3/3 FORCED FORCED ✓ (P-b holds)

(L2 is 10 tokens under dependence_check.py's tokenisation, not the 9 estimated in §1's table.)

The headline run is not forced. All three models rendered 世間紙張還多 as "there is still plenty of paper in the world". Not one produced a great deal of — the lead's phrase, and the Yangs'. Every model's longest contiguous match with the 12-token run is 5, well under the threshold of 7. Prediction P-a was wrong, and it was wrong in the direction that costs something: 「世間紙張還多」 has more freedom in English than the lead judged when writing the prediction, and the lead and the published translation took the same one of the available options.

L4 is the counter-case that makes the instrument readable. "to be of some use to society" came back verbatim from all three, so a 7-token shared run at that site is genuinely forced by 於社會上有些用處. The probe therefore separates the two loci rather than blessing or condemning them together, which is the only reason its verdict on L1 is worth anything. L3 is forced and uninteresting — both names stand in Roman script in the Chinese.

Applied to the artifact, per the frozen branch: T-yingyi-jiejixing-R04-v1 is declared contamination: high. The declaration is on a measurement and a control, not on a prior.

What this does and does not establish. It establishes that the longest shared run is not explained by the source constraining the English, which is the defence the run would otherwise have. It does not establish the mechanism: recall of the Yangs' sentence, a shared idiom-frequency prior, and the two translators independently making the same one of several free choices all predict this result. The probe is one-sided by design (§3) and is used only in the refusing direction.

Why this is worth more than the declaration it produced. The project's standing rule that the lead cannot serve as an independent third translator where a canonical English version exists has rested since S031 on the Turgenev/Garnett 21-token run — asserted to be unforced, never tested. RS-20260726-ovid-period-form tested the same defence on Latin and refused it. This is the second independent refusal, on a different pair, and the first on a lead translation. It also shows the test is cheap: three sentences, three calls, $0.017, and it can be run on any measured run the project has ever recorded. Note (bcp).

5. Provenance