Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: wiki/findings/results/RS-20260728-forced-run.md · rendered 2026-09-09

Page metadata (front matter)
typeresult
idRS-20260728-forced-run
statusactive
created2026-07-28
updated2026-07-30
senses—
internal-judgment-onlytrue
linksworkshop/translations/yingyi-jiejixing/contamination.md, workshop/translations/yingyi-jiejixing/R04-v1/translation.md, workshop/translations/yingyi-jiejixing/R06-v1/translation.md, wiki/findings/results/RS-20260726-ovid-period-form.md, wiki/method-notes.md, wiki/base/sources/S-luxun-yingyi.md

The forcing defence, tested on a lead translation for the first time, and refused

S044. E-20260728-forced-run. 3 calls, $0.0166679. Design and failure branch frozen in git at 750779e before dispatch; scored mechanically by runs/score.py.

1. What was asked

Every contamination figure this project has published is a longest common run between a lead translation and a published one, and every one of them has faced the same objection: the source forces it. The project has answered that objection twice by argument and never by measurement on a lead translation. RS-20260726-ovid-period-form refused it on Latin, between published pairs. CLAUDE.md's standing rule — the lead cannot serve as the independent third translator for any work whose standard English translation is canonical — rests on a 21-token run between the lead's Turgenev and Garnett's, which was declared unavoidable and never tested.

This session produced a case where the test could be run: T-yingyi-jiejixing-R06-v1, a single-pass draft of Lu Xun 1930 §五, frozen in git at a563d0b before any measurement, against the Yangs' 1959 English, freely readable in full.

shared 7-grams shared 12-grams shared 15-grams longest run
lead draft (1,533 tok) vs Yang & Yang §V (1,417 tok) 13, over four distinct sites 1 0 12 tokens

The question, made testable: give three independent panel models the four Chinese sentences, blind, and see whether they produce the runs. A locus counts as FORCED at ≥2 of 3 outputs containing the run verbatim or ≥7 of its tokens contiguously.

2. Result

locus English run Chinese tok P1 P2 P3 verdict
L1 but there is still a great deal of paper in the world 然而世間紙張還多 12 5 5 5 NOT FORCED 0/3
L2 the colour of my teeth in a society like this 罵到牙齒的顏色。在這樣的社會裏 10 2 3 2 NOT FORCED 0/3
L3 the plays of Hauptmann and Lady Gregory 印Hauptmann和Gregory夫人的劇本 8 4 7 7 FORCED 2/3
L4 to be of some use to society 我也願意於社會上有些用處 7 7 7 7 FORCED 3/3

All three models wrote plenty of paper. None wrote a great deal of. The lead did, and so did the Yangs. The headline run is not forced by the source, and the pre-registered failure branch binds: T-yingyi-jiejixing-R04-v1 is declared contamination: high — the project's first high declaration made on a measurement rather than on a prior.

L4 is why the verdict on L1 means anything. At a different site in the same passage, all three independent outputs reproduced the lead's seven tokens verbatim. So the instrument discriminates between loci in the same text rather than returning one answer for the whole passage, which is the only property that would let it refuse a defence rather than merely fail to support one. L3 is forced and uninteresting: both names stand in the Roman alphabet in the Chinese.

3. One prediction failed, and it is the more interesting half

The design predicted P-a: L1 is FORCED — 「世間紙張還多」 was judged to have almost no freedom in English. It was wrong. Three independent renderings took the same other option, and the option the lead took is the published translator's. Whatever explains that — recall, a shared frequency prior over a great deal of against plenty of, or two translators freely converging — the forcing explanation is gone, and the forcing explanation is the one the project has been relying on.

P-b (L4 forced) and P-c (L2 not forced) both held, so the instrument is not simply saying "nothing is forced".

4. What this licenses

Licensed. A shared run between a lead translation and a published one is not explained by source constraint unless the constraint has been demonstrated on somebody else, and demonstrating it costs three sentences, three calls and under two cents.

NARROWED 2026-07-28 (S045), RS-20260728b-forced-run-ru. The sentence above must read "…on somebody else who does not already hold the comparator". Limit 3 below was tested on Garnett and it is live and unrefuted: asked to recall the published English, all three models returned UNKNOWN at 24 of 24 cells — a refusal floor, not an absence — while the same three models, asked instead to identify the passages, named Turgenev for 8, 7 and 8 of 8 and the individual prose poem for 6, 4 and 5 of 8, from three lines each. Free recall does not work as a control on canonical material. What is owed is an elicitation with a floor at chance rather than at zero — forced choice against a distractor, where a model that knows nothing scores 50% and cannot retreat to UNKNOWN. The method's refusing direction is unaffected; what is not established is that a FORCED verdict on canonical material means source constraint.

Not licensed, and these are limits rather than hedges:

  1. The mechanism. Recall of the Yangs, a shared idiom prior, and independent free convergence all predict this result. The probe refuses a defence; it identifies nothing.
  2. Independence, ever. The test is one-sided by construction and is used only in the refusing direction. A FORCED verdict does not clear a run — it removes one reading of it.
  3. Three model outputs are not three translators, and all three share whatever the Yangs contributed to their training. This weakens the forced verdicts (L3, L4) more than the unforced ones: models that had memorised the Yangs would be more likely to reproduce the run, not less, so 0/3 on L1 is the harder result to explain away.
  4. n = 1 translation, 4 loci. Nothing here says how often the forcing defence fails in general.

5. What is now owed

Every published contamination figure in this project is retestable for under two cents each, and none has been tested. The five prior lead translations with measured runs — above all the 21-token Turgenev run that the standing rule in CLAUDE.md is built on — were each declared on the run alone. This session's result is that a run of comparable size on a comparable pair was not forced. That is not evidence the Turgenev run was unforced; it is the reason to go and look. Filed to wiki/backlog.md (T3).

DISCHARGED for the Turgenev cell, 2026-07-28 (S045) — ARM-forced-defence step 1, RS-20260728b-forced-run-ru. The 21-token run is a forced head with an elective tail: the models produce the front of it and none produces "a little and propping myself on my elbow". The standing rule keeps its support and now rests on a measurement. Across the eight loci, 6 of 8 forced on this page's own rule, 4 of 8 on a two-thirds rule — so the "11–21 tokens at eight of eight loci" framing overstated what the run lengths showed. The remaining measured figures, and a verdict on the five unmeasured declarations, are step 2.

FURTHER NARROWED 2026-07-30 (S065), RS-20260730f-recall-floor. The forced-choice control named just above was built. It does not work either: it passes its positive control at 18/18 on the King James Version, returns chance on Garnett (0.467, p = 0.856, 33% order-flip), and on a non-canonical comparator returns 46 of 48 judgments for the lead's own translation — a confident answer to a different question. Neither free recall nor forced choice measures whether a model holds a published translation, so the licensed sentence's clause "who does not already hold the comparator" remains a condition this project cannot currently verify. Method note (bez).

And a second-order caution, recorded because it nearly went unnoticed. Running the measurement between the frozen draft and the revision primes the translator: the lead then knows which of its own strings match a published version. The four matched sites were left exactly as drafted, and the rule was written into the log before the revision began — but the rule had to be invented on the spot, and it belongs in R04 rather than in one artifact's log. Note (bcp).