Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260808c-sense-tradeoff-de/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260808c-sense-tradeoff-de
statusfrozen
created2026-08-08
updated2026-08-08
trackT3
sensesaccuracy, naturalness, style-correspondence, perceived-source-carriage
internal-judgment-onlytrue
provisionaltrue
linkswiki/arms/ARM-sense-tradeoff.md, wiki/findings/results/RS-20260807e-sense-tradeoff.md, wiki/goodness-senses.md, workshop/regimes/R06-lead-single-pass.md, workshop/regimes/R07-fluency.md, workshop/regimes/R08-resistancy.md, workshop/translations/kohlhaas-lisbeth/R06-v1/translation.md, workshop/translations/kohlhaas-lisbeth/R07-v1/translation.md, workshop/translations/kohlhaas-lisbeth/R08-v1/translation.md, config/models.md

E-20260808c — does the sense relation measured on one Chinese passage hold in German?

ARM-sense-tradeoff step 2. Frozen before dispatch. Tier D is NOT PASSED; every figure this design will produce is provisional and no verdict carries evidential weight (charter §2.4).

The subject-rule sentence, written before the design (continue-prompt.md §4.5):

This unit teaches whether the relation between accuracy and naturalness measured on one Chinese passage — no trade-off, and a cost falling on other senses instead — holds in a second language pair whose resistance to translation is syntactic rather than lexical, and writes the answer into the project's typology of what "good" means.

1. Why step 2 is a run and not a paragraph

ARM-sense-tradeoff step 2 is scoped as write the measured relation into wiki/goodness-senses.md. It cannot honestly be written from step 1 alone. RS-20260807e measured one passage, one language pair, one lead translator; a standing relation between two senses in the project's controlled vocabulary, cited by every future evaluation, cannot rest on n = 1. Either the relation is written with a second measurement under it, or the arm closes with the third of its three permitted sentences — the question is not answerable with this instrument. This design is what decides which.

2. Materials

Heinrich von Kleist, «Michael Kohlhaas» (1810), the Lisbeth span — 728 German words in seven segments, from «Diese Reise war aber von allen erfolglosen Schritten…» to «…wieder bei ihm in Kohlhaasenbrück zu sein.» (workshop/translations/kohlhaas-lisbeth/source-de-segments.txt).

Why this span, declared as a property and not as a preference. Lu Xun's 孔乙己 was chosen at S129 because its resistance is lexical-stratum — classical Chinese inside vernacular speech. This span's resistance is syntactic: Kleist suspends subject from verb across forty-word subordinate chains, pre-modifies with stacked participles, and uses a da…so correlative English lost. If the S129 relation is a fact about translation and not a fact about one kind of difficulty, it should survive the change of locus. It is also a register mixture — narrative, a quoted scripture verse, and a chancery decree — which is what R08's rule R8 and R07's rules F1/F10 disagree about.

Two other spans of this novella are already rendered in this repo (the Herse report, the Luther interview); this span overlaps neither.

The seven arms

arm what it is hand
R08 Venuti's foreignizing rules, executed lead
R06 no rule set, single pass lead
R07 Venuti's domesticating / fluency rules lead
WRONG R06 + one declared content error per segment derived
CLUNKY R06 + one declared syntactic mangling per segment derived
IND-R08 / IND-R07 a non-panel model given the two rule sets verbatim and nothing else independent

7 arms × 7 segments = 49 items. The lead's rendering order was R06 → R08 → R07, crossed against S129's R06 → R07 → R08; the independent hand is dispatched R07 first, crossed against S129's R08 first. The two runs therefore do not share an order confound (RS-20260807e limit 4). Both control arms derive from R06 by the string operators in materials/operators.json, each asserted to apply exactly once per segment.

3. The contamination gate, run before this design was written, and what it changed

tools/dependence_check.py over the three lead arms and two public-domain published hands of the identical span — Oxenford & Feiling 1844 and Frances H. King 1913–14 — extracted programmatically and never read by the translator (workshop/translations/kohlhaas-lisbeth/comparators/).

pair shared 7-grams 12-grams longest run
KING ~ OXEN (two independent published hands) 21 5 16
R06 ~ published (both hands) 110 31 21
R07 ~ published 43 7 15
R08 ~ published 16 0 9
R07 ~ R08 4 0 9

This is a gate, not a flag, and it moved the design before anything was dispatched.

  1. R06 is demoted out of the primary. The unruled arm shares five times the 7-grams and six times the 12-grams of two independent published hands, including a 21-token run ("the castellan the groom said had not been at home they had therefore been obliged to put up at an inn") that the translator did not read. Any accuracy advantage for R06 is confounded with recall of published English. It stays in the run as the control base and as a declared- confounded secondary, and it supplies no primary figure.
  2. The primary is R08 vs R07, and it is cleaner here than at S129. The two arms share 4 7-grams, 0 12-grams and a longest run of 9 — against S129's 16-token run between its two contrast arms. R08 sits below the human–human baseline on every measure.
  3. The confound is measured per segment, which buys the control in §5.

4. Procedure

Three seats — P1, P2, P5 (the Tier D jurors, as at S129, so this extends that instrument rather than replacing it) — score all 49 items, blind to arm, in balanced shuffled blocks, on four senses, 1–7, with the source German present. A second, source-blind pass then scores naturalness alone on the English with no German present; that is the figure every naturalness claim uses. Sense definitions verbatim from wiki/goodness-senses.md; the naturalness register anchor is unmarked / literary-contemporary, identical to S129.

A pre-run adversarial critic (non-panel) reads this design before any scoring call. Amendments are recorded here with their number before dispatch.

5. Registered predictions and gates

Gates, read first. If C1 or C2 fails the primary is withheld.

gate criterion
C1 jury sees content damage Δaccuracy(R06−WRONG) > 0, positive at ≥ 2 of 3 seats
C2 jury sees English damage Δnaturalness(R06−CLUNKY), blind, > 0, positive at ≥ 2 of 3 seats
C3 cross-sense specificity leak ratio |Δ off-target| / |Δ on-target| < 0.50 for both controls (the S129 §8 form, registered here rather than argued afterwards)
F1 scale usage every seat uses ≥ 4 distinct integers on every sense

Primary — R1/R2/R3, the three halves of the S129 relation, each independently falsifiable.

The relation is reported as replicated only if all three hold. Any other combination is reported as what it is.

Secondary S1 — the programme tax on accuracy, with its confound declared in advance. S129 §5 found the unruled R06 scoring above both ruled arms on accuracy. Registered: Δaccuracy(R06−R07) > 0 and Δaccuracy(R06−R08) > 0.

Control C4, registered before the run. The seven segments split at the median of R06's published-overlap: high = {S1, S4, S5, S6} (25, 33, 19, 17 shared 7-grams), low = {S2, S3, S7} (12, 2, 2). If Δaccuracy(R06−R07) is positive on the high-overlap segments and ≤ 0 on the low-overlap segments, the advantage is attributed to recall and S1 is NOT reported as a regime effect. If it is positive on both, recall does not explain it.

Secondary S2 — specifiability. S129 §6 found that one unbriefed hand given the two rule sets verbatim produced two texts three readers could not tell apart. Registered as replicated if |Δ| < 0.75 on every sense between IND-R08 and IND-R07. Fails if any sense separates them by ≥ 0.75.

Not registered, and why. S129's P3 — the frozen translator's log predicting where readers see renderings part — is not carried forward. It failed at S129 (ρ = −0.152, P = 0.408), its successor here would collide seven registered sites onto seven segments, and it is a question about the project's own instrument rather than about the translation (the subject rule). The R06 log's registered forks stand as a record; nothing in this design scores them.

6. Failure criteria

7. Known limits, stated before the numbers exist

  1. One passage, one language pair, one lead hand, one independent model — as at S129. Two passages in two pairs is two, not a generalisation.
  2. The lead knew the S129 result while translating. This is worse than S129's "the lead knew the hypotheses": the lead knew the answer. It is mitigated only by the contamination gate having demoted the arm the lead would most plausibly have flattered, and by R08/R07 being written under frozen public rule sets that constrain most of the choices; it is not removed.
  3. naturalness is measured against a register anchor R08 was written to violate — S129 limit 7, kept identical here for comparability, and no more defensible than it was there.
  4. The lead read both published hands' English of the immediately preceding span while selecting material. Common to all three lead arms; declared on each artifact.
  5. Three seats sharing a 2026 training distribution.
  6. Tier D NOT PASSED.

8. Pre-run critic — NEEDS-REDESIGN, 5 BLOCKING, 5 ADVISORY

nvidia/nemotron-3-ultra-550b-a55b (non-panel), $0.0281226, before any scoring call. Ten findings, ten dispositions. Eight accepted, two overruled with the reason written. Nothing below was decided after a number existed.

# sev finding disposition
1 BLOCKING R1 accepts a null by failing to reject it. No equivalence margin, no power. ACCEPTED — A1, A2
2 BLOCKING WRONG/CLUNKY are string edits of the contaminated R06, so C1/C2 inherit it AMENDED — A3; the rebuild overruled
3 BLOCKING The lead knew the S129 result while translating, chose the passage, wrote 3 of 7 arms and both operators ACCEPTED as unfixable — A4, A5, A8
4 BLOCKING A contemporary naturalness anchor makes R2 near-tautological for a deliberately archaising arm ACCEPTED — A6; the anchor change overruled
5 BLOCKING The scoring prompt announces that some items are modified ACCEPTED — A7
6 ADVISORY This is a generalization test, not a replication ACCEPTED — A8
7 ADVISORY C4's 4-vs-3 segment split has no power ACCEPTED — A9
8 ADVISORY The blind pass follows the source-visible pass on the same seats; carryover is certain ACCEPTED — A10
9 ADVISORY R3 is a 3-seat sign test; S2 runs 4 tests uncorrected ACCEPTED — A11
10 ADVISORY The R06 log shows the lead closing, in the unruled arm, the very ambiguity R08's rule keeps open RECORDED — A12

The amendments, frozen before dispatch

9. Budget

Pre-flight worst case $1.20, built from max_tokens and not from expected output (note (abc)): critic 16,000; 9 scoring bodies at 10,000; 3 blind bodies at 12,000; 2 translate bodies at 4,000. UTC day 2026-08-08 stands at $1.584951693 of $5.00 before this run.