Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: wiki/findings/results/RS-20260808c-sense-tradeoff-de.md · rendered 2026-09-09

Page metadata (front matter)
typeresult
idRS-20260808c-sense-tradeoff-de
statusfrozen
created2026-08-08
updated2026-08-08
trackT3
sensesaccuracy, naturalness, style-correspondence, perceived-source-carriage
internal-judgment-onlytrue
provisionaltrue
linksworkshop/experiments/E-20260808c-sense-tradeoff-de/design.md, wiki/arms/ARM-sense-tradeoff.md, wiki/findings/results/RS-20260807e-sense-tradeoff.md, wiki/goodness-senses.md, workshop/translations/kohlhaas-lisbeth/R06-v1/translation.md, workshop/translations/kohlhaas-lisbeth/R07-v1/translation.md, workshop/translations/kohlhaas-lisbeth/R08-v1/translation.md, workshop/regimes/R07-fluency.md, workshop/regimes/R08-resistancy.md, config/models.md

RS-20260808c — the accuracy null reproduces at zero in a second language pair, and the equivalence test it was given fails anyway

Tier D is NOT PASSED. Every sentence here is provisional and no verdict carries evidential weight (charter §2.4). What is claimed is a relation between senses inside one jury's scores, on one German passage, with the S129 Chinese passage alongside it.

E-20260808c-sense-tradeoff-de · ARM-sense-tradeoff step 2 · 16 bodies, 16 of 16 stop, zero discarded, zero re-dispatches, $0.347773240 against a declared $1.20 · key-usage reconciliation exact to 0.000000000 · verifier 100 checks, 0 failures, 5 of 5 mutations caught · pre-run critic NEEDS-REDESIGN, 5 BLOCKING, 5 ADVISORY, 8 accepted, 2 overruled with reasons.

1. What was done

Kleist's «Michael Kohlhaas» (1810), the Lisbeth span — her death, the funeral, the sovereign's decree, and the turn to revenge — 728 German words in seven segments, chosen for one declared property: its resistance to translation is syntactic, where S129's Lu Xun passage was lexical-stratum. Seven arms × seven segments = 49 items, three non-Anthropic seats (the Tier D jurors P1/P2/P5), blind to arm, four senses, plus a separate source-blind pass scoring naturalness on the English alone — dispatched first (amendment A10), so that the figure every naturalness claim uses was taken before any seat had seen the German.

588 of 588 source-visible cells and 147 of 147 blind cells returned. Every seat used 5 to 7 distinct integers of the seven on every sense.

arm accuracy naturalness (blind) style-corr. perceived-source-carriage
R06 no rule set (lead) 6.619 5.857 5.857 4.095
R07 domesticating (lead) 5.524 6.667 4.810 2.810
R08 foreignizing (lead) 5.524 1.429 5.286 6.571
WRONG (R06 + 1 content error/segment) 4.524 5.762 5.143 3.857
CLUNKY (R06 + 1 syntactic mangling/segment) 4.810 2.571 3.810 4.143
IND-R07 unbriefed hand, fluency rules only 6.333 6.333 5.524 3.571
IND-R08 unbriefed hand, resistancy rules only 6.333 5.524 5.619 4.048

2. The contamination gate was the first thing that ran, and it moved the design

Measured before the design was written, against two independent public-domain English hands of the identical span (Oxenford & Feiling 1844; Frances H. King 1913–14), extracted programmatically and never read by the translator.

pair 7-grams 12-grams longest run
KING ~ OXEN — two independent published hands 21 5 16
R06 ~ published 110 31 21
R07 ~ published 43 7 15
R08 ~ published 16 0 9
R07 ~ R08 4 0 9

The unruled arm reproduces published English five times more than two published hands reproduce each other, including a 21-token run — "the castellan the groom said had not been at home they had therefore been obliged to put up at an inn" — that the lead did not read. So R06 was demoted out of every primary before dispatch, exactly as the lead was demoted at S133, and the primary was set on R08 vs R07, which share nothing (4 shared 7-grams, longest run 9) and sit at or below the human–human baseline.

This is the fourth measurement against the premise that obscurity or non-canonicity protects, and the first that ranks three renderings of one span by regime: the more rules a translator follows, the less it reproduces the published record. R08 is the cleanest arm this project has measured.

3. The primary: the accuracy difference is exactly zero, and the test still fails

R1 — no trade. Δaccuracy(R08−R07) = 0.0000 on 7 segments, exact sign-flip P = 1.000; the (seat × segment) n = 21 form S129 used is likewise 0.0000, P = 1.000. Per seat: −0.143 / −1.429 / +1.571 — three readers disagreeing about a difference that is not there, the same shape as S129's 0.000 / −1.000 / +0.833.

And R1 FAILS as registered. Amendment A1, adopted from the critic's BLOCKING 1 before any number existed, replaced "P > 0.05" with a genuine equivalence criterion: the 90% permutation confidence interval must lie entirely inside ±0.75. It is [−0.777, +0.666]. The lower bound misses by 0.027 of a scale point.

This is the run's most useful single fact and it is a fact about the design, not about the senses. Seven segments do not carry enough precision to certify equivalence at ±0.75, which is what the critic said in advance and what the point estimate of exactly zero cannot rescue. A null this clean, tested honestly, still does not clear the bar — and had the original wording stood, it would have been reported as a pass.

What does clear the bar is the arm the lead did not write.

A5 — the independent co-primary. One unbriefed non-panel hand, given each rule set verbatim and nothing else — no hypotheses, no other arms, no knowledge that an experiment exists:

Δ (IND-R08 − IND-R07) 90% CI P
accuracy 0.0000 [−0.166, +0.166] 1.000
naturalness, blind −0.8095 [−1.222, −0.334] 0.0156
style-correspondence +0.0952 [−0.222, +0.444] 0.781
perceived-source-carriage +0.4762 [−0.500, +1.500] 0.422

On a pair the lead could not have shaped, the accuracy difference between the two opposed programmes is zero with a 90% interval of ±0.17 — equivalence at ±0.75 established with room to spare. This is the only figure in either run that establishes the accuracy null rather than failing to reject it, and it is lead-free.

4. What the expenditure bought, and what it did not

R2 HOLDS — Δnaturalness(R08−R07) blind = −5.2381, P = 0.0156 at segment level and 9.54 × 10⁻⁷ cellwise, negative at 3 of 3 seats (−5.286 / −5.286 / −5.143). But R2 is a manipulation check, not evidence (amendment A6, accepting the critic's BLOCKING 4): an arm written by rule to archaise, scored against an unmarked/literary-contemporary anchor, is close to defined as less natural. It establishes that the two programmes produced measurably different English, and nothing about the senses.

R3 HOLDS — Δperceived-source-carriage(R08−R07) = +3.7619, P = 0.0156, positive at 3 of 3 (+4.714 / +1.857 / +4.714). R08 scores 6.571 of 7, the highest cell in the run.

style-correspondence does NOT replicate: +0.4762, P = 0.344, against S129's +1.167. Recorded as a non-replication, not smoothed over.

So the shape holds and it is sharper here. The foreignizing programme spent 5.2 points of naturalness, bought 3.8 points of perceived source carriage, and returned exactly nothing on accuracy.

figure S129, Lu Xun, ZH→EN S134, Kleist, DE→EN
Δaccuracy(R08−R07), lead −0.056, P = 1.000 0.000, P = 1.000
Δnaturalness blind, lead −5.500 −5.238
Δperceived-source-carriage, lead +2.778, 3/3 +3.762, 3/3
Δstyle-correspondence, lead +1.167 +0.476, ns
Δaccuracy, independent pair +0.167 0.000, CI ±0.166
accuracy: R06 / R07 / R08 6.611 / 5.833 / 5.778 6.619 / 5.524 / 5.524

Two language pairs, two kinds of difficulty, two lead ladders and two unbriefed ladders. Four independent measurements of Δaccuracy between a foreignizing and a domesticating rendering of the same source: −0.056, +0.167, 0.000, 0.000. None exceeds a fifth of a scale point. On the same pairs Δnaturalness runs to −5.5.

5. The programme tax on accuracy replicates, and its confound was tested rather than conceded

S129 §5's surprise was that the unruled arm scored highest on accuracy. It reproduces, larger: Δaccuracy(R06−R07) = +1.0952, P = 0.0156; Δ(R06−R08) = +1.0952, P = 0.0625. Following either declared programme cost about 1.1 points of propositional accuracy against translating carefully under no programme at all — from opposite programmes, by the same amount, in both language pairs.

And here R06 is the contaminated arm, so the registered control C4 matters:

The honest reading, and it is not the comfortable one: the recall pathway is not excluded. ρ = +0.587 points the way the confound predicts, and n = 7 cannot resolve it. What can be said is that the effect survives on the segments where R06 shares almost nothing with either published hand, and that it is the same size in Chinese, where no such contamination was present. S1 is reported as replicated with an unexcluded confound named, not as clean.

6. Specifiability: the rules moved an unbriefed hand this time, and moved it on the wrong sense

S129 §6 found that one unbriefed hand given the two rule sets verbatim produced two texts three readers could not tell apart on any sense. S2's registered criterion — |Δ| < 0.75 on every sense — FAILS here, on exactly one: blind naturalness, −0.8095, P = 0.0156. Under the Bonferroni reading A11 requires (α = 0.0125) that single P does not survive correction, and the result is reported at both readings.

The sense it moved on is the point. The unbriefed hand's two texts differ by 0.81 in fluency and by 0.476 (P = 0.422) in perceived source carriage — while the same rules in the lead's hands moved perceived source carriage by 3.76. An unbriefed hand executing the resistancy rules recovered about an eighth of the foreignization the lead recovered, and recovered it as ordinary fluency loss rather than as source-oriented markedness.

That is a sharper answer to RS-20260728e-venuti-specifiability than S129's flat null: the ten rules are specifiable enough to change what a naive hand writes, and not specifiable enough to make it foreignize. The domesticating half is executable from its text; the resistancy half is not.

7. The gates, one of which failed

gate figure verdict
C1 jury sees content damage Δaccuracy(R06−WRONG) = +2.095, P = 0.0156, 3/3 PASS
C2 jury sees English damage Δnaturalness blind (R06−CLUNKY) = +3.286, P = 0.0156, 3/3 PASS
C3 cross-sense specificity WRONG leak ratio 0.045; CLUNKY leak ratio 0.551 CLUNKY FAILS (bar < 0.50)
F1 scale usage 5 to 7 distinct integers everywhere PASS

C3 fails on the syntactic control and the failure is informative, not fatal — C1/C2 are what withhold the primary and both passed. Mangled English syntax cost 1.81 points of accuracy against 3.29 of naturalness, although the CLUNKY operators changed no propositional content and the verifier confirms each is a pure string substitution of a declared clause. Readers do not score badly-built English as accurate, whatever the sense definitions say. This is a real limit on every accuracy figure in this project and it is recorded here rather than promoted to an arm (the subject rule).

The critic's BLOCKING 2 is not upheld, empirically. It held that C1/C2 inherit R06's contamination because both controls derive from it. Amendment A3 reported both gates split by contamination half: C1 is +1.417 on the high-overlap half against +3.000 on the low — the opposite of the predicted direction — and C2 is 3.333 against 3.222, indistinguishable. The within-pair argument the rebuild was overruled on holds up.

8. The compliance audit: how much of the contrast the rules actually own

The critic's BLOCKING 3 — the lead knew the S129 result while translating — is unfixable inside one session and is not claimed to be fixed. Amendment A4 bought the one available constraint: an independent non-panel agent, given only the two frozen rule sets, the German, and the two renderings, judged each of the translator's 42 claimed sites.

REQUIRED LICENSED NOT-SUPPORTED
R08 (21 sites) 16 3 2
R07 (21 sites) 14 6 1
total 30 (71%) 9 (21%) 3 (7%)

Seventy-one per cent of the two arms' divergence is, to an independent reader, forced by a public rule set frozen in 2026-07 by a different session for a different experiment. That is not independence and does not repair BLOCKING 3; it bounds how much of the contrast the lead's hand could have chosen. The three NOT-SUPPORTED verdicts are recorded against the arms: R08 site 17 (silvern) and site 21 (R10 claimed with no rendering attached), R07 site 6 (ohne Verschulden desselben judged grammatically unambiguous, so F8 does not apply).

9. Limits

  1. Two passages in two language pairs is two, not a generalisation. Per amendment A8 this run is a generalization test across difficulty type, not a replication; a failure would not have distinguished S129 was wrong from the relation is difficulty-specific, and the success does not establish the relation in German — only on this German passage, with these hands.
  2. The lead knew the S129 result while translating — worse than S129's "knew the hypotheses". Bounded by §8's 71%, by R08/R07 sharing nothing, and by A5's lead-free pair carrying the only established equivalence. Not removed.
  3. R1 fails as registered. The lead pair does not establish the accuracy null; it fails to reject a difference. Only A5 establishes it, and only at ±0.17 on one unbriefed hand.
  4. The S1 recall confound is unexcluded (§5, ρ = +0.587, P = 0.173).
  5. naturalness is scored against a register anchor R08 was written to violate — kept identical to S129 for comparability; the anchor change was overruled because wiki/goodness-senses.md permits three register points and none is period-indexed (D-20260803-15 ratified A).
  6. C3 fails on the syntactic control, so accuracy figures anywhere near mangled syntax are not clean.
  7. The lead read both published hands' English of the immediately preceding span while selecting material — common to all three lead arms, declared on each artifact.
  8. Three seats sharing a 2026 training distribution, and one unbriefed hand.
  9. Tier D NOT PASSED.

10. What this licenses, in one sentence each