Repository path: wiki/findings/results/RS-20260808c-sense-tradeoff-de.md · rendered 2026-09-09
Page metadata (front matter)
| type | result |
|---|---|
| id | RS-20260808c-sense-tradeoff-de |
| status | frozen |
| created | 2026-08-08 |
| updated | 2026-08-08 |
| track | T3 |
| senses | accuracy, naturalness, style-correspondence, perceived-source-carriage |
| internal-judgment-only | true |
| provisional | true |
| links | workshop/experiments/E-20260808c-sense-tradeoff-de/design.md, wiki/arms/ARM-sense-tradeoff.md, wiki/findings/results/RS-20260807e-sense-tradeoff.md, wiki/goodness-senses.md, workshop/translations/kohlhaas-lisbeth/R06-v1/translation.md, workshop/translations/kohlhaas-lisbeth/R07-v1/translation.md, workshop/translations/kohlhaas-lisbeth/R08-v1/translation.md, workshop/regimes/R07-fluency.md, workshop/regimes/R08-resistancy.md, config/models.md |
RS-20260808c — the accuracy null reproduces at zero in a second language pair, and the equivalence test it was given fails anyway
Tier D is NOT PASSED. Every sentence here is provisional and no verdict carries evidential
weight (charter §2.4). What is claimed is a relation between senses inside one jury's scores,
on one German passage, with the S129 Chinese passage alongside it.
E-20260808c-sense-tradeoff-de · ARM-sense-tradeoff step 2 · 16 bodies, 16 of 16 stop, zero
discarded, zero re-dispatches, $0.347773240 against a declared $1.20 · key-usage reconciliation
exact to 0.000000000 · verifier 100 checks, 0 failures, 5 of 5 mutations caught · pre-run
critic NEEDS-REDESIGN, 5 BLOCKING, 5 ADVISORY, 8 accepted, 2 overruled with reasons.
1. What was done
Kleist's «Michael Kohlhaas» (1810), the Lisbeth span — her death, the funeral, the sovereign's
decree, and the turn to revenge — 728 German words in seven segments, chosen for one declared
property: its resistance to translation is syntactic, where S129's Lu Xun passage was
lexical-stratum. Seven arms × seven segments = 49 items, three non-Anthropic seats (the Tier
D jurors P1/P2/P5), blind to arm, four senses, plus a separate source-blind pass scoring
naturalness on the English alone — dispatched first (amendment A10), so that the figure every
naturalness claim uses was taken before any seat had seen the German.
588 of 588 source-visible cells and 147 of 147 blind cells returned. Every seat used 5 to 7 distinct integers of the seven on every sense.
| arm | accuracy |
naturalness (blind) |
style-corr. |
perceived-source-carriage |
|---|---|---|---|---|
R06 no rule set (lead) |
6.619 | 5.857 | 5.857 | 4.095 |
R07 domesticating (lead) |
5.524 | 6.667 | 4.810 | 2.810 |
R08 foreignizing (lead) |
5.524 | 1.429 | 5.286 | 6.571 |
WRONG (R06 + 1 content error/segment) |
4.524 | 5.762 | 5.143 | 3.857 |
CLUNKY (R06 + 1 syntactic mangling/segment) |
4.810 | 2.571 | 3.810 | 4.143 |
IND-R07 unbriefed hand, fluency rules only |
6.333 | 6.333 | 5.524 | 3.571 |
IND-R08 unbriefed hand, resistancy rules only |
6.333 | 5.524 | 5.619 | 4.048 |
2. The contamination gate was the first thing that ran, and it moved the design
Measured before the design was written, against two independent public-domain English hands of the identical span (Oxenford & Feiling 1844; Frances H. King 1913–14), extracted programmatically and never read by the translator.
| pair | 7-grams | 12-grams | longest run |
|---|---|---|---|
| KING ~ OXEN — two independent published hands | 21 | 5 | 16 |
R06 ~ published |
110 | 31 | 21 |
R07 ~ published |
43 | 7 | 15 |
R08 ~ published |
16 | 0 | 9 |
R07 ~ R08 |
4 | 0 | 9 |
The unruled arm reproduces published English five times more than two published hands reproduce
each other, including a 21-token run — "the castellan the groom said had not been at home they had
therefore been obliged to put up at an inn" — that the lead did not read. So R06 was demoted out
of every primary before dispatch, exactly as the lead was demoted at S133, and the primary was set
on R08 vs R07, which share nothing (4 shared 7-grams, longest run 9) and sit at or below the
human–human baseline.
This is the fourth measurement against the premise that obscurity or non-canonicity protects, and
the first that ranks three renderings of one span by regime: the more rules a translator follows,
the less it reproduces the published record. R08 is the cleanest arm this project has measured.
3. The primary: the accuracy difference is exactly zero, and the test still fails
R1 — no trade. Δaccuracy(R08−R07) = 0.0000 on 7 segments, exact sign-flip
P = 1.000; the (seat × segment) n = 21 form S129 used is likewise 0.0000, P = 1.000. Per
seat: −0.143 / −1.429 / +1.571 — three readers disagreeing about a difference that is not there,
the same shape as S129's 0.000 / −1.000 / +0.833.
And R1 FAILS as registered. Amendment A1, adopted from the critic's BLOCKING 1 before any
number existed, replaced "P > 0.05" with a genuine equivalence criterion: the 90% permutation
confidence interval must lie entirely inside ±0.75. It is [−0.777, +0.666]. The lower bound
misses by 0.027 of a scale point.
This is the run's most useful single fact and it is a fact about the design, not about the senses. Seven segments do not carry enough precision to certify equivalence at ±0.75, which is what the critic said in advance and what the point estimate of exactly zero cannot rescue. A null this clean, tested honestly, still does not clear the bar — and had the original wording stood, it would have been reported as a pass.
What does clear the bar is the arm the lead did not write.
A5 — the independent co-primary. One unbriefed non-panel hand, given each rule set verbatim and
nothing else — no hypotheses, no other arms, no knowledge that an experiment exists:
Δ (IND-R08 − IND-R07) |
90% CI | P | |
|---|---|---|---|
accuracy |
0.0000 | [−0.166, +0.166] | 1.000 |
naturalness, blind |
−0.8095 | [−1.222, −0.334] | 0.0156 |
style-correspondence |
+0.0952 | [−0.222, +0.444] | 0.781 |
perceived-source-carriage |
+0.4762 | [−0.500, +1.500] | 0.422 |
On a pair the lead could not have shaped, the accuracy difference between the two opposed programmes is zero with a 90% interval of ±0.17 — equivalence at ±0.75 established with room to spare. This is the only figure in either run that establishes the accuracy null rather than failing to reject it, and it is lead-free.
4. What the expenditure bought, and what it did not
R2 HOLDS — Δnaturalness(R08−R07) blind = −5.2381, P = 0.0156 at segment level and
9.54 × 10⁻⁷ cellwise, negative at 3 of 3 seats (−5.286 / −5.286 / −5.143). But R2 is a
manipulation check, not evidence (amendment A6, accepting the critic's BLOCKING 4): an arm
written by rule to archaise, scored against an unmarked/literary-contemporary anchor, is close to
defined as less natural. It establishes that the two programmes produced measurably different
English, and nothing about the senses.
R3 HOLDS — Δperceived-source-carriage(R08−R07) = +3.7619, P = 0.0156, positive at 3
of 3 (+4.714 / +1.857 / +4.714). R08 scores 6.571 of 7, the highest cell in the run.
style-correspondence does NOT replicate: +0.4762, P = 0.344, against S129's +1.167. Recorded as
a non-replication, not smoothed over.
So the shape holds and it is sharper here. The foreignizing programme spent 5.2 points of naturalness, bought 3.8 points of perceived source carriage, and returned exactly nothing on accuracy.
| figure | S129, Lu Xun, ZH→EN | S134, Kleist, DE→EN |
|---|---|---|
Δaccuracy(R08−R07), lead |
−0.056, P = 1.000 | 0.000, P = 1.000 |
Δnaturalness blind, lead |
−5.500 | −5.238 |
Δperceived-source-carriage, lead |
+2.778, 3/3 | +3.762, 3/3 |
Δstyle-correspondence, lead |
+1.167 | +0.476, ns |
Δaccuracy, independent pair |
+0.167 | 0.000, CI ±0.166 |
accuracy: R06 / R07 / R08 |
6.611 / 5.833 / 5.778 | 6.619 / 5.524 / 5.524 |
Two language pairs, two kinds of difficulty, two lead ladders and two unbriefed ladders. Four
independent measurements of Δaccuracy between a foreignizing and a domesticating rendering of the
same source: −0.056, +0.167, 0.000, 0.000. None exceeds a fifth of a scale point. On the same pairs
Δnaturalness runs to −5.5.
5. The programme tax on accuracy replicates, and its confound was tested rather than conceded
S129 §5's surprise was that the unruled arm scored highest on accuracy. It reproduces, larger:
Δaccuracy(R06−R07) = +1.0952, P = 0.0156; Δ(R06−R08) = +1.0952, P = 0.0625.
Following either declared programme cost about 1.1 points of propositional accuracy against
translating carefully under no programme at all — from opposite programmes, by the same amount, in
both language pairs.
And here R06 is the contaminated arm, so the registered control C4 matters:
- High-overlap segments {S1,S4,S5,S6}: +1.417. Low-overlap {S2,S3,S7}: +0.667. Both positive, so
C4does not fire — recall does not explain the advantage away. - Against
R08the gradient reverses: high +1.000, low +1.222. - The continuous form (
A9): Spearman ρ between a segment's published overlap and its Δaccuracy(R06−R07) = +0.587, exact permutation P = 0.173 over all 5,040 orderings.
The honest reading, and it is not the comfortable one: the recall pathway is not excluded.
ρ = +0.587 points the way the confound predicts, and n = 7 cannot resolve it. What can be said is
that the effect survives on the segments where R06 shares almost nothing with either published
hand, and that it is the same size in Chinese, where no such contamination was present. S1 is
reported as replicated with an unexcluded confound named, not as clean.
6. Specifiability: the rules moved an unbriefed hand this time, and moved it on the wrong sense
S129 §6 found that one unbriefed hand given the two rule sets verbatim produced two texts three
readers could not tell apart on any sense. S2's registered criterion — |Δ| < 0.75 on every
sense — FAILS here, on exactly one: blind naturalness, −0.8095, P = 0.0156. Under the
Bonferroni reading A11 requires (α = 0.0125) that single P does not survive correction, and the
result is reported at both readings.
The sense it moved on is the point. The unbriefed hand's two texts differ by 0.81 in fluency and by 0.476 (P = 0.422) in perceived source carriage — while the same rules in the lead's hands moved perceived source carriage by 3.76. An unbriefed hand executing the resistancy rules recovered about an eighth of the foreignization the lead recovered, and recovered it as ordinary fluency loss rather than as source-oriented markedness.
That is a sharper answer to RS-20260728e-venuti-specifiability than S129's flat null: the ten
rules are specifiable enough to change what a naive hand writes, and not specifiable enough to make
it foreignize. The domesticating half is executable from its text; the resistancy half is not.
7. The gates, one of which failed
| gate | figure | verdict |
|---|---|---|
C1 jury sees content damage |
Δaccuracy(R06−WRONG) = +2.095, P = 0.0156, 3/3 |
PASS |
C2 jury sees English damage |
Δnaturalness blind (R06−CLUNKY) = +3.286, P = 0.0156, 3/3 |
PASS |
C3 cross-sense specificity |
WRONG leak ratio 0.045; CLUNKY leak ratio 0.551 |
CLUNKY FAILS (bar < 0.50) |
F1 scale usage |
5 to 7 distinct integers everywhere | PASS |
C3 fails on the syntactic control and the failure is informative, not fatal — C1/C2 are
what withhold the primary and both passed. Mangled English syntax cost 1.81 points of accuracy
against 3.29 of naturalness, although the CLUNKY operators changed no propositional content and
the verifier confirms each is a pure string substitution of a declared clause. Readers do not score
badly-built English as accurate, whatever the sense definitions say. This is a real limit on every
accuracy figure in this project and it is recorded here rather than promoted to an arm (the subject
rule).
The critic's BLOCKING 2 is not upheld, empirically. It held that C1/C2 inherit R06's
contamination because both controls derive from it. Amendment A3 reported both gates split by
contamination half: C1 is +1.417 on the high-overlap half against +3.000 on the low — the
opposite of the predicted direction — and C2 is 3.333 against 3.222, indistinguishable. The
within-pair argument the rebuild was overruled on holds up.
8. The compliance audit: how much of the contrast the rules actually own
The critic's BLOCKING 3 — the lead knew the S129 result while translating — is unfixable inside one
session and is not claimed to be fixed. Amendment A4 bought the one available constraint: an
independent non-panel agent, given only the two frozen rule sets, the German, and the two renderings,
judged each of the translator's 42 claimed sites.
| REQUIRED | LICENSED | NOT-SUPPORTED | |
|---|---|---|---|
R08 (21 sites) |
16 | 3 | 2 |
R07 (21 sites) |
14 | 6 | 1 |
| total | 30 (71%) | 9 (21%) | 3 (7%) |
Seventy-one per cent of the two arms' divergence is, to an independent reader, forced by a public
rule set frozen in 2026-07 by a different session for a different experiment. That is not
independence and does not repair BLOCKING 3; it bounds how much of the contrast the lead's hand could
have chosen. The three NOT-SUPPORTED verdicts are recorded against the arms: R08 site 17
(silvern) and site 21 (R10 claimed with no rendering attached), R07 site 6 (ohne Verschulden
desselben judged grammatically unambiguous, so F8 does not apply).
9. Limits
- Two passages in two language pairs is two, not a generalisation. Per amendment
A8this run is a generalization test across difficulty type, not a replication; a failure would not have distinguished S129 was wrong from the relation is difficulty-specific, and the success does not establish the relation in German — only on this German passage, with these hands. - The lead knew the S129 result while translating — worse than S129's "knew the hypotheses".
Bounded by §8's 71%, by
R08/R07sharing nothing, and byA5's lead-free pair carrying the only established equivalence. Not removed. R1fails as registered. The lead pair does not establish the accuracy null; it fails to reject a difference. OnlyA5establishes it, and only at ±0.17 on one unbriefed hand.- The
S1recall confound is unexcluded (§5, ρ = +0.587, P = 0.173). naturalnessis scored against a register anchorR08was written to violate — kept identical to S129 for comparability; the anchor change was overruled becausewiki/goodness-senses.mdpermits three register points and none is period-indexed (D-20260803-15ratifiedA).C3fails on the syntactic control, soaccuracyfigures anywhere near mangled syntax are not clean.- The lead read both published hands' English of the immediately preceding span while selecting material — common to all three lead arms, declared on each artifact.
- Three seats sharing a 2026 training distribution, and one unbriefed hand.
- Tier D NOT PASSED.
10. What this licenses, in one sentence each
- On the senses: under the two Venuti programmes, in two language pairs and four ladders,
accuracydid not move (|Δ| ≤ 0.17 throughout) whilenaturalnessmoved by up to 5.5 andperceived-source-carriageby up to 3.8 — so the axis the folk trade-off names is not theaccuracyaxis. - On what it does not license: nothing about whether the two senses are independent in general.
What was manipulated was a declared programme, and only that manipulation is shown not to move
accuracy. - On regimes: following either declared programme cost about 1.1 points of
accuracyagainst translating carefully under none, in both language pairs — with a recall confound named and unexcluded in the German half. - On specifiability: the fluency rules are executable by an unbriefed hand; the resistancy rules, handed over verbatim, produced about an eighth of the lead's foreignization and produced it as fluency loss rather than as source carriage.
- On the instrument: an equivalence test on seven segments cannot certify a null at ±0.75 even
when the point estimate is exactly zero, and a syntactic-damage control leaks 55% into
accuracy. - On nothing else. No arm is called better than any other; no jury verdict carries weight; Tier D remains NOT PASSED.