Repository path: workshop/experiments/E-20260808c-sense-tradeoff-de/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260808c-sense-tradeoff-de |
| status | frozen |
| created | 2026-08-08 |
| updated | 2026-08-08 |
| track | T3 |
| senses | accuracy, naturalness, style-correspondence, perceived-source-carriage |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-sense-tradeoff.md, wiki/findings/results/RS-20260807e-sense-tradeoff.md, wiki/goodness-senses.md, workshop/regimes/R06-lead-single-pass.md, workshop/regimes/R07-fluency.md, workshop/regimes/R08-resistancy.md, workshop/translations/kohlhaas-lisbeth/R06-v1/translation.md, workshop/translations/kohlhaas-lisbeth/R07-v1/translation.md, workshop/translations/kohlhaas-lisbeth/R08-v1/translation.md, config/models.md |
E-20260808c — does the sense relation measured on one Chinese passage hold in German?
ARM-sense-tradeoff step 2. Frozen before dispatch. Tier D is NOT PASSED; every figure this
design will produce is provisional and no verdict carries evidential weight (charter §2.4).
The subject-rule sentence, written before the design (continue-prompt.md §4.5):
This unit teaches whether the relation between
accuracyandnaturalnessmeasured on one Chinese passage — no trade-off, and a cost falling on other senses instead — holds in a second language pair whose resistance to translation is syntactic rather than lexical, and writes the answer into the project's typology of what "good" means.
1. Why step 2 is a run and not a paragraph
ARM-sense-tradeoff step 2 is scoped as write the measured relation into
wiki/goodness-senses.md. It cannot honestly be written from step 1 alone. RS-20260807e
measured one passage, one language pair, one lead translator; a standing relation between two
senses in the project's controlled vocabulary, cited by every future evaluation, cannot rest on
n = 1. Either the relation is written with a second measurement under it, or the arm closes with the
third of its three permitted sentences — the question is not answerable with this instrument. This
design is what decides which.
2. Materials
Heinrich von Kleist, «Michael Kohlhaas» (1810), the Lisbeth span — 728 German words in seven
segments, from «Diese Reise war aber von allen erfolglosen Schritten…» to «…wieder bei ihm in
Kohlhaasenbrück zu sein.» (workshop/translations/kohlhaas-lisbeth/source-de-segments.txt).
Why this span, declared as a property and not as a preference. Lu Xun's 孔乙己 was chosen at S129
because its resistance is lexical-stratum — classical Chinese inside vernacular speech. This
span's resistance is syntactic: Kleist suspends subject from verb across forty-word subordinate
chains, pre-modifies with stacked participles, and uses a da…so correlative English lost. If the
S129 relation is a fact about translation and not a fact about one kind of difficulty, it should
survive the change of locus. It is also a register mixture — narrative, a quoted scripture verse,
and a chancery decree — which is what R08's rule R8 and R07's rules F1/F10 disagree about.
Two other spans of this novella are already rendered in this repo (the Herse report, the Luther interview); this span overlaps neither.
The seven arms
| arm | what it is | hand |
|---|---|---|
R08 |
Venuti's foreignizing rules, executed | lead |
R06 |
no rule set, single pass | lead |
R07 |
Venuti's domesticating / fluency rules | lead |
WRONG |
R06 + one declared content error per segment |
derived |
CLUNKY |
R06 + one declared syntactic mangling per segment |
derived |
IND-R08 / IND-R07 |
a non-panel model given the two rule sets verbatim and nothing else | independent |
7 arms × 7 segments = 49 items. The lead's rendering order was R06 → R08 → R07,
crossed against S129's R06 → R07 → R08; the independent hand is dispatched R07 first,
crossed against S129's R08 first. The two runs therefore do not share an order confound
(RS-20260807e limit 4). Both control arms derive from R06 by the string operators in
materials/operators.json, each asserted to apply exactly once per segment.
3. The contamination gate, run before this design was written, and what it changed
tools/dependence_check.py over the three lead arms and two public-domain published hands of the
identical span — Oxenford & Feiling 1844 and Frances H. King 1913–14 — extracted programmatically
and never read by the translator (workshop/translations/kohlhaas-lisbeth/comparators/).
| pair | shared 7-grams | 12-grams | longest run |
|---|---|---|---|
| KING ~ OXEN (two independent published hands) | 21 | 5 | 16 |
R06 ~ published (both hands) |
110 | 31 | 21 |
R07 ~ published |
43 | 7 | 15 |
R08 ~ published |
16 | 0 | 9 |
R07 ~ R08 |
4 | 0 | 9 |
This is a gate, not a flag, and it moved the design before anything was dispatched.
R06is demoted out of the primary. The unruled arm shares five times the 7-grams and six times the 12-grams of two independent published hands, including a 21-token run ("the castellan the groom said had not been at home they had therefore been obliged to put up at an inn") that the translator did not read. Any accuracy advantage forR06is confounded with recall of published English. It stays in the run as the control base and as a declared- confounded secondary, and it supplies no primary figure.- The primary is
R08vsR07, and it is cleaner here than at S129. The two arms share 4 7-grams, 0 12-grams and a longest run of 9 — against S129's 16-token run between its two contrast arms.R08sits below the human–human baseline on every measure. - The confound is measured per segment, which buys the control in §5.
4. Procedure
Three seats — P1, P2, P5 (the Tier D jurors, as at S129, so this extends that instrument rather
than replacing it) — score all 49 items, blind to arm, in balanced shuffled blocks, on four senses,
1–7, with the source German present. A second, source-blind pass then scores naturalness alone
on the English with no German present; that is the figure every naturalness claim uses. Sense
definitions verbatim from wiki/goodness-senses.md; the naturalness register anchor is
unmarked / literary-contemporary, identical to S129.
A pre-run adversarial critic (non-panel) reads this design before any scoring call. Amendments are recorded here with their number before dispatch.
5. Registered predictions and gates
Gates, read first. If C1 or C2 fails the primary is withheld.
| gate | criterion |
|---|---|
C1 jury sees content damage |
Δaccuracy(R06−WRONG) > 0, positive at ≥ 2 of 3 seats |
C2 jury sees English damage |
Δnaturalness(R06−CLUNKY), blind, > 0, positive at ≥ 2 of 3 seats |
C3 cross-sense specificity |
leak ratio |Δ off-target| / |Δ on-target| < 0.50 for both controls (the S129 §8 form, registered here rather than argued afterwards) |
F1 scale usage |
every seat uses ≥ 4 distinct integers on every sense |
Primary — R1/R2/R3, the three halves of the S129 relation, each independently falsifiable.
R1— no trade. |Δaccuracy(R08−R07)| < 0.75 and exact permutation P > 0.05. Fails if the foreignizing arm buys accuracy (or loses it) by three quarters of a scale point at P ≤ 0.05 — which is what the folk trade-off predicts and what S129 did not find.R2— one-way expenditure. Δnaturalness(R08−R07), blind, < −1.00 at P < 0.05.R3— what the expenditure buys. Δperceived-source-carriage(R08−R07) > +1.00, positive at ≥ 2 of 3 seats.
The relation is reported as replicated only if all three hold. Any other combination is reported as what it is.
Secondary S1 — the programme tax on accuracy, with its confound declared in advance.
S129 §5 found the unruled R06 scoring above both ruled arms on accuracy. Registered:
Δaccuracy(R06−R07) > 0 and Δaccuracy(R06−R08) > 0.
Control
C4, registered before the run. The seven segments split at the median ofR06's published-overlap: high = {S1, S4, S5, S6} (25, 33, 19, 17 shared 7-grams), low = {S2, S3, S7} (12, 2, 2). If Δaccuracy(R06−R07) is positive on the high-overlap segments and ≤ 0 on the low-overlap segments, the advantage is attributed to recall andS1is NOT reported as a regime effect. If it is positive on both, recall does not explain it.
Secondary S2 — specifiability. S129 §6 found that one unbriefed hand given the two rule sets
verbatim produced two texts three readers could not tell apart. Registered as replicated if
|Δ| < 0.75 on every sense between IND-R08 and IND-R07. Fails if any sense separates them
by ≥ 0.75.
Not registered, and why. S129's P3 — the frozen translator's log predicting where readers see
renderings part — is not carried forward. It failed at S129 (ρ = −0.152, P = 0.408), its
successor here would collide seven registered sites onto seven segments, and it is a question about
the project's own instrument rather than about the translation (the subject rule). The R06 log's
registered forks stand as a record; nothing in this design scores them.
6. Failure criteria
C1orC2fails → the primary is withheld and the run reports the instrument failure.- Fewer than 90% of cells returned → the affected block is re-dispatched once; still short → the primary is withheld.
C4fires →S1is reported as unresolved-by-this-design, not as a null and not as an effect.- A seat naming the work or the translator is recorded; recognition is not a gate here, because no claim in this design turns on the seats not knowing what book it is.
7. Known limits, stated before the numbers exist
- One passage, one language pair, one lead hand, one independent model — as at S129. Two passages in two pairs is two, not a generalisation.
- The lead knew the S129 result while translating. This is worse than S129's "the lead knew
the hypotheses": the lead knew the answer. It is mitigated only by the contamination gate having
demoted the arm the lead would most plausibly have flattered, and by
R08/R07being written under frozen public rule sets that constrain most of the choices; it is not removed. naturalnessis measured against a register anchorR08was written to violate — S129 limit 7, kept identical here for comparability, and no more defensible than it was there.- The lead read both published hands' English of the immediately preceding span while selecting material. Common to all three lead arms; declared on each artifact.
- Three seats sharing a 2026 training distribution.
- Tier D NOT PASSED.
8. Pre-run critic — NEEDS-REDESIGN, 5 BLOCKING, 5 ADVISORY
nvidia/nemotron-3-ultra-550b-a55b (non-panel), $0.0281226, before any scoring call.
Ten findings, ten dispositions. Eight accepted, two overruled with the reason written. Nothing
below was decided after a number existed.
| # | sev | finding | disposition |
|---|---|---|---|
| 1 | BLOCKING | R1 accepts a null by failing to reject it. No equivalence margin, no power. |
ACCEPTED — A1, A2 |
| 2 | BLOCKING | WRONG/CLUNKY are string edits of the contaminated R06, so C1/C2 inherit it |
AMENDED — A3; the rebuild overruled |
| 3 | BLOCKING | The lead knew the S129 result while translating, chose the passage, wrote 3 of 7 arms and both operators | ACCEPTED as unfixable — A4, A5, A8 |
| 4 | BLOCKING | A contemporary naturalness anchor makes R2 near-tautological for a deliberately archaising arm |
ACCEPTED — A6; the anchor change overruled |
| 5 | BLOCKING | The scoring prompt announces that some items are modified | ACCEPTED — A7 |
| 6 | ADVISORY | This is a generalization test, not a replication | ACCEPTED — A8 |
| 7 | ADVISORY | C4's 4-vs-3 segment split has no power |
ACCEPTED — A9 |
| 8 | ADVISORY | The blind pass follows the source-visible pass on the same seats; carryover is certain | ACCEPTED — A10 |
| 9 | ADVISORY | R3 is a 3-seat sign test; S2 runs 4 tests uncorrected |
ACCEPTED — A11 |
| 10 | ADVISORY | The R06 log shows the lead closing, in the unruled arm, the very ambiguity R08's rule keeps open |
RECORDED — A12 |
The amendments, frozen before dispatch
A1—R1becomes an equivalence test.R1holds only if the 90% permutation confidence interval on Δaccuracy(R08−R07) lies entirely inside ±0.75 (two one-sided tests at α = 0.05). A wide interval straddling zero now failsR1instead of passing it. This is the amendment that could most easily cost the run its headline, and it was made before the numbers existed.A2— the unit is stated twice. The primary is computed at segment level (n = 7, seats averaged), which is the true experimental unit; the (seat × segment) n = 21 figure S129 used is reported alongside for comparability, and labelled as the repeated-measures version it is.A3—C1/C2are reported split by contamination. The rebuild from a clean base is overruled:C1andC2are within-pair contrasts (R06againstR06-plus-an-operator), so a recall advantage present in both members subtracts out. The critic's mechanism requires recall to interact with damage detection, which is not argued and is testable — so both gates are also reported on the high-overlap {S1,S4,S5,S6} and low-overlap {S2,S3,S7} halves. If the two halves disagree in sign, finding 2 is upheld and the gates are reported as compromised.A4— a rule-compliance audit. An independent non-panel agent, given only the two frozen rule sets and the two lead renderings, judges for each of the 41 logged sites whether the cited rule actually requires that rendering. This converts the lead chose into the rule chose, verifiably, for whatever fraction it confirms. It does not address passage choice and does not repair finding 3.A5— the independent pair is a registered co-primary forR1. The lead's knowledge cannot operate onIND-R07/IND-R08. Both outcomes are registered: if the independent pair also shows |Δaccuracy| < 0.75 while Δnaturalnessmoves, that is lead-free support; if the independent pair is flat on every sense, that replicates S129 §6 instead andR1gets no independent support. The lead pair alone cannot establishR1.A6—R2is demoted from primary to manipulation check. The critic is right that an arm written by rule to archaise, scored against a contemporary anchor, is close to defined as less natural.R2therefore verifies only that the two programmes produced measurably different English; it is not evidence about the senses. The anchor change is overruled:wiki/goodness-senses.mdpermits an evaluation to name exactly three register points and none of them is period-indexed, and inventing a fourth here would both break comparability with S129 and install an untested anchor by the back door (D-20260803-15ratifiedA, no change).A7— the prompt no longer says any item is modified. The sentence "some are modified versions of another rendering" is struck; the items are presented only as renderings from different sources.A8— the run is relabelled. This is a generalization test across difficulty type, not a replication. No sentence in the result may say "the relation holds in German"; the most it may say is "on this German passage, with this lead hand". A failure will not distinguish S129 was wrong from the relation is difficulty-specific, and the result must say so.A9—C4gains its continuous form and loses its authority. Spearman ρ between each segment'sR06-vs-published 7-gram count and that segment's Δaccuracy(R06−R07), n = 7, reported with its exact permutation P. The 4-vs-3 split is reported as description with its power stated as absent.A10— the blind pass is dispatched FIRST. Reversing the order removes the carryover the critic names in the only direction that matters: the blindnaturalnessfigure is now taken before any seat has seen the German. Free, and it repairs the sole figure everynaturalnessclaim uses.A11—R3gets the same exact permutation test asR1/R2rather than a 2-of-3 sign test; seat agreement is reported as description.S2's four senses are reported with a Bonferroni reading (α = 0.0125) stated alongside the raw figures.A12— recorded, not repaired.R06log fork 17 closes «kraft der ihm angebotenen Macht» whereR08site 19 keeps it open by rule R4. The critic reads this as the lead shaping the diagnostic site. Note the direction: closing an ambiguity is the choice that would raiseR06's accuracy, which is the secondary the contamination gate already demoted — so it cuts againstS1, not for it.
9. Budget
Pre-flight worst case $1.20, built from max_tokens and not from expected output (note (abc)):
critic 16,000; 9 scoring bodies at 10,000; 3 blind bodies at 12,000; 2 translate bodies at 4,000.
UTC day 2026-08-08 stands at $1.584951693 of $5.00 before this run.