Repository path: workshop/experiments/E-20260807e-sense-tradeoff/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260807e-sense-tradeoff |
| status | frozen |
| created | 2026-08-07 |
| updated | 2026-08-07 |
| track | T3 |
| links | wiki/arms/ARM-sense-tradeoff.md, wiki/goodness-senses.md, workshop/translations/kongyiji/R06-v1/translation.md, workshop/translations/kongyiji/R07-v1/translation.md, workshop/translations/kongyiji/R08-v1/translation.md, workshop/translations/kongyiji/contamination.md, framework/tierD-repaired-rules.md, config/models.md, wiki/base/sources/S-venuti-invisibility.md |
| senses | accuracy, naturalness, style-correspondence, perceived-source-carriage |
| internal-judgment-only | true |
| provisional | true |
E-20260807e — does buying accuracy cost naturalness?
Frozen before dispatch. No API call had been made when this page was committed. Amendments made
after the critic pass are numbered A<n> at the foot and dated; nothing above them is edited in
place.
1. Question
wiki/goodness-senses.md lists accuracy and naturalness as two senses of good and says
nothing about how they relate. Every account in the project's Tier 2 base assumes they trade off —
Schleiermacher's two paths, the belles infidèles, Venuti's domestication/foreignization. Is
that assumption true of measured renderings of a literary passage?
2. Materials
Lu Xun 〈孔乙己〉 (1919), six segments, 1,201 Chinese characters — workshop/translations/kongyiji/.
Selected for one declared property: the story runs on classical-Chinese material inside vernacular
speech (之乎者也, 君子固窮, 多乎哉?不多也, 竊 against 偸, 秀才, the character 回), which is where a
translator must choose between carrying what the source has and writing English that moves.
Five arms, 30 items (5 × 6 segments). Three are renderings; two are operator-derived controls.
| arm | what | provenance |
|---|---|---|
R08 |
resistancy — Venuti's ten foreignizing rules, frozen 2026-07-28, consulted at every choice | lead, single pass, log frozen |
R06 |
lead single pass, no rule set — the unmarked middle | lead, single pass, log frozen |
R07 |
fluency — Venuti's ten domesticating rules, frozen 2026-07-28 | lead, single pass, log frozen |
WRONG |
R06 + one declared content error per segment. English left fluent |
analysis/build_controls.py |
CLUNKY |
R06 + one declared syntactic mangling per segment. Content left intact |
analysis/build_controls.py |
The ladder is matched by construction: same translator, same session, same source, same reading,
same revision depth (all three are single-pass R06-family regimes). The rule set is the only
variable. The controls are byte-identical to R06 outside their declared operator, asserted by
the build script (0 failures; every operator matches exactly once and restoring it reproduces R06).
Dose is LIGHT, and the reason is on the record. Tier D failed cross-sense specificity at the
heavy dose (drop(naturalness) = +1.12 against ≤ 0.75, S086) and passed at three sites (+0.50). One
operator per ~150-word segment is the light end. This is an imported constraint, not a rediscovered
one.
Contamination: measured, none — workshop/translations/kongyiji/contamination.md. Lead R06
against the one reachable published rendering: 6 shared 7-grams, 0 twelve-grams, longest run 11
tokens, against a reference pair of two independent hands at 4 / 0 / 10.
3. Seats and blinding
Three seats, all non-Anthropic, the Tier D jurors of S020/S034/S086 so the instrument is extended
rather than replaced (config/models.md panel v1): J1 = P1, J2 = P2, J3 = P5. Roles here, slugs
logged as provenance in runs/.
- Each seat scores all 30 items on four senses, integer 1–7.
- Each item is shown with its Chinese source segment. Required for
accuracy,style-correspondenceandperceived-source-carriage; a declared confound fornaturalness(§7). - No arm labels, no regime names, no mention that arms or programmes exist. Seats are told the set contains renderings by different translators, in no order.
- Items are shuffled per seat under a fixed seed and dispatched in three balanced blocks of ten: each block carries exactly two items from each arm, and at most two from any one segment. No block places two arms of the same segment adjacent.
4. The senses, with the strings quoted verbatim to the seats
accuracy, naturalness, style-correspondence, perceived-source-carriage, from
wiki/goodness-senses.md, first paragraph of each entry.
naturalness: the revised, target-text-only string is used. Recorded here becauseS113requires it: both strings are in the repository and give 1.083 and 0.750 on the same twelve units.- Register anchor named, as the entry requires:
unmarked/literary-contemporary. perceived-source-carriage: the seats HAVE the source. The entry requires this to be stated, and carries a measured limit that is not removed — manufactured oddity was attributed to the source at 0.40 with the source present, so a high score here is weak evidence of actual source licence.
5. Predictions, registered
Contrasts are paired by segment. Δ_s(X−Y) is the mean over the six segments of
score(X, segment) − score(Y, segment) for seat s.
P1— THE TRADE-OFF (primary).Δ(accuracy, R08−R07) > 0andΔ(naturalness, R08−R07) < 0, at ≥ 2 of 3 seats, and with both pooled contrasts in those directions, and at least one of the two pooled contrasts clearing an exact paired sign-flip permutation test at P ≤ 0.05 over the 18 (segment × seat) paired units — all 2¹⁸ = 262,144 sign assignments enumerated, no sampling.P2— the middle sits between.R06lies betweenR08andR07on both senses, pooled. Descriptive; does not gate.P3— the translator knew where it bites. Spearman ρ between the frozen segment rank in theR06log's fork register (1 = densest FORK) and the segment's divergence magnitude|Δ(accuracy, R08−R07)| + |Δ(naturalness, R08−R07)|, pooled over seats, is negative, at exact permutation P ≤ 0.05 over all 720 orderings. The register was frozen and committed before this design existed.P4— the rule sets did what they say.Δ(perceived-source-carriage, R08−R07) > 0at 3 of 3 seats. IfP4fails, the two rule sets did not produce a foreignising/domesticating difference a reader can see, andP1is reported as a null about these two texts and not about the poles.
Why a halo cannot manufacture P1. A general "this is the better text" halo moves all senses
together. P1 requires them to move apart, in opposite directions, on the same pair of
texts. So the most likely instrument artefact in this project's experience pushes against the
primary rather than towards it.
6. Controls and failure criteria — evaluated BEFORE the primary is read
C1— can the jury see content damage? PooledΔ(accuracy, R06−WRONG) ≥ 0.75and positive at 3 of 3 seats. IfC1fails,P1is WITHHELD.C2— can the jury see English damage? PooledΔ(naturalness, R06−CLUNKY) ≥ 0.75and positive at 3 of 3 seats. IfC2fails,P1is WITHHELD.C3— cross-sense specificity, the Tier D condition imported verbatim at its own threshold.Δ(naturalness, R06−WRONG) ≤ 0.75andΔ(accuracy, R06−CLUNKY) ≤ 0.75. If either is exceeded, the two senses are not separable on this instrument at this dose, andP1is WITHHELD with the numbers reported. This is the criterion Tier D failed at S086; it may fail again here, and that outcome is registered as a first-class result, not as a reason to move the bar.F1— scale usage. If a seat uses fewer than 3 distinct integers across its 30 scores on a sense, that (seat, sense) cell is reported and excluded from the pooled figure, with the exclusion declared. It is not silently kept.F2— completeness. Any missing or unparseable cell → the block is re-dispatched once. If a cell is still missing, the segment is dropped from every arm and the drop is reported. No figure is computed over an unbalanced set.F3— dose visibility. IfWRONGandCLUNKYare not distinguishable fromR06on any sense at any seat (all four pooled deltas < 0.25), the operators did not land and the whole run is reported as an instrument null.
No criterion above is weakened after it fires. If one fires, the run reports what it measured
and the arm's step 2 decides whether to repair once or close retired.
7. Declared limits, before the run
- The arms are not independent renderings. Measured:
R06~R07share 51 seven-grams, 9 twelve-grams and a 16-token run;R06~R08share 17 / 0 / 9. They are one rendering and two rule-driven departures from it. Shared wording can only shrink a contrast, never manufacture one, so every measured difference is conservative — but a null is therefore weak evidence of no effect, and will be reported as such. - One passage, one language pair, one translator. Nothing here generalises beyond ZH→EN literary prose rendered by this lead.
naturalnessis defined on the target text alone and the seats can see the source. They are instructed to score it on the English alone; nothing enforces that. Anynaturalnessfigure here is measured under source visibility and is not comparable to one measured without it.- Tier D is NOT PASSED. No verdict here carries evidential weight; every sentence of the result
is
provisional. The claim the design reaches for is the weakest available: a relation between two senses inside one jury's scores, not a ranking of translations and not a statement that any arm is better. - The lead wrote all five arms and registered
P3about its own log.P3is therefore a prediction by the constructor of the data — the circularityRS-20260806d's critic named. It is mitigated only by the register being frozen and committed before the design existed, and it is reported as a weaker prediction thanP1for that reason. R08is deliberately hard to read. If it scores low onnaturalnessthat is the regime working, not a defect; the question is whether it buys anything onaccuracy.
8. Pre-flight budget
Day 2026-08-07 stands at $2.431963030 of $5.00 before this session. Declared ceiling for this run: $0.65, worst case computed from the caps the requests actually permit (note (abc)), not from expected output.
| stage | calls | cap | worst case |
|---|---|---|---|
| critic | 1 | 16,000 | $0.16 |
| scoring | 3 seats × 3 blocks = 9 | 8,000 | $0.34 |
re-dispatch allowance (F2, note (bhf)/(bhq)) |
≤ 4 | 20,000 | $0.15 |
| total | ≤ 14 | $0.65 |
9. Procedure
- Freeze this page. Commit.
- One independent adversarial pre-run critic pass, non-panel slug, before any other call. Findings applied as numbered amendments; accepted and overruled findings both recorded with reasons.
- Build items and blocks (
build_items.py, fixed seed, balance asserted). - Dispatch scoring blocks. Every raw body written to
runs/before anything is computed. - Parse, then run
analyse.py, thenverify.py, which re-derives every reported number by an independent code path and includes mutation tests. - Evaluate
C1–C3,F1–F3first, then readP1–P4.
Amendments
Applied after the pre-run critic pass (nvidia/nemotron-3-ultra-550b-a55b, non-panel, provider
BaseTen, $0.0276492, runs/critic.txt). Verdict NEEDS-REDESIGN, 4 BLOCKING, 6 ADVISORY.
All ten findings accepted; two accepted in modified form with the reason written; one remedy
overruled. Nothing above this line was edited in place.
A1 (critic BLOCKING 1, accepted in full) — naturalness gets a source-blind pass, and it is the
primary measurement of that sense. The critic is right that showing the Chinese while asking for
a target-text-only judgement does not merely weaken the measurement, it manufactures the predicted
direction: a visible source confirms that R08's oddities are source-driven and that R07 is
free, which is exactly the sign P1 predicts. Stage blind is added: every item is shown to
every seat a second time, English only, no Chinese, no other sense, and naturalness is scored
there. The source-visible naturalness figure is still collected and is reported beside it as
a secondary, because the difference between the two is itself a measurement. P1 is evaluated on
the source-blind figure.
A2 (critic BLOCKING 2, accepted in modified form; the critic's own remedy OVERRULED) — an
unbriefed independent hand is added as a replication arm. The critic's remedy is that the lead
must produce none of the renderings or the design be retracted. That is overruled: charter §3 and
amendment A4 make lead translation a first-class labeled subject, and deleting it would delete the
craft limb this arm exists to have. The substantive concern — that the translator knew the
hypotheses and could construct the effect — is met instead by adding a hand that did not. Two new
arms, IND-R08 and IND-R07: a non-panel model (so no seat scores its own prose) is given the
six source segments and one rule set verbatim, and nothing else — not the hypotheses, not the
other arms, not the fork log, not the word experiment. P1 is evaluated on BOTH pairs and both
are reported. Where they disagree, the independent pair governs, and the lead pair is reported as
what it is: a translator who knew what was being looked for.
A3 (critic BLOCKING 4, accepted in modified form) — P3 moves to the independent pair. The
critic is right that P3 on the lead pair is the constructor predicting its own construction.
P3 is therefore re-registered on IND-R08 vs IND-R07: the lead's frozen fork register must
predict where a hand that never saw it diverges. On the lead pair the same correlation is computed
and reported as exploratory, not registered. This is the shape RS-20260807c's critic forced at
S127, where the lead-free version of the same prediction held at exact P = 0.01282 and the
lead-inclusive one did not.
A4 (critic BLOCKING 3, accepted) — a cheap dose pilot gates the main dispatch. One seat scores
the twelve control items and the six R06 items before the main run. If pooled
Δ(accuracy, R06−WRONG) or Δ(naturalness, R06−CLUNKY) is below 0.50, a second operator per
segment is added to that control arm and the arm is rebuilt before anything else is dispatched. The
pilot seat's scores are discarded and not pooled into any reported figure.
A5 (critic ADVISORY 6, accepted) — C1/C2/C3 decoupled. The knife-edge is real: C1 used
≥ 0.75 as a floor and C3 used the same number as a ceiling, so a control that barely passes
detection is likely to fail specificity. Registered now:
C1/C2 detection ≥ 0.50, positive at 3 of 3 seats.
C3 leakage: the off-target drop must be < 0.50 × the on-target drop, for both controls.
The Tier D absolute number (off-target drop ≤ 0.75) is still computed and reported so the
figure stays comparable with S086 and S113 — but the registered gate is the relative one, because
a fixed absolute ceiling rewards a weak dose.
A6 (critic ADVISORY 8, accepted) — the test statistic is written down.
T = mean over the 18 (seat × segment) units of [score(R08, seat, segment) − score(R07, seat,
segment)]. The null is generated by flipping the sign of each of the 18 paired differences
independently; all 2¹⁸ = 262,144 assignments are enumerated; the reported P is two-sided. The same
statistic over the independent pair uses its own 18 units.
A7 (critic ADVISORY 7, accepted) — the scoring prompt stops saying something false. "The
renderings come from different translators" is not true of WRONG and CLUNKY. Replaced with:
"The renderings come from different sources: some are by different translators, some are modified
versions of another rendering."
A8 (critic ADVISORY 10, accepted) — F3 counts eight deltas, not four: four senses × two
controls, all < 0.25 ⇒ instrument null.
A9 (critic ADVISORY 9, accepted) — the conservativeness claim is withdrawn. §7 limit 1 said shared wording "can only shrink a contrast, never manufacture one". That assumes passive reuse and fails if the translator deliberately diverged at fork sites when switching regimes, which is precisely what a rule set instructs. Replaced with: the three lead arms share a translator and a session; shared wording may shrink or inflate a contrast depending on the translator's strategy, and no direction is assumed. This is why A2's independent pair, not the lead pair, carries the weight.
A10 (critic ADVISORY 5, accepted) — order is a named confound and is counterbalanced where it can
be. The lead's order was R06 → R07 → R08 and is confounded with regime. The independent
hand is dispatched in the opposite order, R08 before R07, so the two pairs do not share an
order confound. With one hand per pair, order remains confounded within each pair and is carried
as a limit.
A11 — budget ceiling raised from $0.65 to $1.10. A1, A2 and A4 add a blind pass (3 calls), two translation calls, a pilot call, and grow the scoring blocks from 30 items to 42. Day headroom after the critic call is $2.540. Worst case recomputed from the caps the requests permit: critic $0.028 spent · 2 translations (cap 4,000) $0.06 · pilot (cap 8,000) $0.05 · 9 scoring blocks (cap 10,000) $0.52 · 3 blind blocks (cap 12,000) $0.21 · re-dispatch allowance ≤ 4 (cap 24,000) $0.23 = $1.10.