Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260807e-sense-tradeoff/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260807e-sense-tradeoff
statusfrozen
created2026-08-07
updated2026-08-07
trackT3
linkswiki/arms/ARM-sense-tradeoff.md, wiki/goodness-senses.md, workshop/translations/kongyiji/R06-v1/translation.md, workshop/translations/kongyiji/R07-v1/translation.md, workshop/translations/kongyiji/R08-v1/translation.md, workshop/translations/kongyiji/contamination.md, framework/tierD-repaired-rules.md, config/models.md, wiki/base/sources/S-venuti-invisibility.md
sensesaccuracy, naturalness, style-correspondence, perceived-source-carriage
internal-judgment-onlytrue
provisionaltrue

E-20260807e — does buying accuracy cost naturalness?

Frozen before dispatch. No API call had been made when this page was committed. Amendments made after the critic pass are numbered A<n> at the foot and dated; nothing above them is edited in place.

1. Question

wiki/goodness-senses.md lists accuracy and naturalness as two senses of good and says nothing about how they relate. Every account in the project's Tier 2 base assumes they trade off — Schleiermacher's two paths, the belles infidèles, Venuti's domestication/foreignization. Is that assumption true of measured renderings of a literary passage?

2. Materials

Lu Xun 〈孔乙己〉 (1919), six segments, 1,201 Chinese characters — workshop/translations/kongyiji/. Selected for one declared property: the story runs on classical-Chinese material inside vernacular speech (之乎者也, 君子固窮, 多乎哉?不多也, 竊 against 偸, 秀才, the character 回), which is where a translator must choose between carrying what the source has and writing English that moves.

Five arms, 30 items (5 × 6 segments). Three are renderings; two are operator-derived controls.

arm what provenance
R08 resistancy — Venuti's ten foreignizing rules, frozen 2026-07-28, consulted at every choice lead, single pass, log frozen
R06 lead single pass, no rule set — the unmarked middle lead, single pass, log frozen
R07 fluency — Venuti's ten domesticating rules, frozen 2026-07-28 lead, single pass, log frozen
WRONG R06 + one declared content error per segment. English left fluent analysis/build_controls.py
CLUNKY R06 + one declared syntactic mangling per segment. Content left intact analysis/build_controls.py

The ladder is matched by construction: same translator, same session, same source, same reading, same revision depth (all three are single-pass R06-family regimes). The rule set is the only variable. The controls are byte-identical to R06 outside their declared operator, asserted by the build script (0 failures; every operator matches exactly once and restoring it reproduces R06).

Dose is LIGHT, and the reason is on the record. Tier D failed cross-sense specificity at the heavy dose (drop(naturalness) = +1.12 against ≤ 0.75, S086) and passed at three sites (+0.50). One operator per ~150-word segment is the light end. This is an imported constraint, not a rediscovered one.

Contamination: measured, none — workshop/translations/kongyiji/contamination.md. Lead R06 against the one reachable published rendering: 6 shared 7-grams, 0 twelve-grams, longest run 11 tokens, against a reference pair of two independent hands at 4 / 0 / 10.

3. Seats and blinding

Three seats, all non-Anthropic, the Tier D jurors of S020/S034/S086 so the instrument is extended rather than replaced (config/models.md panel v1): J1 = P1, J2 = P2, J3 = P5. Roles here, slugs logged as provenance in runs/.

4. The senses, with the strings quoted verbatim to the seats

accuracy, naturalness, style-correspondence, perceived-source-carriage, from wiki/goodness-senses.md, first paragraph of each entry.

5. Predictions, registered

Contrasts are paired by segment. Δ_s(X−Y) is the mean over the six segments of score(X, segment) − score(Y, segment) for seat s.

Why a halo cannot manufacture P1. A general "this is the better text" halo moves all senses together. P1 requires them to move apart, in opposite directions, on the same pair of texts. So the most likely instrument artefact in this project's experience pushes against the primary rather than towards it.

6. Controls and failure criteria — evaluated BEFORE the primary is read

No criterion above is weakened after it fires. If one fires, the run reports what it measured and the arm's step 2 decides whether to repair once or close retired.

7. Declared limits, before the run

  1. The arms are not independent renderings. Measured: R06~R07 share 51 seven-grams, 9 twelve-grams and a 16-token run; R06~R08 share 17 / 0 / 9. They are one rendering and two rule-driven departures from it. Shared wording can only shrink a contrast, never manufacture one, so every measured difference is conservative — but a null is therefore weak evidence of no effect, and will be reported as such.
  2. One passage, one language pair, one translator. Nothing here generalises beyond ZH→EN literary prose rendered by this lead.
  3. naturalness is defined on the target text alone and the seats can see the source. They are instructed to score it on the English alone; nothing enforces that. Any naturalness figure here is measured under source visibility and is not comparable to one measured without it.
  4. Tier D is NOT PASSED. No verdict here carries evidential weight; every sentence of the result is provisional. The claim the design reaches for is the weakest available: a relation between two senses inside one jury's scores, not a ranking of translations and not a statement that any arm is better.
  5. The lead wrote all five arms and registered P3 about its own log. P3 is therefore a prediction by the constructor of the data — the circularity RS-20260806d's critic named. It is mitigated only by the register being frozen and committed before the design existed, and it is reported as a weaker prediction than P1 for that reason.
  6. R08 is deliberately hard to read. If it scores low on naturalness that is the regime working, not a defect; the question is whether it buys anything on accuracy.

8. Pre-flight budget

Day 2026-08-07 stands at $2.431963030 of $5.00 before this session. Declared ceiling for this run: $0.65, worst case computed from the caps the requests actually permit (note (abc)), not from expected output.

stage calls cap worst case
critic 1 16,000 $0.16
scoring 3 seats × 3 blocks = 9 8,000 $0.34
re-dispatch allowance (F2, note (bhf)/(bhq)) ≤ 4 20,000 $0.15
total ≤ 14 $0.65

9. Procedure

  1. Freeze this page. Commit.
  2. One independent adversarial pre-run critic pass, non-panel slug, before any other call. Findings applied as numbered amendments; accepted and overruled findings both recorded with reasons.
  3. Build items and blocks (build_items.py, fixed seed, balance asserted).
  4. Dispatch scoring blocks. Every raw body written to runs/ before anything is computed.
  5. Parse, then run analyse.py, then verify.py, which re-derives every reported number by an independent code path and includes mutation tests.
  6. Evaluate C1–C3, F1–F3 first, then read P1–P4.

Amendments

Applied after the pre-run critic pass (nvidia/nemotron-3-ultra-550b-a55b, non-panel, provider BaseTen, $0.0276492, runs/critic.txt). Verdict NEEDS-REDESIGN, 4 BLOCKING, 6 ADVISORY. All ten findings accepted; two accepted in modified form with the reason written; one remedy overruled. Nothing above this line was edited in place.

A1 (critic BLOCKING 1, accepted in full) — naturalness gets a source-blind pass, and it is the primary measurement of that sense. The critic is right that showing the Chinese while asking for a target-text-only judgement does not merely weaken the measurement, it manufactures the predicted direction: a visible source confirms that R08's oddities are source-driven and that R07 is free, which is exactly the sign P1 predicts. Stage blind is added: every item is shown to every seat a second time, English only, no Chinese, no other sense, and naturalness is scored there. The source-visible naturalness figure is still collected and is reported beside it as a secondary, because the difference between the two is itself a measurement. P1 is evaluated on the source-blind figure.

A2 (critic BLOCKING 2, accepted in modified form; the critic's own remedy OVERRULED) — an unbriefed independent hand is added as a replication arm. The critic's remedy is that the lead must produce none of the renderings or the design be retracted. That is overruled: charter §3 and amendment A4 make lead translation a first-class labeled subject, and deleting it would delete the craft limb this arm exists to have. The substantive concern — that the translator knew the hypotheses and could construct the effect — is met instead by adding a hand that did not. Two new arms, IND-R08 and IND-R07: a non-panel model (so no seat scores its own prose) is given the six source segments and one rule set verbatim, and nothing else — not the hypotheses, not the other arms, not the fork log, not the word experiment. P1 is evaluated on BOTH pairs and both are reported. Where they disagree, the independent pair governs, and the lead pair is reported as what it is: a translator who knew what was being looked for.

A3 (critic BLOCKING 4, accepted in modified form) — P3 moves to the independent pair. The critic is right that P3 on the lead pair is the constructor predicting its own construction. P3 is therefore re-registered on IND-R08 vs IND-R07: the lead's frozen fork register must predict where a hand that never saw it diverges. On the lead pair the same correlation is computed and reported as exploratory, not registered. This is the shape RS-20260807c's critic forced at S127, where the lead-free version of the same prediction held at exact P = 0.01282 and the lead-inclusive one did not.

A4 (critic BLOCKING 3, accepted) — a cheap dose pilot gates the main dispatch. One seat scores the twelve control items and the six R06 items before the main run. If pooled Δ(accuracy, R06−WRONG) or Δ(naturalness, R06−CLUNKY) is below 0.50, a second operator per segment is added to that control arm and the arm is rebuilt before anything else is dispatched. The pilot seat's scores are discarded and not pooled into any reported figure.

A5 (critic ADVISORY 6, accepted) — C1/C2/C3 decoupled. The knife-edge is real: C1 used ≥ 0.75 as a floor and C3 used the same number as a ceiling, so a control that barely passes detection is likely to fail specificity. Registered now: C1/C2 detection ≥ 0.50, positive at 3 of 3 seats. C3 leakage: the off-target drop must be < 0.50 × the on-target drop, for both controls. The Tier D absolute number (off-target drop ≤ 0.75) is still computed and reported so the figure stays comparable with S086 and S113 — but the registered gate is the relative one, because a fixed absolute ceiling rewards a weak dose.

A6 (critic ADVISORY 8, accepted) — the test statistic is written down. T = mean over the 18 (seat × segment) units of [score(R08, seat, segment) − score(R07, seat, segment)]. The null is generated by flipping the sign of each of the 18 paired differences independently; all 2¹⁸ = 262,144 assignments are enumerated; the reported P is two-sided. The same statistic over the independent pair uses its own 18 units.

A7 (critic ADVISORY 7, accepted) — the scoring prompt stops saying something false. "The renderings come from different translators" is not true of WRONG and CLUNKY. Replaced with: "The renderings come from different sources: some are by different translators, some are modified versions of another rendering."

A8 (critic ADVISORY 10, accepted) — F3 counts eight deltas, not four: four senses × two controls, all < 0.25 ⇒ instrument null.

A9 (critic ADVISORY 9, accepted) — the conservativeness claim is withdrawn. §7 limit 1 said shared wording "can only shrink a contrast, never manufacture one". That assumes passive reuse and fails if the translator deliberately diverged at fork sites when switching regimes, which is precisely what a rule set instructs. Replaced with: the three lead arms share a translator and a session; shared wording may shrink or inflate a contrast depending on the translator's strategy, and no direction is assumed. This is why A2's independent pair, not the lead pair, carries the weight.

A10 (critic ADVISORY 5, accepted) — order is a named confound and is counterbalanced where it can be. The lead's order was R06 → R07 → R08 and is confounded with regime. The independent hand is dispatched in the opposite order, R08 before R07, so the two pairs do not share an order confound. With one hand per pair, order remains confounded within each pair and is carried as a limit.

A11 — budget ceiling raised from $0.65 to $1.10. A1, A2 and A4 add a blind pass (3 calls), two translation calls, a pilot call, and grow the scoring blocks from 30 items to 42. Day headroom after the critic call is $2.540. Worst case recomputed from the caps the requests permit: critic $0.028 spent · 2 translations (cap 4,000) $0.06 · pilot (cap 8,000) $0.05 · 9 scoring blocks (cap 10,000) $0.52 · 3 blind blocks (cap 12,000) $0.21 · re-dispatch allowance ≤ 4 (cap 24,000) $0.23 = $1.10.