Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260725-tierD-ladder/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260725-tierD-ladder
statusfrozen
created2026-07-25
updated2026-07-25
linksconfig/models.md, config/budget.md, wiki/program.md, wiki/goodness-senses.md, wiki/base/sources/S-berman-tendances.md, workshop/translations/berman-tendances/R04-v1/translation.md, wiki/base/anchors/A-shaw-spider-thread/A-shaw-spider-thread.md, wiki/base/anchors/A-yosano-yomogiu/A-yosano-yomogiu.md, workshop/translations/genji-yomogiu/R04-v1/translation.md, workshop/experiments/E-20260725-tierD-ladder/critic.md, workshop/experiments/README.md

Frozen design (v2) — Tier D detection calibration: a pilot, not a gate

v1 was rejected by the independent critic pass (critic.md, verdict NEEDS-REDESIGN, twenty-one blockers). This is the rewrite; 26 of 29 points are implemented and 3 are accepted as stated limitations, with dispositions tabulated in critic.md. A run may proceed only against this version.

No senses: field: this page designs an evaluation of the jury, not of a translation, and asserts no evaluative claim about any translation.

0. What this run can and cannot be — read first

This run cannot pass Tier D. Charter §5 and wiki/program.md Slate A item 4 make a held-out arm — reference against an independent same-quality translation, chance expected — a mandatory control. The critic established by measurement that the project's stored materials cannot supply one: the only candidate pair is Beowulf, where Gummere 1909 is alliterative verse (56 half-lines, 5 archaic tokens), Kirtlan 1913 is prose with 17 archaic tokens, and the lead's is plain modern prose with none. Verse against prose across a century of English is not a same-quality control; it can only produce an uninterpretable number.

v1 responded by quietly rewriting the control's success criterion inside the experiment page that benefited from it. That is exactly the move charter §2.8 forbids. v2 does not relax the control. It declares the control unsatisfiable from stored materials, drops the arm, and declines to claim a Tier D pass at all. config/models.md stays NOT CALIBRATED whatever these arms show, and finding a genuine same-quality pair becomes the named blocker on a real Tier D pass, recorded in NEXT.md.

What the run is: the first empirical characterisation of the instrument Tier D will need — the false-alarm rate against quality-neutral editing, the panel's actual use of a numeric scale, and detection and cross-sense specificity on four operators. Every finding is provisional.

1. Question

On which goodness senses can the non-Anthropic panel detect deliberate, sense-targeted degradation of a translation, and does the damaged sense fall further than the senses that were not damaged?

2. Materials

Two passages, two provenances, two source languages. All texts are public domain or lead-authored and already in the repository; nothing new is fetched and nothing new is stored.

id source reference translation provenance source language
A Akutagawa Ryūnosuke, 蜘蛛の糸 §一 (1918), 805 chars — workshop/canon/kumonoito/source.txt Glenn W. Shaw, "The Spider's Thread" §I (1930), 314 words — wiki/base/anchors/A-shaw-spider-thread/ published modern Japanese
B Murasaki Shikibu, 『源氏物語』「蓬生」§1‑3 (c. 1008), 494 chars — wiki/base/anchors/A-yosano-yomogiu/ the lead's own T-genji-yomogiu-R04-v1 §1‑3, 325 words lead classical Japanese

Word counts are measured, not estimated (critic A1 corrected v1's 318 and 380).

Why B carries the lead provenance. The design requires references of both provenances. Among the five available lead translations, the Genji §1‑3 passage is the only literary narrative prose into English whose source is stored beside it and whose source is not already committed to another role in this design. The Beowulf lead translation is excluded on a second ground the critic supplied: its own frozen log records its plainness as a deliberate choice to occupy one pole of that study's scale, which makes it an unrepresentative reference. (v1 gave a different and false reason — struck.)

What the jury is shown, and what was stripped. Each item is source + two unattributed English texts. Provenance headers naming Glenn W. Shaw, licence blocks, title lines, footnote markers and chapter headings are all removed by materials/build.py; the critic verified independently that none survives into the built files. Prose is unwrapped to one paragraph per line so that hard-wrap differences cannot fingerprint a text.

Not used, and why: the classical-Japanese caveat stands. B's source is harder for the jury to read than A's, so B's accuracy scores are expected to be noisier. Every metric is reported per passage as well as pooled.

3. Perturbation operators

Three of four come from Berman's tendances déformantes, read in French this session and frozen at T-berman-tendances-R04-v1 before this design existed (freeze commit 0aeeb93; the hash changed from 0111d38 when the branch was rebased to reset commit authorship — same tree, same ordering). The fourth does not, for a reason the frozen mapping establishes.

op catalogue source what is done target sense(s) prediction on naturalness
O1 Berman 5, l'appauvrissement qualitatif content words replaced by same-meaning words lacking the original's sonorous / iconic richness literary-quality flat
O2 Berman 1 + 3, rationalisation, allongement punctuation and clause order recomposed to canonical English discursive order; long periods split within their paragraph; verbs nominalised; connectives added style-correspondence flat or up
O3 Berman 10 + 11, destruction des réseaux vernaculaires / des locutions et idiotismes source-marking and period-marked idiom removed; every edit preserves the referent voice + cultural-mediation flat or up
O4 the project's own catalogue — the documented Shaw Christianisation (A-shaw-spider-thread) and Swann (1974) on Turney's lapses semantic errors: wrong referent, wrong word sense, invented detail, dropped negation. Style untouched. accuracy flat

O4 is not from Berman, and that is a finding. The frozen mapping records that none of the twelve tendencies targets accuracy — Berman's quarrel is with the translator who writes well. A Berman-only operator set cannot test the sense a jury is most likely to detect.

Four operators, not twelve. The frozen log records that Berman's twelve collapse to roughly five independent operations. Tendencies 8, 9 and 12 are excluded on a stated ground: a signifier network, a systematism and a superimposition of languages cannot be shown, let alone dosed, inside 320 words.

O3 was rebuilt after the critic pass. v1's O3 silently injected accuracy damage (an omission of "pistils and stamens", 八月 rendered "September", 総角 rendered "local kids", a stereoscope becoming a window) and crashed register ("Hold on now… it's got a soul too", which is Berman's vulgarisation, tendency 4, not 10/11). v2's O3 keeps every referent and stays inside the reference's own register; cultural-mediation is now scored, matching the frozen mapping.

Dose. One dose, 8 edit sites per item. v1's light dose is dropped: the light sets were nested subsets that retained the loudest edit, so they could not have measured a dose effect (critic D24), and the budget could not have afforded them anyway (E26). The dose axis is deferred explicitly, in NEXT.md, rather than by arithmetic accident.

Every edit site is logged verbatim in materials/perturbations.md, committed before the run, with the items.json SHA-256 recorded in the log and in every raw response file.

4. Arms and stages

10 items. The jury never learns that anything was modified, which text is the reference, or who produced either.

stage arm items purpose
1 sham A-sham, B-sham dispatched first, standalone. 8 substitutions per item, the same count as a targeted arm, chosen as free variation. Measures (a) the false-alarm rate against quality-neutral editing and (b) — the reason it runs first — how these three jurors actually use a 1–7 scale, which the project has never measured.
2 targeted A×{O1,O2,O3,O4}, B×{O1,O2,O3,O4} = 8 detection and cross-sense specificity

Order swap is mandatory on every item: each pair is dispatched twice, slots swapped. S010 measured slot preference at 0.63–0.85; an unswapped run is uninterpretable.

Stage 2 runs only if stage 1 clears two pre-registered conditions: 1. Scale usage. Pooled score standard deviation ≥ 0.75 and at least 4 distinct integers used across the six senses. If the jurors compress into two or three adjacent values, a 0.75-point specificity margin is not measurable and stage 2 is not dispatched. 2. Budget. Stage 1's actual billed cost leaves the stage-2 worst case (§8) inside the day's headroom.

On the sham's neutrality — conceded, not assumed. The critic's structural point is right: the reference is a considered text, so any eight substitutions tend to move off a local optimum, and a perfectly neutral sham is not constructible. v1's sham was worse than that — measurably archaizing (upon ×2, every one, about the grounds) and in one place sense-shifting (glanced ≠ looked). Those are removed. What remains is free variation (in bloom→blooming, which→that, until→till, remembered→recalled, stayed on→remained), word- and paragraph- and sentence-count neutral by measurement. The sham is therefore reported as an upper bound on the false-alarm rate, and its direction is a result rather than an assumption. It no longer holds a global veto (v1's [0.30, 0.70] band voided a perfectly neutral run ~15% of the time on 12 votes).

5. Jury and procedure

6. Pre-registered firing rule

The unit of analysis is (juror × passage), with the two orderings averaged within the unit — not the individual vote. v1 pooled 12 votes as if independent; the calibration-v1 critic raised that, it was accepted, S014 implemented it, and v1 reverted it (critic B7).

For each operator there are 6 units (3 jurors × 2 passages). A unit is scored:

Under a null of independent coin-flip preferences, P(+1) = 0.25.

Detection fires for operator O iff at least 5 of 6 units are +1 and none is −1. Exact one-sided P = 0.00464; across four operator cells the family-wise rate is ≈ 0.018. (v1's rule — 9 of 12 pooled votes — had P = 0.073 per cell and a 0.365 chance of at least one cell firing by accident.)

Specificity fires for O iff both: 1. the mean drop on O's target sense(s) exceeds the mean drop across the non-target senses excluding naturalness by ≥ 0.75 scale points; and 2. the largest per-sense drop, excluding naturalness, is one of O's targets.

naturalness is excluded from the baseline because the design predicts it will rise for O2 and O3, which would mechanically inflate the very margin it is meant to be independent of (critic B9). Both the excluding and the including figures are computed and reported; only the excluding one fires the rule. Condition 2 is scale-free and carries the actual content of "specificity"; condition 1 is calibrated against the scale usage stage 1 measures.

The naturalness probe has a rule and a consequence (v1 stated a prediction that could not fail — the third occurrence of that defect in this project): for O2 and O3, drop(naturalness) must be ≤ 0.25 scale points. If it exceeds that, the pre-registered reading is "the jury lowers every sense together", and no specificity claim is made for that operator regardless of its margin.

The sham's three-way interpretation rule, on its own 6 units: - ≥5 of 6 units +1 → the jury prefers unedited text as such; every detection result in stage 2 is confounded with edit-presence and no detection claim is licensed from this run. - ≤1 of 6 units +1 → the jury prefers edited text; same conclusion, opposite sign. - otherwise → the false-alarm rate is acceptable at this resolution, and stage 2's detection results stand on their own terms.

Gate status, pre-committed for every outcome (critic F34): Tier D remains NOT PASSED in all cases, because the held-out control is absent (§0). If detection fires and specificity fires on a sense, the recorded claim is "the jury detects S-damage at 8 sites, pending a held-out control". If detection fires and specificity fails, the claim is "detects damage, not sense-calibrated on S". If detection fails, the claim is "no detection at this dose and this power".

7. Predictions, written before the run

  1. accuracy (O4) fires detection. Most likely of the four.
  2. literary-quality (O1) shows the weakest detection — the edits preserve meaning and grammar.
  3. Specificity fails on at least two of the four operators: the jury says "worse" broadly. This is the modal outcome in the lead's estimate and is the thing the old Tier P design could never measure. internal-judgment-only.
  4. Detection is weaker on passage B than on A, because the jury reads classical Japanese less well.
  5. The sham lands in the middle band. The critic predicts the opposite — that the reference wins the sham arm outright, because any edit of a considered text drifts downward. Both predictions are on the record; the sham arm settles it.
  6. naturalness does not fall by more than 0.25 for O2 and O3.
  7. The pooled scale SD exceeds 0.75 and at least four distinct integers are used.

8. Budget

Headroom, UTC 2026‑07‑25: $1.600774 ($5.00 − $3.399226). Key-usage snapshot at session start: 18.290925792.

The estimate is stated as a range with a worst case, built from S014's measured per-call figures on a per-sense scoring task — means for the central estimate, maxima for the worst case (critic E27; the only run in this project to land inside its estimate is the one that did this).

stage items payloads calls central (means, $0.04146/payload) worst case (maxima, $0.06467/payload)
1 — sham 2 4 12 $0.166 $0.259
2 — targeted 8 16 48 $0.663 $1.035
total 10 20 60 $0.829 $1.293

Worst case $1.293 < headroom $1.601, with $0.31 (19%) of margin. Hard abort at $1.40, checked before every dispatch using the worst call seen so far as the margin. Because the sham is stage 1 and standalone, an abort can only ever truncate targeted arms — never the control that gates them (critic E28: v1's dispatch order would have dropped the sham first).

9. Failure criteria — what voids or qualifies the run

10. Verification

tools/verify_tierD.py recomputes every reported number from runs/ and the items manifest only, re-parsing each model's own text rather than trusting the runner's cached parse. It also checks the perturbation log by reconstruction — rebuilding each variant from its reference plus the logged edits and asserting byte equality, which catches edits that are present but unlogged. v1's check only tested the other direction and, the critic showed, missed exactly that: three edits in the built files were absent from the log at the time of review.

11. Known threats, stated in advance