Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260805e-naturalness-wording/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260805e-naturalness-wording
statusfrozen
created2026-08-05
updated2026-08-05
linkswiki/arms/ARM-tierP.md, wiki/decisions/resolved/D-20260802-13-naturalness-licence-clause.md, wiki/decisions/votes/2026-08-02/D-20260802-13-ratification-record.md, wiki/findings/results/RS-20260802-tierD-verdict.md, workshop/experiments/E-20260801f-tierD-run/design.md, framework/tierD-repaired-rules.md, wiki/goodness-senses.md, config/models.md, config/budget.md, workshop/translations/o-enfermeiro/R04-v1/translation.md, workshop/translations/o-enfermeiro/R06-v1/translation.md, workshop/experiments/E-20260805e-naturalness-wording/materials/perturbations.md, workshop/experiments/E-20260805e-naturalness-wording/critic.md

Frozen design (v2) — does the failing number belong to the wording or to the prose?

ARM-tierP step 3, the arm's declared closing condition, and the discharge of wiki/backlog.md's only owed row — condition 4 of D-20260802-13's ratifying vote: re-run Tier D's failed drop(naturalness) specificity test under the revised wording before any revised-sense verdict is called calibrated.

v1 was frozen, then put through the independent pre-run critic pass (critic.md, seat x-ai/grok-4.5, verdict NEEDS-REDESIGN, sixteen findings, eight BLOCKING). This is the amended version: all sixteen accepted, fourteen in full and two in part with the refused remedy reasoned. A run may proceed only against this version, and nothing had been dispatched when the amendments were made.

No senses: field: this page designs an evaluation of the instrument, not of a translation, and asserts no evaluative claim about any translation. The lead's two new references are scored by the panel, blind and unattributed; the lead judges nothing (charter §5).

0. What this run can and cannot be — read first

  1. Tier D is NOT PASSED and nothing here changes that. This run re-measures one cell of one condition. It does not dispatch the sham, the held-out arm or the prior positive control, and a specificity condition is one of three.
  2. It is a re-measurement, not a new instrument. Payloads, jurors, orderings, unit definition, scoring and the six senses are carried from E-20260801f; the OLD arm's payloads are byte-identical to that run's stage 4.
  3. The manipulated variable is the whole naturalness criterion string, and it changes in three places at once. §3.2 lists them. No sentence this run produces may attribute an outcome to the struck escape clause, to the target-only frame, or to the named register anchor separately. Everything is attributed to the §3.2 REVISED string as a whole (three inseparable edits).
  4. So the obligation is discharged as a BUNDLE TEST and not otherwise. Condition 4 of D-20260802-13's vote asks for the test under the revised wording; that is what runs here. A clause-specific discharge — whether the escape clause in particular drives the number — is out of scope and needs the single-edit design recorded in §10. This is the first BLOCKING finding of the critic pass and it is accepted without hedging.
  5. A null is the more interesting outcome and is registered as such (§6, P4p). If the number does not move, the S034/S086 figure is a fact about how O4-damaged prose reads, not about how the question was worded — and that is a claim about translation, not about the apparatus.

1. The subject-rule sentence, written before the unit was designed

What does this unit teach about translating literature or evaluating translations? — It tests whether a reader's sense that an English translation has gone wrong as English can be held apart from the accuracy errors that made it go wrong, and whether that inseparability is a property of the damaged prose or of the words the question was asked in. Two prior runs found that eight accuracy errors in a 300–440-word passage stop reading as eight accuracy errors and start reading as prose that has failed.

The unit is nonetheless method work under wiki/tracks.md's subject rule, and takes the rule's stated exception rather than pretending otherwise: a named deliverable is blocked. ARM-tierP cannot write the Tier P statement RS-20260804i §12 makes available while an obligation stands that forbids calling any revised-sense verdict calibrated or comparing it with a Tier D result. It is the arm's last step and its declared closing condition.

2. Question

Does drop(naturalness) at the heavy 8-site accuracy dose — the single number on which Tier D failed at S034 (1.11) and at S086 (1.12) — survive the naturalness criterion string that D-20260802-13 installed; and does it survive on materials outside the Russian corpus both prior runs shared?

3. Materials

3.1 Six items, two provenances

item source reference words provenance
T8-GA Turgenev, «Свидание» Garnett 299 carried byte-identical from E-20260801f (S086)
T8-HB Turgenev, «Свидание» Hapgood 341 carried byte-identical
T8-K Korolenko, «Сон Макара» lead 394 carried byte-identical
T8-L Korolenko, «Лес шумит» lead 397 carried byte-identical
N8-A Machado de Assis, «O enfermeiro» ¶20–23 lead, this session 394 new
N8-B Machado de Assis, «O enfermeiro» ¶43–46 lead, this session 394 new

Each item's second text is its reference plus O4 × 8, the accuracy operator reused verbatim from E-20260725-tierD-ladder §3. materials/build.py records a SHA-256 of every carried string and re-derives nothing. The two new edit tables are materials/perturbations.md: eight sites each, two of each of O4's four types, sites numbered in textual order with the order asserted by the build.

The two new items are not independent of each other. Same story, same translator, same hour. They are two passages, not two works, and no figure treats them as independent draws.

3.2 The two criterion strings, verbatim

OLD — exactly the string tools/run_tierD.py put to the jury at S034 and S086:

Reads as fluent, idiomatic English prose WITHIN ITS OWN EVIDENT REGISTER AND PERIOD; nothing rings as translationese unless the source rings strange in the same place. Judge only whether the English is idiomatic on its own terms: do NOT penalise a text for being of an older or a more recent idiom than the other, and do not reward whichever is closer to present-day usage. Threshold competence, not distinction.

REVISED — condensed from wiki/goodness-senses.md §naturalness as it now stands, by the same practice that produced the OLD string from the entry as it then stood:

Reads as fluent, idiomatic English prose, scored ON THE TARGET TEXT ALONE. This criterion measures DISTANCE FROM TARGET-LANGUAGE NORM — not whether a departure is justified, faithful, purposeful, or good overall. A low score is not by itself a verdict that the translation is bad; it records markedness against the register named here. THE REGISTER ANCHOR FOR THIS ITEM IS: <anchor>. Judge only whether the English is idiomatic for that register: do NOT penalise a text for being of an older or a more recent idiom than the other, and do not reward whichever is closer to present-day usage. Threshold competence, not distinction.

Three edits, inseparable, and §0.3 binds every sentence about them:

  1. The source-conditional escape clause is gone — nothing rings as translationese unless the source rings strange in the same place — struck by D-20260802-13 and replaced on the list by a separate sense, perceived-source-carriage.
  2. A target-only frame is added — scored on the target text alone, distance from target-language norm, not whether a departure is justified, faithful, purposeful, or good overall.
  3. A register anchor is named rather than left to the juror's discovery of its own evident register and period. The critic's finding 2 is right that this is a second treatment and not a clarification: it can suppress register-uncertainty penalties, pull scores toward the named band, or inflate drops where the label mismatches the prose. The remedy is not to hide it but to name it as edit 3, to take the label out of the lead's hands (§3.3), and to report the effect split by anchor group (§6.2).

Held constant on purpose: the do-not-penalise-idiom-age instruction and threshold competence, not distinction are carried into REVISED unchanged, so the contrast is not also a contrast between an instructed item and a bare one. The other five sense definitions are byte-identical across arms, and analysis/verify.py asserts that the two payloads differ in exactly one line.

3.3 The register anchor is assigned blind, by a third party — stage 0

The v1 rule assigned the anchor from the reference's provenance. The critic's finding 3 is that a date of production is not a register, and it is accepted with the critic's own preferred remedy.

Stage 0, dispatched before anything else. A seat that is neither a juror (P1, P2, P5) nor the critic (P3) — P4 moonshotai/kimi-k3 — is shown the six undamaged reference texts, unlabelled, in a fixed order, told nothing about the source, translator, date, experiment or hypothesis, and asked to assign each exactly one of the three anchors wiki/goodness-senses.md permits: period-idiomatic, contemporary-vernacular, unmarked/literary-contemporary.

Whatever it returns is used, verbatim, in the REVISED string for that item. It is recorded in materials/anchors.json with the seat's one-line reason for each, and the raw body is preserved. A returned label outside the permitted three aborts the run. The lead makes no register judgment anywhere in this design.

Seat amendment A3. The seat was to be P4 moonshotai/kimi-k3; it returned finish_reason: length with 7,997 reasoning tokens and empty content at cap 8,000, $0.129576 (preserved as runs/stage0-anchors.attempt1-P4.json; note (bhf), tenth firing). Raising a $15/M-output cap after a seat has spent its whole allowance on hidden reasoning is not a fix, and this is a labelling task, so the seat was changed to mistralai/mistral-medium-3-5 — also probed-but-not-selected, also neither a juror nor the critic, and not a hidden-reasoning seat. It answered on first dispatch for $0.006381.

What stage 0 returned, before any juror was paid:

item reference assigned anchor
T8-GA Garnett, 1890s unmarked/literary-contemporary
T8-HB Hapgood, 1903 unmarked/literary-contemporary
T8-K lead, 2026 unmarked/literary-contemporary
T8-L lead, 2026 unmarked/literary-contemporary
N8-A lead, 2026 unmarked/literary-contemporary
N8-B lead, 2026 unmarked/literary-contemporary

Three consequences, all of them recorded here rather than discovered later.

  1. The anchor is a constant across all six items, so the anchor×item and anchor×cell confounds the critic's findings 3 and 10 identified do not arise in this run. Edit 3 of §3.2 is still an edit — the REVISED string names a register where OLD asked the juror to find one — but it names the same register everywhere.
  2. §6.2's anchor split is vacuous and is reported as vacuous. The provenance split {T8-GA, T8-HB} against {T8-K, T8-L} still runs.
  3. A blind reader shown Garnett's 1890s English and Hapgood's 1903 English does not call either of them period-idiomatic, on the seat's own reasons ("polished and literary without period or colloquial markers", "refined and literary, lacking period or vernacular traits"). That is one seat on one payload and licenses nothing; it is recorded because the design would otherwise have assigned those two items period-idiomatic by the provenance rule the critic struck, and the difference between the two rules is not hypothetical.

3.4 Contamination on the new references, measured before the design was written

tools/dependence_check.py, each new reference against Isaac Goldberg's complete 1921 English of the same story (Brazilian Tales, PG #21040, 4,220 words), reporting the longest run and the shared 7-gram count as CLAUDE.md requires:

7-grams 12-grams 15-grams longest run
N8-A 21 8 5 19 tokens
N8-B 6 0 0 11 tokens

The critic's finding 7 is that "the comparison is a text against a damaged copy of itself" blocks the independence objection and not the measurement objection: overlap can still move a number, above all if an edit site sits inside or beside the shared run. The check it asks for was run before dispatch. The 19-token run — "him back to life; it was too late — the aneurism had burst, and the colonel was dead. I went" — occupies characters 1332–1424 of N8-A's reference. No edit site intersects it. The nearest is site 7 (I drew back in terror, characters 1228–1249), which ends 83 characters before the run begins; the next nearest is site 8, 419 characters after it.

N8-A and N8-B are reported separately throughout, so that the item carrying 21 shared 7-grams and the item carrying 6 can be compared rather than pooled silently.

4. Arms, jury, and dispatch

4.1 Twenty-four payloads, 72 calls, plus stage 0

stage items criterion string payloads calls units
0 6 references, undamaged — — 1 —
1a 4 carried OLD 8 24 12
1b 4 carried REVISED 8 24 12
2a 2 new OLD 4 12 6
2b 2 new REVISED 4 12 6

A payload is one item in one ordering; a call is one dispatch of a payload to one juror. Order swap is mandatory: every item is dispatched twice with the slots swapped.

Stage order is 1a → 1b → 2a → 2b, so that what truncates first is what can be lost.

The seat probe is the longest payload in the run, not a synthetic one — critic finding 11, and note (bja): a probe bounds a seat on the payload shape it probed and on no other. T8-L under REVISED, order 0 — measured at 8,572 characters ≈ 2,143 tokens, the longest of the run's 24 payloads — is dispatched to all three jurors first; its three bodies are real data. If any seat returns an empty body or finish_reason: length there, the run stops before anything else is dispatched.

Partial-stage rule (critic finding 11). The obligation cell completes only if all 16 carried payloads score. If any is missing, no sentence from §6.1 may be written; the run reports what it has as an incomplete re-measurement.

4.2 Jury

P1, P2, P5, resolved from config/models.md at run time and logged as provenance — the same three seats as S034 and S086, because the OLD arm is a retest and a retest with a different jury is not one. P3 and P4 are excluded from the jury exactly as in those runs, which is also what frees them for stage 0 and the critic pass; this is a power limitation, not a finding about those models.

Six senses, scored 1–7 for each text, plus a forced overall preference, no ties. Strict JSON. Judgment is not parallelized (charter §6): strictly sequential dispatch.

4.3 Unit and cell statistics

The unit is (juror × item), with the two orderings averaged within it. A unit is +1 if the reference is preferred in both orderings, −1 if the second text is, 0 if split. Under a null of independent coin-flip preferences P(+1) = P(−1) = 0.25, P(0) = 0.5. Identical to E-20260801f §6.3.

drop(sense) for a unit is score_ref(sense) − score_var(sense), averaged over the two orderings. A cell's drop(sense) is the mean over its units. Cells are 12 units (carried) and 6 units (new).

All exact probabilities are computed by enumeration in analysis/rules.py, which imports nothing from tools/.

5. Failure criteria, registered before dispatch

id criterion if it fires
FC1 Detection must survive the criterion-string change. In each cell the targeted firing rule must fire: ≥ 8 of 12 at +1 and none −1 (carried, exact null P 0.000594) or ≥ 5 of 6 at +1 and none −1 (new, exact null P 0.003174). FC1 is an OVERALL-PREFERENCE rule, not a naturalness-targeted or accuracy-targeted one — a construct mismatch inherited from E-20260801f and named here rather than re-imported silently (critic finding 14). It is therefore accompanied by a reported check: drop(accuracy) > 0 in the same cell drop(naturalness) in a cell where FC1 did not fire is uninterpretable and is reported without a verdict. Every row of §6.1 is gated on FC1 firing in BOTH carried cells (critic finding 8). If FC1 fires but the accuracy check does not, the construct mismatch is flagged on the result page
FC2 Spillover, bounded proportionally to this run's own accuracy drop (critic finding 9 — the v1 bound reused the naturalness pass bar with no argument): |drop(accuracy)ᴼᴸᴰ − drop(accuracy)ᴿᴱⱽ| ≤ 0.25 × drop(accuracy)ᴼᴸᴰ on the carried cell. The absolute difference is reported too the payload difference is not confined to the naturalness line; the contrast is confounded and is reported as such
FC3a Retest, reported and not a gate (critic findings 4 and 13): |drop_natᴼᴸᴰ, this run − 1.12|, the OLD carried cell against RS-20260802-tierD-verdict §2. This is a statistic, not a threshold. v1 used it as a resolvability gate, which the critic showed is a random throttle: a lucky retest near 1.12 would license almost any wording effect and an unlucky one would suppress a large one reported in every case; licenses nothing on its own
FC3b Resolvability, defined entirely within-run. The effect is resolvable iff the permutation test of §6 P2p fires and |effect| ≥ δ_min = 0.25 (frozen here; the same scale as P4p, so the two are complementary) if it does not hold, nothing is licensed about the criterion string; both cell means are reported
FC4 Ceiling. If the reference's naturalness score is 7 in ≥ 50% of a cell's ratings, drops in that cell are floor-limited the cell's drop(naturalness) is reported as a lower bound, not as a value

FC1–FC4 are evaluated and reported whether or not they fire.

6. Predictions, registered before dispatch

6.1 The result → claim map, frozen

Every row is void unless FC1 fires in both carried cells and all 16 carried payloads scored. Otherwise the only licensed sentence is uninterpretable under FC1 or incomplete under the partial-stage rule.

# this run's OLDᶜ REVISEDᶜ FC3b resolvable what may be written
1 > 0.75 ≤ 0.75 yes On these stored scores, drop(naturalness) at threshold 0.75 would have decided the specificity condition FAILED at the 8-site dose under the OLD string and PASSED under the §3.2 REVISED string. Attribution is to the three-edit bundle only
2 > 0.75 ≤ 0.75 no The threshold is crossed by an effect no larger than δ_min or not distinguishable from unit-level noise. Both means reported; nothing licensed about the criterion string
3 > 0.75 > 0.75 yes The bundle moved the number and did not move it across the threshold. Both means and the effect reported
4 > 0.75 > 0.75 no Both measured means exceed 0.75 and the OLD–REVISED gap is within this run's own unit-level noise. Not licensed: that the number was shown stable, that it "survives", or anything about the struck clause (critic finding 6)
5 ≤ 0.75 either either Row 1's second clause is unavailable: this run's own OLD cell did not reproduce the failure, so no sentence may say the condition would have FAILED under the OLD string here. Report FC3a prominently; the run becomes a retest result about drop_nat's stability and says so
6 any ≤ 0.75 and |effect| ≤ 0.25 — P1p and P4p both hold. The threshold is crossed with essentially no movement between the two strings — i.e. the pass, if any, is not the criterion string's doing. This row overrides rows 1–3 wherever it applies

Charter §5.5 binds every sentence: on these stored scores, statistic S at threshold T would have decided D.

6.2 Pre-registered exploratory split (critic finding 10)

The wording effect is reported split by assigned anchor group and separately by provenance {T8-GA, T8-HB} against {T8-K, T8-L}. Stage 0 assigned every item the same anchor, so the anchor split is vacuous and is reported as vacuous; the provenance split runs. The primary stays pooled, and §6.1 may not claim homogeneity across groups. This is descriptive: 12 units do not support an interaction test.

7. Budget

Built from max_tokens, not from an assumed output length (note (abc)).

juror S034 measured max completion max_tokens here S086's cap headroom
P1 886 3,500 3,500 3.9×
P2 2,839 5,000 6,000 1.8×
P5 5,893 8,000 10,000 1.4×

P2's and P5's caps are lower than S086's, for headroom, and this is declared rather than silent. A cap cannot change the content of an accepted body; it can only make truncation likelier, and a truncated body is never scored — it is retried at a raised cap as an amendment written into config/budget.md before the call (note (bgt)). Worst-case input 2,600 tokens (measured longest payload 2,143).

juror price in / out per M worst case per call
P1 $1.00 / $6.00 $0.02360
P2 $1.50 / $7.50 $0.04140
P5 $1.65 / $3.30 — worst plausible provider, not list $0.03069

Per payload $0.09569.

stage payloads calls worst case, reserved
pre-run critic — 3 (2 dead) spent: $0.15947
0 blind anchors — 2 (1 dead) spent: $0.135957
1a + 1b carried 16 48 $1.531
2a + 2b new 8 24 $0.766
retry allowance 2 6 $0.191
total 24 82 $2.747

Against $2.9276 of headroom at session start. The reservation is per stage and is checked before the stage is entered; a stage that does not fit is not entered and the deferral goes to NEXT.md. Central estimate from S086's realised per-call mean on this exact task: $0.0054 × 72 ≈ $0.39.

8. Verification

analysis/verify.py recomputes every reported number from the stored raw bodies and asserts, separately (critic finding 16):

  1. every carried string against the SHA-256 in items.json;
  2. every OLD payload against tools.run_tierD.build_prompt on the same item — the retest claim depends on this;
  3. every REVISED payload differing from its OLD twin in exactly one line, that line beginning - naturalness:, and containing the anchor stage 0 assigned to that item;
  4. every stored body's prompt_sha256 against a rebuild of its payload;
  5. the exact null probabilities by enumeration against closed form, and the permutation test by full enumeration;
  6. the per-request cost sum.

Then mutation tests: each writes a corrupted copy, asserts the bytes on disk changed, and asserts that a named reported number changes.

9. Known limits, stated before the run

  1. One cell of one condition. Tier D is not re-run; no verdict on Tier D follows.
  2. Three edits, one measurement (§3.2, §0.3–0.4). Bundle attribution only.
  3. The units are not independent. 12 units are 3 jurors × 4 items; a juror with fixed taste contributes correlated units, so exact null probabilities are lower bounds, as at S086.
  4. The new cell is 6 units on 2 non-independent passages and is anchored by whatever stage 0 assigns; it can support a direction, not a rate, and P3p is descriptive unless its rule fires.
  5. The retest is at temperature 0.2, not 0. drop_nat is not a fixed property of the payload, which is why FC3a reports and FC3b gates from within-run structure only.
  6. The REVISED string has never been put to a jury on this task. If it behaves oddly, this run cannot distinguish the revision does something from this condensation of the revision does something; the condensation is printed verbatim in §3.2 so the distinction stays available.
  7. O4 as implemented changes length: the variants run +2.3% to +4.4% longer, and 5–6 of each 8 sites change the word count. Any inseparability claim is about this operator including its surface disruption, not about pure semantic error (critic finding 15).

10. What a clause-specific discharge would need, recorded so it is not lost

The cheapest non-confounded design, from the critic's finding 1: OLD against OLD-minus-the-escape- clause only — same register phrasing, no anchor token, no target-only frame; one contrast, the same payloads, 8 payloads / 24 calls on the carried cell. That would answer does the escape clause drive drop(naturalness)? This run does not ask it.

A process consequence, moved out of §6.1 because it is not a licensed result (critic finding 6): whatever this run returns, a Tier D design that reports a condition-3 number should state which naturalness string it was measured under. The two strings now both exist in the repository.