Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: framework/v0.3/entries/HB-evaluating.md · rendered 2026-09-09

Page metadata (front matter)
typeentry
idHB-evaluating
statusdraft
created2026-09-09
updated2026-09-09
sensesaccuracy, naturalness, perceived-source-carriage, voice, style-correspondence, affect, cultural-mediation, consistency
pairsJA→EN, RU→EN, FR→EN, PT→EN
provisionaltrue
internal-judgment-onlytrue
linksframework/v0.3/README.md, wiki/goodness-senses.md, wiki/findings/sense-dossier.md, wiki/findings/results/RS-20260802-tierD-verdict.md, wiki/findings/results/RS-20260906-tierD-verdict-v3.md, wiki/findings/results/RS-20260805e-wording-or-prose.md, wiki/findings/results/RS-20260802c-regime-scoring.md, wiki/findings/results/RS-20260803-a4-set.md, wiki/findings/results/RS-20260803f-craft-carriers.md, wiki/findings/results/RS-20260907-panel-judging-2.md, framework/tierD-repaired-rules.md, framework/control-arm-spec.md, config/models.md, wiki/findings/claims/CL-20260726-jury-detects-not-localises.md, wiki/decisions/resolved/D-20260905-01-tierD-primary-dose.md, wiki/method-notes.md, wiki/plan.md

Evaluating a translation: naming the sense, the jury that never calibrated, and what a provisional score licenses

Standing. Written at the project's close (2026-09-09, S257) as a consolidation of the record without the translation limb the procedure's step 6 requires — no fresh passage was translated under this entry, so it stays status: draft and its §5 Application reads 'not applied'. Every number here is a count recomputed from stored panel outputs — evidence class X2 for the count, X3 (panel-scored quality) for anything the count would be taken to say about a translation, and X3 is inadmissible, permanently. Tier D is NOT PASSED and, since 2026-09-06 (S250), EXHAUSTED (config/models.md; RS-20260906-tierD-verdict-v3): the one further redesign Tom authorized on 2026-09-04 failed, and by that authorization's stated consequence no further repair is authorized. The jury is not, and will not be, calibrated. Every panel score in this repository is provisional and internal-judgment-only for good, and the handbook's guidance rests on workshop findings alone. Nothing here says one rendering is better than another on a jury's word; the human bearing is two catalogued reception records (D-20260725-07; S-botchan), and the lead's own coding is marked where it occurs.

1. The problem

To evaluate a translation is to say in what way it is good, for whom, and on whose authority. The project answered the first with a controlled vocabulary of senses (wiki/goodness-senses.md) so that "good" can never be one number; the second with a purpose: parameter every evaluation must declare; and the third with a five-seat, non-Anthropic panel (config/models.md) that was to score the senses blind — but only after two certifications: Tier D (does the jury detect deliberately introduced damage, and confine it to the damaged sense?) and Tier P (does it reproduce a documented human ranking of two good translations, per sense?). Neither ever passed. A translator or pipeline now has the senses, a shelf of human anchors, mechanical checks, and a panel whose scores describe and never decide.

2. What the human record does — and what the jury did with it

This family has no published hands at loci; its analogue is the documented human judgment the project catalogued.

record · material what the human record does what the jury did with it source
The 1904 Nation / 1906 Athenaeum verdict on Garnett vs Hapgood, ratified D-20260725-07 (scoped to the Memoirs of a Sportsman cycle; admissible for accuracy and cultural-mediation only) · Turgenev «Свидание» split at a landmark; «Певцы», six loci splits the verdict by dimension — Hapgood the more accurate, Garnett the better English, neither better overall Tier D's held-out arm never fired (units to Garnett / split / Hapgood: 3/3/0 S034, 3/3/0 S086, 2/4/0 S250, rule ≥5 of 6; no unit ever went to Hapgood). Per sense, Garnett scored higher on naturalness on both items in all three runs; Hapgood higher on accuracy on both items at S034 and S086, split at S250. Tier P run 2: naturalness reproduced (Garnett 30–6), accuracy did not (19–17, one seat reversed), and all three seats named both translators — the blind was never blind; run 3: unattributed, accuracy 18–6 for Hapgood; told the names falsely, accuracy held 23–1, English-style fell 20–4 → 14–10, a coin RS-20260726d §4; RS-20260802 §4; RS-20260906 §6; RS-20260804c; RS-20260804i §1
The Botchan reception record (S-botchan) · three matched opening spans, JA→EN, Morri 1918 / Turney 1972 / Cohn 2005 is organised around whether the comedy lands — an affect difference Tier P run 1 (S014): the real test (Turney v Cohn) at record-fit voice 0.55, affect 0.42, literary-quality 0.43 (chance 0.5), naturalness 0.80 and register-cued; jurors recognised the work on 93% of probes RS-20260725-calibration-caseA
  1. The human records this project holds do not rank; they split by dimension. That converted Tier P from a ranking test into a dissociation test — and the half a jury holds under a false name is the accuracy half; the style half washes out.
  2. A jury's per-sense output matched the human record's direction on naturalness in every Tier D held-out run (three runs, both items) and in Tier P runs 2–3, and this never amounted to a pass: two items, six units per run, canonicity uncontrolled by construction, and — S034's own description (RS-20260726d §4) — the forced overall preference to the canonical translator every time.

3. What this project's own practice found

3.1 The calibration ladder — every run, its gates, its numbers

run · page design fired failed
S020 RS-20260725-tierD-ladder · $0.755 JA→EN: Shaw 1930 + the lead's Genji; 8 sites; P1/P2/P5; no held-out arm — pre-committed to no pass sham in band; accuracy detection 6/6, specificity +3.10 against 0.75 literary-quality detected 6/6, not localised (+1.08 flat on three untargeted senses); style-correspondence's +0.96 withheld by a pre-registered rule; de-marking (O3) split by passage — Shaw 0/6, lead 6/6 — and cultural-mediation on the de-marked Shaw rose 4.00 → 6.67
S034 RS-20260726d-tierD-heldout · $0.726 RU→EN: Garnett 1897 / Hapgood 1904 + the lead's Korolenko; first run with all three controls; source shown detection 9/9 at 8 and at 3 sites; specificity margin +2.04 at both doses; held-out 3 Garnett / 3 split / 0 Hapgood drop(naturalness) 1.11 at 8 sites (bar ≤ 0.75); the sham fired its lower branch (0 of 6 at +1), a branch with null probability 0.534, so no detection claim was licensed. NOT PASSED twice over
S040 / S050 / S055 RS-20260727b-tierD-rules, RS-20260728g, RS-20260729c · $0.28 the repair, no jury run sham branches matched by construction; power computed for the first time (≈15 units); stage-1 scale gate declared unrepairable → replaced by a prior positive control — (a specification: framework/tierD-repaired-rules.md R1–R5)
S083 + S086 RS-20260801f-tierD-stages12, RS-20260802-tierD-verdict · $0.812 (96 calls) the repaired instrument, RU→EN (Garnett/Hapgood + two lead items), P1/P2/P5, source shown, heavy dose primary positive control 2/1/0; sham 6/7/2 of 15, in band; held-out 3/3/0; heavy 12/0/0; light 12/0/0; light specificity: accuracy +2.88, margin +2.08, drop(naturalness) +0.50 heavy specificity: accuracy +4.17, margin +2.16, drop(naturalness) +1.12 against ≤ 0.75 — NOT PASSED on one number; the design had no row for heavy fails, light passes
S113 RS-20260805e-wording-or-prose · $0.887 the failed number re-run under the revised naturalness string (D-20260802-13), payloads byte-identical, plus two PT→EN items retest 1.083; revised string 0.750 exactly the gap is not distinguishable from unit noise (P 0.138) and is the reference coming down; on the PT items drop(naturalness) is 2.250 / 1.917. No verdict changed
S247 → S250 D-20260905-01 → RS-20260906-tierD-verdict-v3 · $0.733 (96 calls) the one authorized redesign: light dose primary, specificity as a 3:1 ratio (note (bke)), all eight senses, fresh materials, source withheld except in the positive control, P3 seated for P5 (note (bne)) positive control 3/0/0; heavy 12/0/0; held-out 2/4/0 (as predicted) sham 9/5/1 — upper branch, first in four runs (dispositive alone); light 9/2/1, one unit at −1, per-juror 0 of 3 — does not fire; specificity fails at both doses because consistency moved more than accuracy (+4.583 vs +3.500 heavy, +2.292 vs +2.125 light). Checkable mechanism, not a registered test: the three items whose light draw held a dropped-negation site fired 9/9. EXHAUSTED

Two things the ladder settled about the instrument, never about a translation: it detects damage at ceiling and does not confine it to the damaged sense at the heavy dose (CL-20260726-jury-detects-not-localises, status active, unrevised since S034; its registered falsifier is met at the light dose by S250 and the page has not been re-examined); and shown no source, its accuracy sense reads as consistency — catching only the subset of accuracy damage that produces an internal contradiction (note (btb)).

3.2 What a provisional panel score licenses — the judging installments

What a provisional score therefore licenses, all provisional, none evidential: (i) that the format separates grossly damaged prose from clean prose (≈4.3 points on accuracy, RU→EN, one operator); (ii) a resolution of about 0.23 scale points; (iii) that craft flattening at a large dose (37 sites in 669 words) is visible at 1.5 points; (iv) stability over time; (v) that naturalness reads period texture. It does not license: ranking two competent translations (their spread sits inside 0.17–1.33 points); any sense-localised reading at a heavy dose; any statement about readers (the seats are models); any framework recommendation.

3.3 The senses, as the dossier leaves them

Eight senses stand (wiki/goodness-senses.md): four site-level (accuracy, naturalness, style-correspondence, cultural-mediation), the home of 9 of 14 decision classes over 380 logged decisions and 48 of 53 sense assignments in a second corpus; four whole-text (voice, affect, consistency, plus the retired literary-quality), the home of 0 of 14 classes and 5 of 53 assignments — reachable only by reception records and cross-site censuses, never by a translator's log (sense-dossier §0; one coder's figures; an independent seat agreed on 9 of 24). literary-quality was retired (D-20260801-11) with no evidence of any kind after 83 sessions; purpose-fit became the parameter declared-purpose (D-20260801-10) after 16 of 18 crib-vs-reading-edition divergences decomposed into trades between existing senses. naturalness is scored on the target alone since D-20260802-13 — under a clause-free item, perceived licence did not move its score (licensed strangeness 1.667 vs matched unlicensed 2.000 vs fluent 5.000), and the old licence clause was struck as unusable by a no-source jury (the run never scored under it, so "inert" is not shown) — and its purpose-indexed refinement was measured and ratified no change (D-20260803-15). perceived-source-carriage is a statement about a reader, never a source–target relation: a rendering carrying none of the source's forms scored higher on it than one carrying all, and it has never been through Tier D. accuracy scored in a call with other senses carries a measured +0.38 inflation. affect must say which half it scores.

4. The options

Evaluating without a calibrated jury is a choice among instruments that each answer a different question, and the evidence prices each:

5. Guidance

For a translator

  1. Before judging a rendering, write down the senses you are judging and the reader you assume, by id, one purpose line; do not total them (wiki/goodness-senses.md usage rules 1–2). — evidenced (JA→EN, RU→EN, FR→EN) as reporting discipline; untested as anything more.
  2. Treat a low naturalness reading as markedness, not badness: state the register anchor, and expect a deliberately period or textured rendering to score lowest there and nowhere else. — evidenced (JA→EN, RU→EN, FR→EN; five items, one panel).
  3. Where a human record exists for the pair, compare against it per dimension and expect a split, not a winner; name its scope. — evidenced (RU→EN, one pair, one cycle; JA→EN, one record).
  4. Freeze your translator's log before anyone evaluates, and do not read it as a list of what you changed — have a second reader diff the two texts against the source. — evidenced (RU→EN, one operator, one log), the lead's own coding, internal-judgment-only.
  5. If you use a panel score, read it only for what §3.2 licenses: gross damage, a 0.23-point resolution, large craft flattening, stability. Never read a difference under about a quarter of a point, never a ranking of two competent renderings, and never a heavy-dose sense breakdown. — evidenced (JA→EN, RU→EN, FR→EN, PT→EN) as a limit; the licence itself is provisional.
  6. Do not build a fidelity judgment on a reader who cannot see the source: shown no source, a model jury's accuracy became consistency (HB-register item 6). — evidenced (RU→EN, one design; JA→EN, two builds).
  7. Measure independence before comparing: contamination on the span (not the work), the lead's self-match, the panel's slot bias (0.75–0.80 on identical texts). — evidenced (JA→EN, RU→EN, FR→EN), mechanical.

For a pipeline

  1. Declare: senses: by id, purpose: one line, the naturalness anchor, the affect half, source-shown or withheld; refuse to run without them. — untested as an enforced step (every result page carried it by hand).
  2. Gate: run tools/dependence_check.py on the span against every published comparator and against any prior lead rendering; discard on the frozen threshold. — evidenced (JA→EN, RU→EN, FR→EN) as executed by the A4 and craft-carrier runs.
  3. Anchor: if wiki/base/anchors/ or a ratified reception record covers the pair and problem, compare per dimension at the loci; else mark anchor: none and the evaluation internal-judgment-only. — untested as a pipeline step.
  4. Score (optional, provisional): three non-Anthropic jurors, blind, authorship stripped, each item alone against its source, two passes, with a byte-identical retest, a paraphrase null, a positive control and a per-juror degeneracy check (SD > 0.30 on ≥ 4 of 6 senses); flag a juror that fails the check and report it as a limit, not pooled silently; score accuracy in its own call where it is load-bearing. — evidenced (JA→EN, RU→EN, FR→EN) for the protocol as run at S094/S253; the degeneracy check evidenced once (S253); the own-call element untested as a step (usage rule 6, from RS-20260808d).
  5. Report: every figure beside its floor; per length-sign stratum for any paired comparison (R1, R4); the sentence what these scores license on the page, bounded as in §3.2. — untested as an automated step; hand-executed on every result page.
  6. File: the scores on the translation page under a heading that states Tier D EXHAUSTED and provisional; never into a handbook entry as evidence. — evidenced (RU→EN, FR→EN, JA→EN) at S253.

Human entry points. Step 1's purpose and anchor choice is the one decision no run has ever made without a person (declared-purpose); step 3's reading of a human record against a text is a critic's act the panel could not reproduce; step 4's verdict on a degenerate juror is a person's; step 6 is a person's decision to publish nothing. If nobody enters: default to purpose: readers of literary fiction in English who cannot read the source (the A4 default), no anchor, the step-4 protocol with its floors, and a filed score marked exactly as above — the pipeline then reports and never recommends.

Application. Not applied: written at close-out without a translation limb.

6. Not evidenced, and open

7. Sources consumed