Repository path: framework/v0.3/entries/HB-evaluating.md · rendered 2026-09-09
Page metadata (front matter)
| type | entry |
|---|---|
| id | HB-evaluating |
| status | draft |
| created | 2026-09-09 |
| updated | 2026-09-09 |
| senses | accuracy, naturalness, perceived-source-carriage, voice, style-correspondence, affect, cultural-mediation, consistency |
| pairs | JA→EN, RU→EN, FR→EN, PT→EN |
| provisional | true |
| internal-judgment-only | true |
| links | framework/v0.3/README.md, wiki/goodness-senses.md, wiki/findings/sense-dossier.md, wiki/findings/results/RS-20260802-tierD-verdict.md, wiki/findings/results/RS-20260906-tierD-verdict-v3.md, wiki/findings/results/RS-20260805e-wording-or-prose.md, wiki/findings/results/RS-20260802c-regime-scoring.md, wiki/findings/results/RS-20260803-a4-set.md, wiki/findings/results/RS-20260803f-craft-carriers.md, wiki/findings/results/RS-20260907-panel-judging-2.md, framework/tierD-repaired-rules.md, framework/control-arm-spec.md, config/models.md, wiki/findings/claims/CL-20260726-jury-detects-not-localises.md, wiki/decisions/resolved/D-20260905-01-tierD-primary-dose.md, wiki/method-notes.md, wiki/plan.md |
Evaluating a translation: naming the sense, the jury that never calibrated, and what a provisional score licenses
Standing. Written at the project's close (2026-09-09, S257) as a consolidation of the record
without the translation limb the procedure's step 6 requires — no fresh passage was translated under
this entry, so it stays status: draft and its §5 Application reads 'not applied'. Every number here
is a count recomputed from stored panel outputs — evidence class X2 for the count, X3 (panel-scored
quality) for anything the count would be taken to say about a translation, and X3 is inadmissible,
permanently. Tier D is NOT PASSED and, since 2026-09-06 (S250), EXHAUSTED (config/models.md;
RS-20260906-tierD-verdict-v3): the one further redesign Tom authorized on 2026-09-04 failed, and by
that authorization's stated consequence no further repair is authorized. The jury is not, and will
not be, calibrated. Every panel score in this repository is provisional and internal-judgment-only
for good, and the handbook's guidance rests on workshop findings alone. Nothing here says one
rendering is better than another on a jury's word; the human bearing is two catalogued reception
records (D-20260725-07; S-botchan), and the lead's own coding is marked where it occurs.
1. The problem
To evaluate a translation is to say in what way it is good, for whom, and on whose authority.
The project answered the first with a controlled vocabulary of senses (wiki/goodness-senses.md) so
that "good" can never be one number; the second with a purpose: parameter every evaluation must
declare; and the third with a five-seat, non-Anthropic panel (config/models.md) that
was to score the senses blind — but only after two certifications: Tier D (does the jury detect
deliberately introduced damage, and confine it to the damaged sense?) and Tier P (does it
reproduce a documented human ranking of two good translations, per sense?). Neither ever passed. A
translator or pipeline now has the senses, a shelf of human anchors, mechanical checks, and a panel
whose scores describe and never decide.
2. What the human record does — and what the jury did with it
This family has no published hands at loci; its analogue is the documented human judgment the project catalogued.
| record · material | what the human record does | what the jury did with it | source |
|---|---|---|---|
The 1904 Nation / 1906 Athenaeum verdict on Garnett vs Hapgood, ratified D-20260725-07 (scoped to the Memoirs of a Sportsman cycle; admissible for accuracy and cultural-mediation only) · Turgenev «Свидание» split at a landmark; «Певцы», six loci |
splits the verdict by dimension — Hapgood the more accurate, Garnett the better English, neither better overall | Tier D's held-out arm never fired (units to Garnett / split / Hapgood: 3/3/0 S034, 3/3/0 S086, 2/4/0 S250, rule ≥5 of 6; no unit ever went to Hapgood). Per sense, Garnett scored higher on naturalness on both items in all three runs; Hapgood higher on accuracy on both items at S034 and S086, split at S250. Tier P run 2: naturalness reproduced (Garnett 30–6), accuracy did not (19–17, one seat reversed), and all three seats named both translators — the blind was never blind; run 3: unattributed, accuracy 18–6 for Hapgood; told the names falsely, accuracy held 23–1, English-style fell 20–4 → 14–10, a coin |
RS-20260726d §4; RS-20260802 §4; RS-20260906 §6; RS-20260804c; RS-20260804i §1 |
The Botchan reception record (S-botchan) · three matched opening spans, JA→EN, Morri 1918 / Turney 1972 / Cohn 2005 |
is organised around whether the comedy lands — an affect difference |
Tier P run 1 (S014): the real test (Turney v Cohn) at record-fit voice 0.55, affect 0.42, literary-quality 0.43 (chance 0.5), naturalness 0.80 and register-cued; jurors recognised the work on 93% of probes |
RS-20260725-calibration-caseA |
- The human records this project holds do not rank; they split by dimension. That converted
Tier P from a ranking test into a dissociation test — and
the half a jury holds under a false name is the
accuracyhalf; the style half washes out. - A jury's per-sense output matched the human record's direction on
naturalnessin every Tier D held-out run (three runs, both items) and in Tier P runs 2–3, and this never amounted to a pass: two items, six units per run, canonicity uncontrolled by construction, and — S034's own description (RS-20260726d§4) — the forced overall preference to the canonical translator every time.
3. What this project's own practice found
3.1 The calibration ladder — every run, its gates, its numbers
| run · page | design | fired | failed |
|---|---|---|---|
S020 RS-20260725-tierD-ladder · $0.755 |
JA→EN: Shaw 1930 + the lead's Genji; 8 sites; P1/P2/P5; no held-out arm — pre-committed to no pass | sham in band; accuracy detection 6/6, specificity +3.10 against 0.75 |
literary-quality detected 6/6, not localised (+1.08 flat on three untargeted senses); style-correspondence's +0.96 withheld by a pre-registered rule; de-marking (O3) split by passage — Shaw 0/6, lead 6/6 — and cultural-mediation on the de-marked Shaw rose 4.00 → 6.67 |
S034 RS-20260726d-tierD-heldout · $0.726 |
RU→EN: Garnett 1897 / Hapgood 1904 + the lead's Korolenko; first run with all three controls; source shown | detection 9/9 at 8 and at 3 sites; specificity margin +2.04 at both doses; held-out 3 Garnett / 3 split / 0 Hapgood | drop(naturalness) 1.11 at 8 sites (bar ≤ 0.75); the sham fired its lower branch (0 of 6 at +1), a branch with null probability 0.534, so no detection claim was licensed. NOT PASSED twice over |
S040 / S050 / S055 RS-20260727b-tierD-rules, RS-20260728g, RS-20260729c · $0.28 |
the repair, no jury run | sham branches matched by construction; power computed for the first time (≈15 units); stage-1 scale gate declared unrepairable → replaced by a prior positive control | — (a specification: framework/tierD-repaired-rules.md R1–R5) |
S083 + S086 RS-20260801f-tierD-stages12, RS-20260802-tierD-verdict · $0.812 (96 calls) |
the repaired instrument, RU→EN (Garnett/Hapgood + two lead items), P1/P2/P5, source shown, heavy dose primary | positive control 2/1/0; sham 6/7/2 of 15, in band; held-out 3/3/0; heavy 12/0/0; light 12/0/0; light specificity: accuracy +2.88, margin +2.08, drop(naturalness) +0.50 |
heavy specificity: accuracy +4.17, margin +2.16, drop(naturalness) +1.12 against ≤ 0.75 — NOT PASSED on one number; the design had no row for heavy fails, light passes |
S113 RS-20260805e-wording-or-prose · $0.887 |
the failed number re-run under the revised naturalness string (D-20260802-13), payloads byte-identical, plus two PT→EN items |
retest 1.083; revised string 0.750 exactly | the gap is not distinguishable from unit noise (P 0.138) and is the reference coming down; on the PT items drop(naturalness) is 2.250 / 1.917. No verdict changed |
S247 → S250 D-20260905-01 → RS-20260906-tierD-verdict-v3 · $0.733 (96 calls) |
the one authorized redesign: light dose primary, specificity as a 3:1 ratio (note (bke)), all eight senses, fresh materials, source withheld except in the positive control, P3 seated for P5 (note (bne)) | positive control 3/0/0; heavy 12/0/0; held-out 2/4/0 (as predicted) | sham 9/5/1 — upper branch, first in four runs (dispositive alone); light 9/2/1, one unit at −1, per-juror 0 of 3 — does not fire; specificity fails at both doses because consistency moved more than accuracy (+4.583 vs +3.500 heavy, +2.292 vs +2.125 light). Checkable mechanism, not a registered test: the three items whose light draw held a dropped-negation site fired 9/9. EXHAUSTED |
Two things the ladder settled about the instrument, never about a translation: it detects damage
at ceiling and does not confine it to the damaged sense at the heavy dose
(CL-20260726-jury-detects-not-localises, status active, unrevised since S034; its registered falsifier is met at the light dose by S250 and the page has not been re-examined); and shown no
source, its accuracy sense reads as consistency — catching only the subset of accuracy damage
that produces an internal contradiction (note (btb)).
3.2 What a provisional panel score licenses — the judging installments
RS-20260802c-regime-scoring(S089, $0.578; five blind R06 → R04 self-revision pairs, JA/RU, P1/P2/P5, order-swapped; a sixth FR pair, revised knowing it would be scored, placed outside the primary): all six senses moved together, +0.333 to +0.467, a spread of 0.134 — not the shape of C12. The null of two identical texts returned a floor of exactly 0.0000 because all 18 cells recognised the identity (note (bhk)); the forced preference agreed with itself in both orders on 7 of 15 pairs, a coin; P2 and P5 chose slot A at 0.75 / 0.80.RS-20260803-a4-set(S094, $0.649; five filed lead translations rated alone against the source, 1–7, three jurors, two passes): every primary cell 5.000–6.667;naturalnessthe lowest mean and the only sense with spread — under its post-D-20260802-13wording it is a markedness meter, and the most archaising item sits at 5.000. Retest floor 0.233; a frozen ten-edit paraphrase null moved the jury 0.139, below its own noise — the working null; the damaged control fell 4.333 onaccuracy. Ceiling fractions P1 0.010, P2 0.771, P5 0.677. A seat reading only the translator's log named the weakest sense at r̄ 2.000 — and a predictor that always saysnaturalnessscores 1.500 (note (bht)).RS-20260803f-craft-carriers(S099, $0.746; Gorky RU→EN, LIVE vs FLAT — two English renderings of the same 494 Russian words, LIVE 669 words, with 37 logged non-propositional properties removed): blind seats named 5 of 6 flattening categories; identity control 0 false alarms. The gate found the operator had deleted content at five sites before amendment (note (bic)). Rated per sense by two jurors:voiceandstyle-correspondence+1.500,affect+1.250,naturalness+0.750,accuracy+0.500, against a retest floor of 0.250 — LIVE at 7.000 on three senses, so the gaps are floors.RS-20260907-panel-judging-2(S253, $0.629; the same seven A4 items, byte-identical, 159 sessions later, plus Chekhov RU→EN and Schwob FR→EN): mean shift 0.127 against a retest floor of 0.230, within 2× the floor on 6 of 6 senses — licensing only this panel scores the same texts consistently over time. P2 is degenerate: 89.8% of its cells at 7, SD > 0.30 on 2 of 6 senses (note (btc)). The paraphrase-floor criterion failed on P1 at exactly 0.5 on 3 of 6 senses. No positive control re-run; nine items judged, ~206 of ~208 candidate files unjudged.
What a provisional score therefore licenses, all provisional, none evidential: (i) that the
format separates grossly damaged prose from clean prose (≈4.3 points on accuracy, RU→EN, one
operator); (ii) a resolution of about 0.23 scale points; (iii) that craft flattening at a large
dose (37 sites in 669 words) is visible at 1.5 points; (iv) stability over time; (v) that
naturalness reads period texture. It does not license: ranking two competent translations (their
spread sits inside 0.17–1.33 points); any
sense-localised reading at a heavy dose; any statement about readers (the seats are models); any
framework recommendation.
3.3 The senses, as the dossier leaves them
Eight senses stand (wiki/goodness-senses.md): four site-level (accuracy, naturalness,
style-correspondence, cultural-mediation), the home of 9 of 14 decision classes over 380 logged
decisions and 48 of 53 sense assignments in a second corpus; four whole-text (voice, affect, consistency,
plus the retired literary-quality), the home of 0 of 14 classes and 5 of 53 assignments —
reachable only by reception records and cross-site censuses, never by a translator's log
(sense-dossier §0; one coder's figures; an independent seat agreed on 9 of 24). literary-quality was retired (D-20260801-11) with no evidence of any kind after 83 sessions; purpose-fit became the parameter declared-purpose
(D-20260801-10) after 16 of 18 crib-vs-reading-edition divergences decomposed into trades between
existing senses. naturalness is scored on the target alone since
D-20260802-13 — under a clause-free item, perceived licence did not move its score (licensed
strangeness 1.667 vs matched unlicensed 2.000 vs fluent 5.000), and the old licence clause was
struck as unusable by a no-source jury (the run never scored under it, so "inert" is not shown) — and its purpose-indexed refinement was measured and ratified no change
(D-20260803-15). perceived-source-carriage is a statement about a reader, never a source–target
relation: a rendering carrying none of the source's forms scored higher on it than one carrying all,
and it has never been through Tier D. accuracy scored in a
call with other senses carries a measured +0.38 inflation. affect must say which half it scores.
4. The options
Evaluating without a calibrated jury is a choice among instruments that each answer a different question, and the evidence prices each:
- Name the sense and the reader. No cost; the only move that makes an evaluation legible: which
of eight questions is being asked, under what
purpose:, whichnaturalnessregister anchor, whichaffecthalf, whether the source was shown, whetheraccuracywas scored alone. What it cannot do: settle anything. - Compare against a catalogued human record or precedent hand. The only human bearing available
(Tier 1 anchors under
wiki/base/anchors/; reception records ratified as decisions). What it buys: a dimension-split verdict a jury cannot fake by name recognition. What it costs: exists for few pairs and works; scoped narrowly; canonicity is not a text property and no measurement controls for it. - Mechanical checks (X2). Contamination measured before anything is designed (
CLAUDE.md§Contamination); length-sign strata and Metric A on every paired comparison (framework/control-arm-spec.mdR1–R5); a self-match check where the lead re-renders (note (bhb), 51 tokens). What they decide: whether a comparison is identifiable and independent — never whether it is good. - A provisional panel score. Licenses §3.2 (i)–(v) and nothing else; must carry a retest floor, a paraphrase null, a positive control and a per-juror degeneracy check, or its numbers cannot be read at all.
- The translator's own report. A log frozen before any evaluation is designed. What it buys: the decisions and alternatives. What it does not buy: a prediction of weakness (a no-information predictor beat the log reader) or a record of what was changed (five deletions in 37 logged sites went unlogged).
5. Guidance
For a translator
- Before judging a rendering, write down the senses you are judging and the reader you assume,
by id, one purpose line; do not total them (
wiki/goodness-senses.mdusage rules 1–2). —evidenced (JA→EN, RU→EN, FR→EN)as reporting discipline;untestedas anything more. - Treat a low
naturalnessreading as markedness, not badness: state the register anchor, and expect a deliberately period or textured rendering to score lowest there and nowhere else. —evidenced (JA→EN, RU→EN, FR→EN; five items, one panel). - Where a human record exists for the pair, compare against it per dimension and expect a split,
not a winner; name its scope. —
evidenced (RU→EN, one pair, one cycle; JA→EN, one record). - Freeze your translator's log before anyone evaluates, and do not read it as a list of what you
changed — have a second reader diff the two texts against the source. —
evidenced (RU→EN, one operator, one log), the lead's own coding,internal-judgment-only. - If you use a panel score, read it only for what §3.2 licenses: gross damage, a 0.23-point
resolution, large craft flattening, stability. Never read a difference under about a quarter of a
point, never a ranking of two competent renderings, and never a heavy-dose sense breakdown. —
evidenced (JA→EN, RU→EN, FR→EN, PT→EN)as a limit; the licence itself isprovisional. - Do not build a fidelity judgment on a reader who cannot see the source: shown no source, a
model jury's
accuracybecameconsistency(HB-registeritem 6). —evidenced (RU→EN, one design; JA→EN, two builds). - Measure independence before comparing: contamination on the span (not the work), the lead's
self-match, the panel's slot bias (0.75–0.80 on identical texts). —
evidenced (JA→EN, RU→EN, FR→EN), mechanical.
For a pipeline
- Declare:
senses:by id,purpose:one line, thenaturalnessanchor, theaffecthalf, source-shown or withheld; refuse to run without them. —untestedas an enforced step (every result page carried it by hand). - Gate: run
tools/dependence_check.pyon the span against every published comparator and against any prior lead rendering; discard on the frozen threshold. —evidenced (JA→EN, RU→EN, FR→EN)as executed by the A4 and craft-carrier runs. - Anchor: if
wiki/base/anchors/or a ratified reception record covers the pair and problem, compare per dimension at the loci; else markanchor: noneand the evaluationinternal-judgment-only. —untestedas a pipeline step. - Score (optional, provisional): three non-Anthropic jurors, blind, authorship stripped, each
item alone against its source, two passes, with a byte-identical retest, a paraphrase null, a
positive control and a per-juror degeneracy check (SD > 0.30 on ≥ 4 of 6 senses); flag a juror that fails the check and report it as a limit, not pooled silently;
score
accuracyin its own call where it is load-bearing. —evidenced (JA→EN, RU→EN, FR→EN)for the protocol as run at S094/S253; the degeneracy checkevidencedonce (S253); the own-call elementuntestedas a step (usage rule 6, fromRS-20260808d). - Report: every figure beside its floor; per length-sign stratum for any paired comparison
(R1, R4); the sentence what these scores license on the page, bounded as in §3.2. —
untestedas an automated step; hand-executed on every result page. - File: the scores on the translation page under a heading that states Tier D EXHAUSTED and
provisional; never into a handbook entry as evidence. —evidenced (RU→EN, FR→EN, JA→EN)at S253.
Human entry points. Step 1's purpose and anchor choice is the one decision no run has ever made
without a person (declared-purpose); step 3's reading of a human record against a text is a
critic's act the panel could not reproduce; step 4's verdict on a degenerate juror is a person's;
step 6 is a person's decision to publish nothing. If nobody enters: default to purpose: readers
of literary fiction in English who cannot read the source (the A4 default), no anchor, the step-4
protocol with its floors, and a filed score marked exactly as above — the pipeline then reports and
never recommends.
Application. Not applied: written at close-out without a translation limb.
6. Not evidenced, and open
- Pairs. Only Japanese, Russian and French have a panel competence screen (plus two Portuguese items at S113); of the sixteen source languages banked by S089, only those three have ever been judged, and ~206 of ~208 candidate files remain unjudged. No human reader ever sat on the jury; every "reader" is a model seat.
- The instrument's limits. Tier D failed on
naturalnessbleed (S034, S086) and then on the sham and the primary cell (S250); the S250 failure cannot separate the source-withholding choice from the P3-for-P5 swap.perceived-source-carriagenever went through Tier D;consistencywas scored once and dominated. - The senses. Venuti's challenge to
naturalnessis recorded and unanswered;affectrests on one reception record;literary-quality's route back is named and untaken. - Open: whether a source-shown instrument at the light dose would fire where the no-source one
did not — a checkable prediction (note (btb)) no design is authorized to test. Open: whether the
jury's de-marking preference at S020 rewards domestication, modernity, or a correct reading of a
1930 crib. Open: the regime comparison
(
wiki/plan.md§W2 step 5, R06 vs R04 vs an unspecified R03) — never run; its home isHB-pipeline.
7. Sources consumed
wiki/goodness-senses.md(senses,declared-purpose, usage rules,D-20260803-15) andwiki/findings/sense-dossier.md§0–§2 — consolidated as §3.3.- Tier P:
RS-20260725-calibration-caseA,RS-20260804c-peer-record,RS-20260804i-name-or-prose(runs 2–3 not in this family's index row) — §2. Tier D:RS-20260725-tierD-ladder,RS-20260726d-tierD-heldout,RS-20260727b-tierD-rules,RS-20260801f-tierD-stages12,RS-20260802-tierD-verdict,RS-20260805e-wording-or-prose,D-20260905-01-tierD-primary-dose,RS-20260906-tierD-verdict-v3,framework/tierD-repaired-rules.md,framework/control-arm-spec.md,config/models.md,CL-20260726-jury-detects-not-localises— §3.1. Judging:RS-20260802c,RS-20260803-a4-set,RS-20260803f,RS-20260907-panel-judging-2;wiki/method-notes.md(btc), (btb), (bke), (bne), (bhb); archive notes (bhk), (bhs), (bht), (bic), (bhu) — §3.2, §4. - v0.2 has no numbered section for this family: its §Standing and §8's last bullet (nothing about quality, still) are its whole content here.
- Withdrawn along the way, one line each: C12 (self-revision buys
naturalnesswithout movingaccuracy) — inadmissible as X3 (framework/v0.1/README.md§1 for the class rule, §4 for C12) and its shape not reproduced (RS-20260802c§1). The light dose is the cleaner instrument setting (CL-20260726,RS-20260802§2) — scoped to source-shown designs byRS-20260906-tierD-verdict-v3§8. The approach is not exhausted (RS-20260802§7) — superseded byRS-20260906§8. S034's sham lower-branch rule and the pooled scale-usage gate — replaced (RS-20260727b;RS-20260728g; the S034 verdict itself not reopened). S020's sham as a false-alarm floor — scoped byRS-20260727b(iii): not the operator's edit kind.RS-20260802cprediction 4 (the nulls are quiet) — the floor measured identity detection (RS-20260802c§3).RS-20260803-a4-set§5's log-predicts-weakness finding — deflated by its own no-information benchmark (same page, (bht)); its prediction-2 criterion — defective, denominator zero ((bhs)).RS-20260803fprediction P5 (naturalnessmoves toward FLAT) — refuted, with the A1 repair as a confound.sense-dossier§0's "0 of 14 classes" as a measurement — reduced to one coder's placement byRS-20260802-voice-warrant.naturalness's source-conditional escape clause — struck (D-20260802-13), and the decision's own "measured inert" reading withdrawn in place (the run never scored under the clause);literary-quality— retired (D-20260801-11);purpose-fit— moved to a parameter (D-20260801-10).