Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260803-a4-set/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260803-a4-set
statusfrozen
created2026-08-03
updated2026-08-03
sensesaccuracy, naturalness, voice, style-correspondence, cultural-mediation, affect
purposeReaders of literary fiction in English who cannot read the source, meeting these texts as reading editions rather than as cribs. Declared because every evaluation must state its purpose parameter (D-20260801-10).
internal-judgment-onlytrue
provisionaltrue
linkswiki/arms/ARM-first-judgment.md, wiki/findings/results/RS-20260802c-regime-scoring.md, wiki/goodness-senses.md, config/models.md, framework/control-arm-spec.md, workshop/translations/clos-des-ames/R04-v1/translation.md, wiki/findings/results/RS-20260802-tierD-verdict.md

E-20260803-a4-set — the A4 promise: five filed lead translations judged on their own terms, on a floor that means something

Frozen 2026-08-03 (S094) after the fresh translation and its log were committed at 9ab4c89 and before any call was dispatched. ARM-first-judgment step 2. Translation limb: T-clos-des-ames-R04-v1 (Arène, FR→EN), translated in session and frozen before this file existed.

0. Standing, stated first because it governs every number below

Tier D is NOT PASSED (RS-20260802-tierD-verdict, S086; config/models.md). No score this design produces carries evidential weight, none may support a framework recommendation, and every artifact it touches stays provisional and internal-judgment-only. ARM-first-judgment was unblocked by that failure and pre-committed to running in exactly this mode.

Two further obligations, both new since step 1 and both binding here. (i) D-20260802-13 struck naturalness's source-conditional escape clause on 2026-08-02; the wording put to the jury in §3 is therefore NOT the wording step 1 used, and the two runs' naturalness numbers are not comparable. (ii) The new sense perceived-source-carriage is deliberately not scored here — its own entry requires a departure-level record which this design does not collect, and it has never been through Tier D. Scoring it as a bare 1–7 item would violate the sense's own rule.

1. The question, and the subject-rule sentence

What do this project's own translations actually score, per sense, when read blind against their sources — and does the translator's own frozen record of what was hard predict where blind readers find the translation weakest?

Fifty-eight filed translations, sixteen source languages, and until S089 not one had ever been judged for quality (wiki/reassessment-2026-08-01.md §1.3). S089 scored six pairs — a draft against its own revision. No translation of this project has ever been judged on its own terms, which is what charter §3 A4 promises and what wiki/tracks.md §Deliverables names as T3's current deliverable.

Subject-rule sentence (wiki/tracks.md; continue-prompt.md §4.5): this unit teaches whether a translator's own written account of the difficulties of a passage locates where readers of the finished translation find it weakest — asked on six translations, five of them made years of sessions before anyone thought of scoring them. That is a claim about the practice of translating and about how translations are evaluated. The jury is the means, not the subject.

The wire between the limbs, in one sentence. The study limb cannot say that any difference between two of these translations is a difference at all without a resolution floor; a floor requires edits that are certainly quality-neutral; and only prose whose every decision the translator has just recorded can be edited with that certainty — so the session translated «Le Clos des Ames» in order to have a text it could paraphrase without changing its quality, and that paraphrase is the null this instrument has never had (RS-20260802c §3, §10.1; note (bhk)).

2. Materials

materials/build.py asserts every slice and refuses to emit a manifest if any assertion fails. materials/items.json carries SHA-256 0d32bd3ca022ec8b931bafc2f06e46e2b96b195d56ce78e2ffee99583b5e0053, recorded by the runner in every raw response file. Sources and translations are the filed artifacts, sliced source-and-target together.

item work pair regime role words
TAK Ōgai, 高瀬舟 (1916), opening JA→EN R04 primary 517
KUS Sōseki, 草枕 ch. VII (1906), first 3 of 12 ¶ JA→EN R06 primary 587
BAR Andreyev, «Баргамот и Гараська» (1898), unit B RU→EN R04 primary 441
MAR Sand, La Mare au Diable ch. II (1846) FR→EN R04 primary 569
PAN Arène, «La Mort de Pan» (1876) FR→EN R04 primary 477
CLO Arène, «Le Clos des Ames» (1876) — secondary FR→EN R04 secondary 503
CLO-P CLO + 10 frozen micro-paraphrase edits — — null 499
BAR-D BAR + variant-F1, 8 documented damage sites RU→EN R04 ctrl-pos 461

Three languages, not sixteen, and the reason is not laziness. Only Russian, French and Japanese have ever been screened for panel competence (config/models.md, S015: six items, ≥5/6 per juror). An accuracy judgment from a juror who cannot read the source is not a weak measurement but a different one, so the banked German, Chinese, classical-Chinese, Polish, Bengali, Italian, Hungarian, Swedish, Danish, Korean, Portuguese, Spanish, Finnish and Dutch artifacts are all excluded by that rule. This is a real narrowing of the "cross-language set" the arm asked for and is declared as one. Within the screened three, the regimes that exist are R04, R06 and R07; the one R07 artifact in a screened language with an aligned source span (T-may-night-golova-R07-v1) is 1,457 words, outside the length band, so the set is R04 ×4 and R06 ×1.

Length band 441–587 words, by construction. KUS is sliced to its first 3 of 12 paragraphs, source and target together, by the pre-registered rule take the fewest leading paragraphs whose English reaches 400 words; every other item is the whole filed span. Absolute rating of a longer text gives more opportunities to find a fault, so length is a confound and it is bounded here rather than argued away.

Two items share an author. PAN and CLO are both Arène, both R04, both the lead. That is a declared limitation on work-spread and a small compensation: it is the one place where a translation made before this design existed and one made inside it can be compared on the same author.

2.1 The null, and the edit list that makes it a null

RS-20260802c §3 established that step 1's null did not work: two byte-identical texts, all eighteen cells recognised the identity, the floor came out at exactly 0.0000, and "above the noise floor" became a sentence the instrument could not say. The remedy its own critic proposed and this design executes: a meaning-preserving micro-paraphrase with a frozen edit list.

CLO-P is CLO with exactly ten substitutions, each asserted by build.py to match once and only once, all listed in materials/edits.json:

# from to
1 the whole of his close his whole close
2 very old and badly kept very old and poorly kept
3 right at the bottom at the very bottom
4 must at one time have formed must once have formed
5 meeting it for the first time coming on it for the first time
6 having long since worn away having long ago worn away
7 in cherry time at cherry time
8 of every bourgeois felicity of all bourgeois felicity
9 has gone on trembling ever since has been trembling ever since
10 in a group around him in a group about him

Every edit is a function-word or near-synonym substitution inside one clause. None changes propositional content, sentence structure, paragraph shape, punctuation, the handling of any culture-bound item, or any decision entered in the translator's log — the thirteen logged decisions were checked one by one against the list before it was frozen. Net −4 words on 503.

Why this null works where step 1's did not: the jurors never see the two texts together. Every call in this design rates one text. CLO and CLO-P arrive in different calls, in different passes, with no shared context, so no juror can recognise that a relationship exists — which is exactly what made step 1's identical-text null measure identity detection instead of resolution.

2.2 The contamination gate, which fired twice before admitting the fresh work

Run before any locus was chosen and before anything was translated (CLAUDE.md; note (bcd)), on a discard rule frozen in advance: longest common run ≥ 12 tokens, or ≥ 1 shared 12-gram → discard.

candidate comparator (whole, never displayed) run 12-grams nulls verdict
Daudet, «Le Phare des Sanguinaires» anon., Letters from My Windmill (PG 30442) 14 3 4/4/3, 0 DISCARDED
Sand, La Mare au Diable ch. XI Ives (PG 12816) 16 13 4/4, 0 DISCARDED
Sand, La Mare au Diable ch. XI Sedgwick (PG 21993) 13 2 3/2, 0 DISCARDED
Arène, «Le Clos des Ames» none exists (two search routes) — — — admitted

The second firing lands on a work this design also scores, and it forced a re-check of a published figure. T-mare-au-diable-R04-v1 is filed contamination: none, measured at S030 on ch. II against those same two comparators. Ch. XI, same translator, same comparators, returns 16 tokens and 13 shared 12-grams. That does not falsify the ch. II figure, so the ch. II figure was recomputed on the filed text before MAR was admitted as an item: materials/gate3.py returns 9 tokens against Sedgwick and 10 against Ives, 0 twelve-grams in both, nulls 2–4 — the published figure reproduces and MAR stays. What the pair of measurements establishes is that contamination: is a property of the span, not of the work, which the project's own front-matter field does not distinguish. That is method work and is recorded in the result's limits, not made this session's subject.

3. Procedure

4. Analysis, fixed before the run

Let x(i, s, j, p) be the score for item i, sense s, juror j, pass p. Let x̄(i, s, j) = mean over p, and X(i, s) = mean over j.

  1. Primary table: X(i, s) for all 8 items × 6 senses, printed whole, no aggregation hidden.
  2. The retest floor R(s) = mean over (i, j) of |x(i,s,j,1) − x(i,s,j,2)|, and R = mean over s. Per-juror floors reported separately. This is the instrument's noise on identical text.
  3. The paraphrase floor P(s) = |X(CLO, s) − X(CLO-P, s)|, and P = mean over s. The comparison of P with R is the null's whole point: if P ≈ R the ten edits are quality-neutral to this jury and R is the operative floor; if P > R the edits are not quality-neutral and P is the operative floor. Both are reported whatever they show.
  4. The positive control D(s) = X(BAR, s) − X(BAR-D, s). Registered gate — see F1.
  5. No difference between two items on a sense is described as a difference unless it exceeds the operative floor, and every reported difference is printed with the floor beside it.
  6. The log-prediction test. For each of the six logged items, rank the six senses by X(i, s) ascending (1 = lowest), ties given the mean rank. Let r(i) be the rank of the sense the log-reading seat named first. The statistic is r̄ = mean over the six items. Under the null of no relation, each r(i) is uniform on {1..6}, so E[r̄] = 3.5; the exact null distribution is computed by full enumeration of 6^6 = 46,656 outcomes and the achieved one-sided p is reported. Criterion, fixed here: r̄ ≤ 2.20 fires. The lead's own prediction is scored the same way and reported beside it, labelled internal-judgment-only.
  7. Nuisance checks, computed and reported whatever they show: (a) Spearman between item word count and X(i, s) over the five primary items, per sense; (b) the fraction of cells at the ceiling (7) and at the floor (1), per juror; (c) pass-1 minus pass-2 mean per juror, as a drift diagnostic; (d) per-juror mean score, since an absolute scale has no order-swap to cancel juror-level offset.
  8. control-arm-spec R3/R4: five primary items against R4's 22–30 per stratum. No estimate, no interval and no stratified summary is claimed anywhere. Every number is descriptive.

5. Registered predictions

# prediction
1 The retest floor R is greater than 0 and less than 1.00 scale points.
2 The paraphrase floor P is within 0.50 of R — the ten edits are quality-neutral to this jury.
3 F1 fires: D(accuracy) ≥ 2.00 and D(accuracy) is the largest of the six D(s).
4 No primary item scores below 4.0 on any sense — this jury does not find any of these translations bad.
5 accuracy has the highest mean over the five primary items of the six senses; affect or style-correspondence the lowest.
6 The log-prediction test fires: r̄ ≤ 2.20 for the P3 seat.
7 The between-item spread on any one sense over the five primary items exceeds the operative floor — i.e. this instrument can tell these five translations apart at all.

6. Failure criteria

7. Cost, and the reserve declared before dispatch

stage calls cap worst case from max_tokens
pre-run critic 1 24,000 $0.19
log-reading seat (P3) 6 6,000 $0.22
rating, pass 1 24 6,000 $0.90
rating, pass 2 24 6,000 $0.90
declared worst case 55 $2.21

Worst cases are built from max_tokens and the highest list price on the seat, per note (abc), and sized from measured reasoning appetite per note (bhq) — the S092/S093 bodies spent 2,056–7,119 completion tokens on reasoning for short answers, so 6,000 is the answer plus that appetite. 2026-08-03 is a fresh UTC day with no rows; the full $5.00 is available. The realistic expectation from step 1's 60 calls at $0.516 is ≈ $0.50.

8. What this design cannot do, written before it runs

  1. The jury is not calibrated. Tier D NOT PASSED. Nothing here carries evidential weight.
  2. The item format is new. Tier D and step 1 both used paired presentation; nothing established about this jury transfers automatically to absolute rating, which is why F1 exists as a gate rather than a diagnostic.
  3. n = 5 primary items, against control-arm-spec R4's 22–30. No estimate is claimed.
  4. Three languages, four regimes reduced to two, two items sharing an author (§2).
  5. The log-prediction stage is partly a readback. Logs written after the project adopted its own sense vocabulary use that vocabulary — T-mort-de-pan-R04-v1's log has section headings named Accuracy against the source, Style correspondence, Naturalness and consistency. A seat reading that log is being handed the answer in the project's own words. Per item, whether the log names senses explicitly is recorded in the result table, and the test is reported both overall and restricted to the logs that do not.
  6. One translator, one agent. Nothing here is about translation in general.

9. Amendments from the independent pre-run critic pass

runs/critic_kimi-k3_try1.txt, seat P4 (moonshotai/kimi-k3), verdict NEEDS-AMENDMENT, eleven findings, three BLOCKING. All eleven accepted; one accepted with its mechanism corrected and half of its remedy refused in writing. Applied before any rating or log call was dispatched. Note (rr)'s streak, broken at S093, resumes. Note (b)'s prescribed fix — cap the reasoning effort on this seat, do not rely on a generous max_tokens — was applied on the first dispatch this time, and the body came back stop at $0.07026 against S093's $0.2215 for zero characters on the same slug.

A1 (finding 1, BLOCKING) — the null's two quantities are put on the same footing. R was a mean absolute per-call retest difference and P a difference of juror-averaged means, so juror disagreement in direction cancels in P and noise shrinks it. P is now computed per juror, P(s, j) = |x̄(CLO, s, j) − x̄(CLO-P, s, j)|, and compared against that same juror's retest differences. The pooled figure is retained as a secondary line only.

A2 (finding 2, BLOCKING) — accepted in its risk, corrected in its mechanism, half its remedy refused with a reason. The finding is that a juror seeing the same source twice with near-identical targets can infer the relationship, converting absolute rating into implicit comparison. The mechanism requires memory across calls and there is none: call.py issues one stateless HTTP request per rating with a single user message and no conversation history, so no juror can know that any other item exists. This is the same correction E-20260802c made to the same argument (its A1). Refused: showing CLO and CLO-P once each, because it would destroy the per-juror pairing A1 requires. Adopted: the dispatch order within a pass is now forced to place CLO and CLO-P, and BAR and BAR-D, at maximal separation, as belt-and-braces; and the result page says recognition is precluded by statelessness, not merely mitigated.

A3 (finding 3, BLOCKING) — the log-prediction null is now conditional on the observed tie structure. Enumerating 6^6 uniform outcomes is invalid when X(i, s) has ties, which ceiling clustering makes likely. The null is now: for each item, the six senses' ranks are the observed rank multiset (ties at mean rank), and under the null the seat's named sense is one of the six uniformly at random. The exact null is the enumeration of the product of the six items' observed multisets — still 6^6 = 46,656 outcomes, but drawn from the real ranks. The uniform-{1..6} figure is not reported.

A4 (finding 4) — the claim is bounded to this translator. Every statement of the log-prediction result is restricted to this lead's logs for these six items; §1's subject-rule sentence is read under that restriction, and no sentence about translators in general may be written from it.

A5 (finding 5) — prediction 3 is aligned with gate F1. It now reads: D(accuracy) ≥ max(2.00, 3 × R(accuracy)) and D(accuracy) is the largest of the six D(s). The registered prediction and the format gate can no longer diverge.

A6 (finding 6) — two predictions tightened and one reframed. Prediction 7 now requires the between-item spread to exceed the operative floor on at least three of six senses, not one. Prediction 4 is restated per juror: no juror's mean for a primary item falls below 4.0 on any sense, with the juror offsets reported beside it. Prediction 1's band is left as registered and is labelled in the result as the weak prediction the critic says it is.

A7 (finding 7) — F2 governs the null's status, and prediction 2 is restated in F2's terms. Prediction 2 now reads: P(s, j) ≤ 2 × R(s, j) on at least 5 of 6 senses for every juror. The aggregate |P − R| ≤ 0.50 form is withdrawn.

A8 (finding 8) — the edit list gets an independent neutrality check before dispatch. A new stage edits, dispatched before any rating call: seat P3 is given the ten edits in context and the six sense definitions, and asked, for each edit, which senses it would move and in which direction. P3 has not seen the design, the predictions or the critic. Its verdict is recorded whatever it says and is reported beside the measured P. It is a check, not a gate: the edit list is frozen and is not revised on it, because revising the null after an opinion about it is the move this project's discipline exists to stop.

A9 (finding 9) — absence of a comparator is not absence of contamination. CLO is admitted because no English rendering exists to collide with; that says nothing about whether the lead carries the French source from training. Recorded in §8 and on T-clos-des-ames-R04-v1. The word "uncontaminated" is not used of CLO anywhere.

A10 (finding 10) — R is computed over the five primary items only. The retest differences of CLO, CLO-P and BAR-D are reported separately, since damage and paraphrase may interact with pass.

A11 (finding 11) — KUS's fragment status is a per-item covariate. KUS is the first 3 of 12 paragraphs; every other item is a whole filed span. No KUS-versus-whole-item comparison is made on affect or style-correspondence, both of which are properties of a completed unit.