Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: framework/v0.3/entries/HB-pipeline.md · rendered 2026-09-09

Page metadata (front matter)
typeentry
idHB-pipeline
statusdraft
created2026-09-09
updated2026-09-09
sensesaccuracy, naturalness, voice, style-correspondence, cultural-mediation, affect, consistency
pairsJA→EN, RU→EN, FR→EN, FA→EN, NO→EN, IT→EN, DE→EN, RO→EN, SR→EN
provisionaltrue
internal-judgment-onlytrue
linksframework/v0.3/README.md, workshop/regimes/README.md, workshop/regimes/R01-single-pass.md, workshop/regimes/R02-draft-revise.md, workshop/regimes/R04-lead-close.md, workshop/regimes/R05-serial-long-work.md, workshop/regimes/R06-lead-single-pass.md, wiki/findings/results/RS-20260724-selfrevise-first.md, wiki/findings/results/RS-20260802c-regime-scoring.md, wiki/findings/results/RS-20260727c-arm-identifiability.md, wiki/findings/results/RS-20260728c-length-matching.md, wiki/findings/results/RS-20260801e-lead-carryover.md, wiki/findings/results/RS-20260907-panel-judging-2.md, wiki/findings/results/RS-20260906-tierD-verdict-v3.md, framework/control-arm-spec.md, framework/closure.md, framework/v0.1/README.md, wiki/method-notes.md, wiki/reassessment-2026-09-04.md, workshop/translations/README.md, workshop/translations/verter/plan.md, workshop/translations/gulistan-bab2/register.md, wiki/arms/ARM-gulistan.md, framework/v0.3/entries/HB-change-of-language.md, framework/v0.3/entries/HB-footing-and-address.md, framework/v0.3/entries/HB-realia.md, framework/v0.3/entries/HB-register.md, framework/v0.3/entries/HB-telling-the-reader.md

The pipeline: the steps a translation went through here, which regime each step is, and where a person can enter

Standing. Written at the project's close (2026-09-09, S257) as a consolidation of the record without the translation limb the procedure's step 6 requires — no fresh passage was translated under this entry, so it stays status: draft and its §5 Application reads "not applied". The evidence is X2 (token overlap, word counts and filing tallies recomputed from stored texts) plus the project's own frozen regime specifications and translator's logs (internal-judgment-only). The two regime comparisons that scored anything (RS-20260724-selfrevise-first, RS-20260802c-regime-scoring) are X3, panel-scored; Tier D is NOT PASSED and now EXHAUSTED (config/models.md, RS-20260906-tierD-verdict-v3), so nothing here says one regime's output is better than another's on a jury's word — what those runs contribute is instrument findings and a provisional direction, labelled as such. No v0.2 section belongs to this family (its 52 numbered sections contain none on regimes or process); the material is the regime pages, v0.1 §1 and §5, and the result pages in §7.

1. The problem

A regime is "a fully specified way of producing a translation" and comparing regimes is "the workshop's basic experiment" (PROJECT.md §3); the framework must "treat human involvement as a designed, parameterized dimension — specifying where human judgment can enter and what its presence or absence changes" (§1). A translator meets this as a sequence of decisions made before the first sentence: which work, checked against what, drafted how, revised or not, logged when, frozen when, judged by whom. A pipeline meets it as the same sequence with gates. The record ran one such sequence 318 times (318 filed translation.md files, every one translated-by: lead, by id, 96 R04, 92 R06, 8 R05) and compared regimes twice by jury and three times by machine. What follows is that sequence as executed.

2. What published translators do

No anchor here describes a published translator's working procedure, so this family has no hands table of the usual kind. Published translators enter the pipeline at one step only — as comparators at the contamination gate — and the table is what that gate measured against them (tools/dependence_check.py: shared 7-grams are context, 12- and 15-grams the signal, plus the longest common run in tokens and name-excluded counts).

pair · lead text comparator longest run shared 12-grams null floor verdict · source
JA→EN · T-kusamakura-vii-bath R06/R04 (1,210/1,206 words) Takahashi 1927, the only reachable English 8 (a doubled proper name) 0 (4 shared 7-grams) 5 (Morri 1918) suspected — bounds contamination from one side only · RS-20260727c-arm-identifiability §8
FR→EN · Maupassant «Menuet», gate translation anon. 1903 (PG #3086) 14 (ordinary prose) 3 3 / 3 / 2 DISCARDED · RS-20260802c §9
FR→EN · Villiers «La Torture par l'espérance», gate translation anon. (PG #29704) 12, nine of them a proper name and a title 1 3 / 2 / 2 DISCARDED — "the rule was frozen and had already fired once" · RS-20260802c §9
FR→EN · Arène «La Mort de Pan» none found (Gutenberg, Internet Archive) — — — admitted on an absence, "weaker than a measurement" · RS-20260802c §9
RO→EN · Slavici «Popa Tanda», selection probe Byng 1921 (PG #38991) 12 (ordinary prose) 1 (10 shared 7-grams) 3 would be suspected; work rejected on length · verter/plan.md §2
SR→EN · Lazarević «Вертер» (14,293 words) none on four documented routes — — — none, not measured · verter/plan.md §2
FA→EN · Gulistan span F (R05) Eastwick 10 (14 shared 7-grams) — — clean against the published hand; DEPENDENT? against the lead's own R52-v1: 536 shared 7-grams, 153 15-grams, a 51-token run → high · ARM-gulistan

The regularities, as what the measurements did: (1) the gate's threshold sits on a scale calibrated once ("0 Ovid → 21 Turgenev", RS-20260727c §8) and discards on a run of ordinary prose above a same-volume null floor, not on fame (Villiers, table); (2) the lead's largest measured overlaps are with itself — 37 contiguous tokens across sessions (RS-20260801e), 51 when a span was re-rendered minutes after reading its own predecessor (notes (bhb), (bta)); (3) with one reachable comparator a low run bounds contamination from one side and the declaration stays suspected; with none, the honest declaration is none, not measured with the routes recorded, which "does not license treating the work as an independence measurement for anything" (verter/plan.md §2).

3. What this project's own practice found

The pipeline as run, step by step. Every step below was executed by the lead agent in session following continue-prompt.md §5 and CLAUDE.md §Contamination, except where a script is named.

  1. Selection gate. Public domain, verified on the manifest before entry (workshop/canon/README.md); from S171, confirm the published comparator exists, in reach, before choosing (note (bmw)); from S244, Tom's long-work criteria (post-1880, rarely translated, 8,000–15,000 words, both sides free "where possible"). Last executed S256: four candidates fetch-verified, two measured, «Вертер» chosen with the missing comparator documented (verter/plan.md).
  2. Contamination measurement, before any locus is chosen — one dependence_check.py run per candidate unit, longest run and shared 7-gram count, because "run length alone has failed as a proxy three times" (CLAUDE.md). Over the whole filing tree: 121 none, 138 suspected, 59 high of 318 declarations (mechanical tally, S257). R05's Known limitations ask the arm to re-measure at least once mid-work; ARM-gulistan did after four spans, and the last re-measure moved a declaration to high (§2, row 7).
  3. Regime choice. R04 (lead close translation, source-only, draft then self-revise, frozen v1.0 S041) is the default; R05 (serial long work) for anything spanning sessions; R06 (lead single pass) is a by-product, not a separate choice: R04 §Procedure 2a freezes the draft as its own artifact, so every R04 run contains an R06 output. The API regimes R01 and R02 (frozen v1.0, S010) ran once, on two Japanese works × two panel translators × three reps, in workshop/experiments/E-20260724-r01r02-selfrevise/runs/ — 0 of 318 filed translations were produced by an API regime.
  4. Translate from the source alone. R04 §1: source and non-translational apparatus only; "it does not read any published rendering of the same passage into any target language before the translation and its log are frozen"; priming, if it occurred, is declared on the artifact. The S256 probe chose a locus whose English had not yet been seen rather than declaring priming after the fact (verter/plan.md §2).
  5. Freeze the draft, then self-revise. "The commit is the freeze" (R06 §4). Executed for the first lead pair at b2a3199 (RS-20260727c §8). What the pass did, measured: a model self-revision changes a mean 0.032 of words and came back shorter in 10 of 12 API pairs (RS-20260724 finding 4; RS-20260727c); the lead's own revisions moved a few words in either direction — −4 on Kusamakura, longer in 4 of 5 at S089 (RS-20260802c), +14 on Kleist (RS-20260728c); a revision of its own draft shares a 51-token run with it (RS-20260801e §2, the positive control) — the two arms of a pair are near-identical by construction.
  6. Translator's log, frozen before any evaluation is designed — "decisions and alternatives, not quality claims" (R04); evaluative sentences carry internal-judgment-only; "a log written after seeing scores is worthless, which is what the freeze protects" (workshop/translations/README.md). The log is self-report: "decisions made without noticing them do not appear" (R04). Tallies over a log are computed by a committed script, never read as prose (note (bey)).
  7. Filing under workshop/translations/<work>/<regime>-v<N>/translation.md with type: translation, translated-by:, contamination: and its basis, lead translations have no run/ (cost $0, never ledgered); tools/build_index.py regenerates the index.
  8. Optional blind panel judging, provisional permanently. Three non-Anthropic jurors, authorship stripped, order-swapped, a meaning-preserving micro-paraphrase null (since S094). Last executed S253 (RS-20260907-panel-judging-2, $0.629): the seven items shared with S094 moved a mean 0.127 scale points across 159 sessions against a retest floor of 0.230; one juror (P2) flagged degenerate by a pre-registered check (89.8% of cells at the ceiling). About 206 of about 208 candidate files remain unjudged.
  9. Collation and register, for serial work (R05, draft v0.1 — never frozen; eight filings). Whole cleaned source committed once, spans named by paragraph index; artifact append-only, revisions only by numbered erratum (Gulistan: five spans, 9,035 English words, D1–D81, errata filed: none); a binding register with an unresolved section (gulistan-bab2/register.md, opened at span B, closed at span F at nineteen rules); "decisions made by drift are recorded as such" (R05 rule 6).

What the regime comparisons found. Five pages, two scored (X3), three mechanical (X2):

4. The options

5. Guidance

For a translator

  1. Confirm a published comparator exists and is reachable before choosing the work; if none is, decide with that documented, not assumed. — evidenced (FR→EN, RO→EN, SR→EN).
  2. Measure contamination on the candidate before choosing a locus: longest common run and shared 7-gram count against the comparator and a same-volume null. Discard on a frozen rule; with one comparator only, declare no lower than suspected; with none, none, not measured with the routes. — evidenced (JA→EN, FR→EN, RO→EN, SR→EN, FA→EN).
  3. Choose R04 by default, R05 for anything that will span sittings; write the draft out and commit it before revising. The commit yields an R06 output for free; a draft reconstructed afterwards is not one. — evidenced (JA→EN; 96R04and 92R06filings).
  4. Translate from the source and non-translational apparatus only; open published English after the freeze; declare any priming. — evidenced (RO→EN selection probe; standing practice).
  5. Log decisions and alternatives as you write, never quality; freeze the log with the commit; compute any tally by script. — evidenced as universal filed practice (every filing carries a frozen log; pairs beyond pairs: not separately verified).
  6. Never judge your own output; if it is scored, every score is provisional, permanently. — evidenced (JA→EN, RU→EN, FR→EN judged S089/S253).
  7. Expect a self-revision to change about 3% of words and to move length by a few words in either direction; do not expect any measurement to show which sense it buys. The 3% is one JA→EN API run; the lead's own revisions ran longer in 4 of 5 pairs at S089, +14 on Kleist, −4 on Kusamakura; the only scored reading, an even small lift, is X3 and licenses nothing. — the length figures evidenced (JA→EN, RU→EN, DE→EN); the sense clause untested.
  8. For serial work: commit the whole cleaned source once; name spans by index; repair a frozen span only by numbered erratum; keep a register with an unresolved section; re-measure contamination mid-work against the published hand and your own earlier spans. — evidenced (FA→EN).
  9. Do not re-render a span for comparison against a prior rendering you have just read and call it independent: draft blind and diff afterwards, or declare the dependency in the design. The self-match reached 51 tokens. — evidenced (FA→EN), note (bta).

For a pipeline

  1. Select: verify PD status and comparator reachability from tool results; emit the search log. — evidenced (RO→EN, SR→EN), executed by the lead in session.
  2. Gate: run dependence_check.py per candidate (7-grams, 12/15-grams, longest run, name-excluded counts, a same-volume null); apply the frozen discard rule; write the declaration and its basis. — evidenced (JA→EN, FR→EN, RO→EN, FA→EN), scripted.
  3. Set the regime parameter: R04 or R05; the R06 arm is emitted by the freeze. — evidenced (JA→EN, FA→EN) for the lead executing it; untested for R03, retrieval variants and cross-model revision.
  4. Render from source only; block published target-language renderings until the freeze; declare priming. — evidenced as lead practice; untested as an enforced context boundary.
  5. Freeze: commit draft + log; commit revision + continued log; cite the commits. — evidenced (JA→EN); a byte-for-byte check of frozen spans against their commit untested (named by R05, never built).
  6. File with the required fields; rebuild the index. — evidenced (the whole filing tree), scripted.
  7. Judge (optional): three blind non-Anthropic seats, order-swapped, paraphrase null, per-juror degeneracy check; label every score provisional. — evidenced (JA→EN, RU→EN, FR→EN), scripted; the null untested as a working floor (failed its bar on one juror S253).
  8. For serial work: commit the collated source; index spans; append-only artifact with errata; maintain register.md; re-run step 2 mid-work against the comparator and the artifact's own earlier spans. — evidenced (FA→EN), executed by the lead.
  9. Compare regimes under control-arm-spec R1–R5 (never pool across length-sign strata; 22–30 items per stratum before any estimate); treat a same-session pair's contrast as an upper bound. — evidenced (JA→EN, RU→EN, DE→EN) as rules applied; R06 vs R04 vs R03 on prose (W2 step 5) — untested, never run.

Human entry points. (a) Step 1's choice of work and declared purpose; (b) step 2's verdict on a borderline gate firing (Villiers — the record chose the frozen rule); (c) step 3's regime and every per-family policy parameter (four of the five entries' step 3s), defaulting to R04 and each entry's stated default; (d) step 5's revision pass, which a human editor could take — the R03 editor role, never specified; (e) step 7, where a human reader could replace or join the panel — never done, and no human-subject data is collected; (f) step 8's unresolved register questions. If nobody enters: what the record shows — every one of 318 filings ran with nobody, on the defaults above; no live protocol for a person's entry was ever written.

Application. Not applied: written at close-out without a translation limb.

6. Not evidenced, and open

7. Sources consumed