Repository path: framework/v0.3/entries/HB-pipeline.md · rendered 2026-09-09
Page metadata (front matter)
The pipeline: the steps a translation went through here, which regime each step is, and where a person can enter
Standing. Written at the project's close (2026-09-09, S257) as a consolidation of the record
without the translation limb the procedure's step 6 requires — no fresh passage was translated under
this entry, so it stays status: draft and its §5 Application reads "not applied". The evidence is
X2 (token overlap, word counts and filing tallies recomputed from stored texts) plus the project's own
frozen regime specifications and translator's logs (internal-judgment-only). The two regime
comparisons that scored anything (RS-20260724-selfrevise-first, RS-20260802c-regime-scoring) are
X3, panel-scored; Tier D is NOT PASSED and now EXHAUSTED (config/models.md,
RS-20260906-tierD-verdict-v3), so nothing here says one regime's output is better than another's on
a jury's word — what those runs contribute is instrument findings and a provisional direction,
labelled as such. No v0.2 section belongs to this family (its 52 numbered sections contain none on regimes
or process); the material is the regime pages, v0.1 §1 and §5, and the result pages in §7.
1. The problem
A regime is "a fully specified way of producing a translation" and comparing regimes is "the
workshop's basic experiment" (PROJECT.md §3); the framework must "treat human involvement as a
designed, parameterized dimension — specifying where human judgment can enter and what its presence
or absence changes" (§1). A translator meets this as a sequence of decisions made before the first
sentence: which work, checked against what, drafted how, revised or not, logged when, frozen when,
judged by whom. A pipeline meets it as the same sequence with gates. The record ran one such sequence
318 times (318 filed translation.md files, every one translated-by: lead, by id, 96 R04, 92 R06, 8 R05) and compared regimes twice by jury and three times by machine. What follows is that sequence as executed.
2. What published translators do
No anchor here describes a published translator's working procedure, so this family has no hands
table of the usual kind. Published translators enter the pipeline at one step only — as
comparators at the contamination gate — and the table is what that gate measured against them
(tools/dependence_check.py: shared 7-grams are context, 12- and 15-grams the signal, plus the
longest common run in tokens and name-excluded counts).
| pair · lead text | comparator | longest run | shared 12-grams | null floor | verdict · source |
|---|---|---|---|---|---|
JA→EN · T-kusamakura-vii-bath R06/R04 (1,210/1,206 words) |
Takahashi 1927, the only reachable English | 8 (a doubled proper name) | 0 (4 shared 7-grams) | 5 (Morri 1918) | suspected — bounds contamination from one side only · RS-20260727c-arm-identifiability §8 |
| FR→EN · Maupassant «Menuet», gate translation | anon. 1903 (PG #3086) | 14 (ordinary prose) | 3 | 3 / 3 / 2 | DISCARDED · RS-20260802c §9 |
| FR→EN · Villiers «La Torture par l'espérance», gate translation | anon. (PG #29704) | 12, nine of them a proper name and a title | 1 | 3 / 2 / 2 | DISCARDED — "the rule was frozen and had already fired once" · RS-20260802c §9 |
| FR→EN · Arène «La Mort de Pan» | none found (Gutenberg, Internet Archive) | — | — | — | admitted on an absence, "weaker than a measurement" · RS-20260802c §9 |
| RO→EN · Slavici «Popa Tanda», selection probe | Byng 1921 (PG #38991) | 12 (ordinary prose) | 1 (10 shared 7-grams) | 3 | would be suspected; work rejected on length · verter/plan.md §2 |
| SR→EN · Lazarević «Вертер» (14,293 words) | none on four documented routes | — | — | — | none, not measured · verter/plan.md §2 |
FA→EN · Gulistan span F (R05) |
Eastwick | 10 (14 shared 7-grams) | — | — | clean against the published hand; DEPENDENT? against the lead's own R52-v1: 536 shared 7-grams, 153 15-grams, a 51-token run → high · ARM-gulistan |
The regularities, as what the measurements did: (1) the gate's threshold sits on a scale calibrated
once ("0 Ovid → 21 Turgenev", RS-20260727c §8) and discards on a run of ordinary prose above a
same-volume null floor, not on fame (Villiers, table); (2) the lead's largest measured overlaps are with
itself — 37 contiguous tokens across sessions (RS-20260801e), 51 when a span was re-rendered
minutes after reading its own predecessor (notes (bhb), (bta)); (3) with one reachable comparator a
low run bounds contamination from one side and the declaration stays suspected; with none, the
honest declaration is none, not measured with the routes recorded, which "does not license
treating the work as an independence measurement for anything" (verter/plan.md §2).
3. What this project's own practice found
The pipeline as run, step by step. Every step below was executed by the lead agent in session
following continue-prompt.md §5 and CLAUDE.md §Contamination, except where a script is named.
- Selection gate. Public domain, verified on the manifest before entry
(
workshop/canon/README.md); from S171, confirm the published comparator exists, in reach, before choosing (note (bmw)); from S244, Tom's long-work criteria (post-1880, rarely translated, 8,000–15,000 words, both sides free "where possible"). Last executed S256: four candidates fetch-verified, two measured, «Вертер» chosen with the missing comparator documented (verter/plan.md). - Contamination measurement, before any locus is chosen — one
dependence_check.pyrun per candidate unit, longest run and shared 7-gram count, because "run length alone has failed as a proxy three times" (CLAUDE.md). Over the whole filing tree: 121none, 138suspected, 59highof 318 declarations (mechanical tally, S257).R05's Known limitations ask the arm to re-measure at least once mid-work;ARM-gulistandid after four spans, and the last re-measure moved a declaration tohigh(§2, row 7). - Regime choice.
R04(lead close translation, source-only, draft then self-revise, frozen v1.0 S041) is the default;R05(serial long work) for anything spanning sessions;R06(lead single pass) is a by-product, not a separate choice:R04§Procedure 2a freezes the draft as its own artifact, so everyR04run contains anR06output. The API regimesR01andR02(frozen v1.0, S010) ran once, on two Japanese works × two panel translators × three reps, inworkshop/experiments/E-20260724-r01r02-selfrevise/runs/— 0 of 318 filed translations were produced by an API regime. - Translate from the source alone.
R04§1: source and non-translational apparatus only; "it does not read any published rendering of the same passage into any target language before the translation and its log are frozen"; priming, if it occurred, is declared on the artifact. The S256 probe chose a locus whose English had not yet been seen rather than declaring priming after the fact (verter/plan.md§2). - Freeze the draft, then self-revise. "The commit is the freeze" (
R06§4). Executed for the first lead pair atb2a3199(RS-20260727c§8). What the pass did, measured: a model self-revision changes a mean 0.032 of words and came back shorter in 10 of 12 API pairs (RS-20260724finding 4;RS-20260727c); the lead's own revisions moved a few words in either direction — −4 on Kusamakura, longer in 4 of 5 at S089 (RS-20260802c), +14 on Kleist (RS-20260728c); a revision of its own draft shares a 51-token run with it (RS-20260801e§2, the positive control) — the two arms of a pair are near-identical by construction. - Translator's log, frozen before any evaluation is designed — "decisions and alternatives, not
quality claims" (
R04); evaluative sentences carryinternal-judgment-only; "a log written after seeing scores is worthless, which is what the freeze protects" (workshop/translations/README.md). The log is self-report: "decisions made without noticing them do not appear" (R04). Tallies over a log are computed by a committed script, never read as prose (note (bey)). - Filing under
workshop/translations/<work>/<regime>-v<N>/translation.mdwithtype: translation,translated-by:,contamination:and its basis, lead translations have norun/(cost$0, never ledgered);tools/build_index.pyregenerates the index. - Optional blind panel judging, provisional permanently. Three non-Anthropic jurors, authorship
stripped, order-swapped, a meaning-preserving micro-paraphrase null (since S094). Last executed S253
(
RS-20260907-panel-judging-2, $0.629): the seven items shared with S094 moved a mean 0.127 scale points across 159 sessions against a retest floor of 0.230; one juror (P2) flagged degenerate by a pre-registered check (89.8% of cells at the ceiling). About 206 of about 208 candidate files remain unjudged. - Collation and register, for serial work (
R05,draftv0.1 — never frozen; eight filings). Whole cleaned source committed once, spans named by paragraph index; artifact append-only, revisions only by numbered erratum (Gulistan: five spans, 9,035 English words,D1–D81, errata filed: none); a binding register with anunresolvedsection (gulistan-bab2/register.md, opened at span B, closed at span F at nineteen rules); "decisions made by drift are recorded as such" (R05rule 6).
What the regime comparisons found. Five pages, two scored (X3), three mechanical (X2):
R01vsR02, JA→EN, S010 (RS-20260724-selfrevise-first; 12 paired items, 2 jurors, uncalibrated): a self-revision was preferred over its own draft onnaturalnessat 0.76;accuracy0.43 andstyle-correspondence0.39 were "under-powered (CIs include 0.5) — direction only". Instrument findings outlived the verdict: jurors leaned to whichever text was shown first at 0.63–0.85. Then (RS-20260727c, X2): the revision arm is identifiable without reading at 0.833 — shorter in 10 of 12 pairs — and the temperature control's pooled ≈0.53 was a cancellation of two length-driven halves. Then (RS-20260728c, DE→EN Kleist, X2): length-matching a revision to its draft made a third text, and stratifying by length sign needs 22–30 items per stratum.framework/control-arm-spec.mdR1–R5 (ratified as amended, S083) is what survives.R06vsR04, JA→EN and RU→EN, S089 (RS-20260802c-regime-scoring; five blind lead pairs, 3 jurors, $0.578, uncalibrated): order-averaged deltas of +0.333 to +0.467 on a 1–7 scale on all six senses — within 0.134 of each other; "a revision that lifts every dimension a little and none in particular". Bounds: two jurors at ceiling on 22/60 and 24/60 draft cells; forced preference agreed across orderings on 7 of 15, "what a coin does"; the result moves on one pair. The one prescriptive claim the project ever carried (C12, "one self-revision pass buysnaturalnesswithout movingaccuracy") is inadmissible as X3 and was invoked 0 times in 63 logged decisions / 126 classifications (framework/v0.1/README.md§5,framework/closure.md§2).- Same-session pairs are not independent (
RS-20260801e, NO→EN Kielland, IT→EN d'Annunzio and Tarchetti, X2): on the two pairs off the floor, the second arm sat two to three times closer in shared 7-grams to an independently written first arm than to the arm beside it; the registered carryover predictions Q1–Q2 failed 0 of 3 with the sign reversed, and Q3 (independent arm above the floor mean at 3 of 3) fell short at 2 of 3. Standing rule (workshop/regimes/README.md): any regime contrast measured on a same-session pair is an upper bound on what the regime change buys.
4. The options
R06— single pass, source-only, no return pass. Keeps: the drafter's first reading, at $0; everyR04run supplies one free. Costs: "It has the drafter's blind spots undiluted instead"; "real errors survive. That is the regime, not a defect of the run" (R06).R04— draft, freeze, self-revise. Keeps: a second reading against the source. On the one scored run the panel's deltas were small and even across six senses (X3 — an instrument observation, not a licence); length moved a few words in either direction. Costs: "self-revision shares the drafter's blind spots" (R02,R04); the lift is unstable to one item and unreadable against a working null.R05— serial, cumulative, register-bound. Keeps: decisions made before knowing what was coming, visible because nothing frozen is silently repaired; terminology and voice bound across sessions. Costs: "the freeze is only as good as the discipline"; "the register can become an alibi"; contamination measured once at selection unless re-run.R01/R02— API single pass, API draft-plus-self-revision. Keeps: a scripted step with frozen prompts, reproducible across seats. Costs: per-call money; the reasoning-token truncation trap; a revision arm identifiable by length alone.R03— multi-agent (translator / editor / source-scholar). Reserved from run one (wiki/program.mdSlate D), named as the branch point for cross-model revision and for human-in-the-loop variants, and never specified; retrieval-augmented variants and human-in-the-loop protocols (charter §3) likewise (wiki/reassessment-2026-09-04.md§1 item 6). Fifty-five of the fifty-nine regimes the reassessment counts (58 spec files) are one-off briefs.- Where a person may enter — the parameter charter §1 asks for, read off the five active entries
(§7): four of the five put their policy parameter (step 3 of the pipeline
half) as a
declared-purposedecision routed to a person (HB-telling-the-reader's is which statement to print, its step 1); three name classifier/policy conflicts and flagged sites as a second entry point; every one states a default that runs when nobody enters.
5. Guidance
For a translator
- Confirm a published comparator exists and is reachable before choosing the work; if none is,
decide with that documented, not assumed. —
evidenced (FR→EN, RO→EN, SR→EN). - Measure contamination on the candidate before choosing a locus: longest common run and shared
7-gram count against the comparator and a same-volume null. Discard on a frozen rule; with one
comparator only, declare no lower than
suspected; with none,none, not measuredwith the routes. —evidenced (JA→EN, FR→EN, RO→EN, SR→EN, FA→EN). - Choose
R04by default,R05for anything that will span sittings; write the draft out and commit it before revising. The commit yields anR06output for free; a draft reconstructed afterwards is not one. —evidenced (JA→EN; 96R04and 92R06filings). - Translate from the source and non-translational apparatus only; open published English after
the freeze; declare any priming. —
evidenced (RO→EN selection probe; standing practice). - Log decisions and alternatives as you write, never quality; freeze the log with the commit;
compute any tally by script. —
evidencedas universal filed practice (every filing carries a frozen log; pairs beyondpairs:not separately verified). - Never judge your own output; if it is scored, every score is
provisional, permanently. —evidenced (JA→EN, RU→EN, FR→EN judged S089/S253). - Expect a self-revision to change about 3% of words and to move length by a few words in either
direction; do not expect any measurement to show which sense it buys. The 3% is one JA→EN API
run; the lead's own revisions ran longer in 4 of 5 pairs at S089, +14 on Kleist, −4 on Kusamakura;
the only scored reading, an even small lift, is X3 and licenses nothing. — the length figures
evidenced (JA→EN, RU→EN, DE→EN); the sense clauseuntested. - For serial work: commit the whole cleaned source once; name spans by index; repair a frozen span
only by numbered erratum; keep a register with an
unresolvedsection; re-measure contamination mid-work against the published hand and your own earlier spans. —evidenced (FA→EN). - Do not re-render a span for comparison against a prior rendering you have just read and call it
independent: draft blind and diff afterwards, or declare the dependency in the design. The
self-match reached 51 tokens. —
evidenced (FA→EN), note (bta).
For a pipeline
- Select: verify PD status and comparator reachability from tool results; emit the search log.
—
evidenced (RO→EN, SR→EN), executed by the lead in session. - Gate: run
dependence_check.pyper candidate (7-grams, 12/15-grams, longest run, name-excluded counts, a same-volume null); apply the frozen discard rule; write the declaration and its basis. —evidenced (JA→EN, FR→EN, RO→EN, FA→EN), scripted. - Set the regime parameter:
R04orR05; theR06arm is emitted by the freeze. —evidenced (JA→EN, FA→EN)for the lead executing it;untestedforR03, retrieval variants and cross-model revision. - Render from source only; block published target-language renderings until the freeze; declare
priming. —
evidencedas lead practice;untestedas an enforced context boundary. - Freeze: commit draft + log; commit revision + continued log; cite the commits. —
evidenced (JA→EN); a byte-for-byte check of frozen spans against their commituntested(named byR05, never built). - File with the required fields; rebuild the index. —
evidenced(the whole filing tree), scripted. - Judge (optional): three blind non-Anthropic seats, order-swapped, paraphrase null, per-juror
degeneracy check; label every score
provisional. —evidenced (JA→EN, RU→EN, FR→EN), scripted; the nulluntestedas a working floor (failed its bar on one juror S253). - For serial work: commit the collated source; index spans; append-only artifact with errata;
maintain
register.md; re-run step 2 mid-work against the comparator and the artifact's own earlier spans. —evidenced (FA→EN), executed by the lead. - Compare regimes under
control-arm-specR1–R5 (never pool across length-sign strata; 22–30 items per stratum before any estimate); treat a same-session pair's contrast as an upper bound. —evidenced (JA→EN, RU→EN, DE→EN)as rules applied;R06vsR04vsR03on prose (W2 step 5) —untested, never run.
Human entry points. (a) Step 1's choice of work and declared purpose; (b) step 2's verdict on a
borderline gate firing (Villiers — the record chose the frozen rule); (c) step 3's regime and every
per-family policy parameter (four of the five entries' step 3s), defaulting to R04 and each
entry's stated default; (d) step 5's revision pass, which a human editor could take — the R03
editor role, never specified; (e) step 7, where a human reader could replace or join the panel —
never done, and no human-subject data is collected; (f) step 8's unresolved register questions.
If nobody enters: what the record shows — every one of 318 filings ran with nobody, on the
defaults above; no live protocol for a person's entry was ever written.
Application. Not applied: written at close-out without a translation limb.
6. Not evidenced, and open
R03was never specified and W2 step 5 (R06vsR04vsR03on two prose passages, judged blind) was never run. Cross-model revision, retrieval-augmented variants and a live human-in-the-loop protocol exist only as names in charter §3.- The scored comparisons are JA→EN (12 API pairs) and JA/RU→EN (five lead pairs, one translator).
Nothing is a sample of a pair, a regime in general, or self-revision in general (
RS-20260802c§10). - The freeze is a convention. No tool checks an earlier span byte-for-byte against its freezing
commit;
R05isdraftv0.1, never frozen, so by the regime README's rule its eight filings could be "piloted but not compared for the record". - Contamination is one-sided or unmeasured wherever fewer than two comparators are reachable;
none, not measuredis a documented absence, not cleanliness. - Panel judging has reached 9 of about 208 candidate files (about 206 remain after S253);
P2is degenerate on the instrument. - Open: a mechanical freeze check (
R05§Known limitations). Open: whether a working null can be built at all, after the paraphrase floor failed on one juror (RS-20260802c§3,RS-20260907§4). Open: whether the lead's self-match (37→51 tokens) is a property of re-render briefs only (notes (bhb), (bta)). Open: «Вертер» — six spans and a cadence declared, no word translated (verter/plan.md).
7. Sources consumed
workshop/regimes/README.md(the regime definition; the same-session upper-bound caution),R01,R02(frozen v1.0 S010, one run each),R04(v1.0 S041, step 2a),R05(v0.1draft, never frozen),R06(v1.0 S041, the pairing property).R03: reserved from run one, never specified;R15has no spec page (two pevtsy filings carryR15-crib/R15-englishids).RS-20260724-selfrevise-first(S010),RS-20260802c-regime-scoring(S089) — the row's two pages; pages the row omitted:RS-20260727c-arm-identifiability,RS-20260728c-length-matching,framework/control-arm-spec.md(ratified S083),RS-20260801e-lead-carryover,RS-20260907-panel-judging-2,framework/v0.1/README.md§1 (evidence classes) and §5 (C12),framework/closure.md§2,verter/plan.md(S256),ARM-gulistanandgulistan-bab2/register.md, archived note (bhk).PROJECT.md§1, §3;CLAUDE.md§Contamination;tools/dependence_check.py;wiki/method-notes.md(bhb), (bta), (bmw), (bey);wiki/program.mdSlate D;wiki/plan.md§W2;wiki/reassessment-2026-09-04.md§1 item 6; the five active entries' pipeline halves. v0.2: no section by title or content.- Withdrawn along the way:
RS-20260724finding 1's inference that "lowering temperature alone moves naturalness only to 0.53" isolates the revision act — withdrawn byRS-20260727c§6 (a cancellation of two length-driven halves). C12 — inadmissible (X3) per v0.1 §5 andclosure.md§2, and its shape not reproduced atRS-20260802c§1.RS-20260802cprediction 2 ("noaccuracymovement detectable above the floor") — registered failure, on a floor that measured identity detection (§3).RS-20260727cP3 (the lead's self-revision would lengthen) — falsified, n = 1.RS-20260801eQ1–Q2 (carryover > 0) — failed 0 of 3, sign reversed; Q3 (independent arm above the floor mean) — 2 of 3.RS-20260727c§11's two prescribed repairs (match, or stratify) — neither adopted whole,RS-20260728c.