Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: wiki/plan.md · rendered 2026-09-09

Page metadata (front matter)
typeprogram
idplan
statusactive
created2026-09-04
updated2026-09-09
linksPROJECT.md, wiki/reassessment-2026-09-04.md, framework/v0.3/README.md, NEXT.md, continue-prompt.md, wiki/tracks.md, wiki/arms/ARM-gulistan.md, config/models.md, workshop/translations/gulistan-bab2/register.md

The plan — workstreams, steps, rotation

CLOSED 2026-09-09 (S257). Tom wound the project up; the Routine is off and no session follows. This page is kept as the record of the allocation machinery and of where each workstream ended; wiki/close-out.md is the close-out record. The ledger below is marked done in the sense closed at the wind-up, not finish line reached.

What this page is for. It replaces the six-track rotation of wiki/tracks.md as the page that allocates sessions (2026-09-04, S244, on Tom's direction — wiki/reassessment-2026-09-04.md). A session does not choose a track; it takes the assignment NEXT.md names, which the previous session wrote by the rotation rule below. This page carries each workstream's finish line, its ordered steps, and the ledger tools/check_state.py reads. Cap 24 KB, enforced by the tool.

Why steps and not arms. Ninety-nine arms were constituted between July 26 and September 4; ninety-eight closed inside two sessions, and the last thirty-five formed one chain on one family of problems, each naming the next as its successor. Arms measured shape; nothing measured whether the work was converging on the charter's deliverables. Steps are written against a finish line, in order, and a finished step is progress by definition. Arm pages stay as the record; one is still live (ARM-gulistan) and its cap and budget are still checked.

The deliverables (charter §1, §3, and Tom's answers of 2026-09-04)

  1. A handbook a translator can use and a pipeline can run — framework/v0.3/, organised by translation problem, every entry carrying both the human instruction and the pipeline step.
  2. An evaluation that can say "better" — the jury's Tier D calibration retaken once under a fresh design (authorized by Tom, charter §11), then judged translations and regime comparisons.
  3. Sustained practice in prose — the current Persian and change-of-language thread finished and closed, then a long prose work translated serially, with the handbook applied to it.

W1 — Handbook v0.3 (deliverable 1) — session type H

Finish line. Every family in framework/v0.3/README.md §Index has an entry at status: active written to the template; framework/v0.2/README.md is frozen as the record and no new numbered section is added to it; a pair-coverage table exists; the handbook has been applied end to end to one whole prose work (W3 step 6) with the followability log filed; and a closing changelog states what v0.3 recommends, what it marks untested, and what it withdrew from v0.2.

Steps (one family per session; an unfinished family continues at the next H session):

# family entry state
A change of language inside the source HB-change-of-language done S244 (the exemplar)
G footing, address and displaced marking (R1) HB-footing-and-address done S246
I realia and cultural mediation HB-realia done S249
H register: elevation, placelessness, period HB-register done S252
K what the translator tells the reader: glosses, notes, prefaces HB-telling-the-reader done S255
F mimetics and categories the target lacks HB-mimetics draft, S257 (close-out)
E sound figures and ornament in prose HB-sound-in-prose draft, S257 (close-out)
J emphasis and typography HB-emphasis-and-typography draft, S257 (close-out)
L long works: registers, collation, recensions HB-long-works draft, S257 (close-out)
M evaluating a translation: the senses and the jury HB-evaluating draft, S257 (close-out)
B rhymed prose and rhyme HB-rhymed-prose draft, S257 (close-out)
C radif and refrain HB-radif-and-refrain draft, S257 (close-out)
D meter, line-end and the translator's formal contract HB-meter-and-line draft, S257 (close-out)
P the pipeline: regimes as steps, where a human enters HB-pipeline draft, S257 (close-out)
Z pair-coverage table, changelog, closing statement in README.md done S257

The section-to-family mapping, the entry template, and the consolidation procedure are in framework/v0.3/README.md. An H session's translation limb is the entry's application: a short fresh passage (300–800 words) that presents the entry's problem, translated under the entry's guidance, with a followability log — did each instruction decide anything, and what did it cost. That log is filed with the translation and cited by the entry's §5.

W2 — Jury calibration, then evaluation (deliverable 2) — session type C

Standing, FINAL as of 2026-09-06 (S250): TIER D CALIBRATION IS EXHAUSTED. The one redesign Tom authorized on 2026-09-04 (charter §11) was ratified and run 2026-09-06 (S250) and FAILED, on two independent dispositive gates — the sham broke into the upper branch for the first time in four runs, and detection did not fire at the primary (light) dose. Per the authorization's own stated consequence, no further repair is authorized: config/models.md records the state; full account RS-20260906-tierD-verdict-v3. Steps 1–3 below are therefore CLOSED. Steps 4–5 (judging translations, regime comparison) proceed on provisional-only panel scores, permanently — never pending recalibration.

Steps.

  1. Design session. Import framework/tierD-repaired-rules.md whole (R1–R5 and the required pre-check). Write a fresh design, frozen before any S086 material is re-read, and state in it: (a) the primary dose, declared with a reason that does not cite the S086 cell values — the candidate reason is the jury's actual job, telling near-competent renderings apart, for which the light dose (three sites in ~300 words) is the relevant capability, with the heavy dose kept as a secondary dose-response cell; (b) specificity stated as a ratio to the on-target effect, with the absolute form reported beside it (note (bke)); (c) naturalness scored under its current wording only (D-20260802-13), and every sense on wiki/goodness-senses.md including perceived-source-carriage, which has never been through Tier D; (d) fresh materials — new reference translations of both provenances (published: stored anchors not used at S034 or S086; lead: two translated in this session, contamination measured first — that is the session's translation limb), new damage draws from published catalogues of translation failure only, a rebuilt sham, held-out and positive-control arm, order swap, per-juror robustness, and a verdict table with a row for heavy fails, light passes; (e) the pre-registered firing rule and the exact licence a pass grants (the jury detects S-damage at dose D, never "calibrated"); (f) the worst case from max_tokens, under $3. Send it to two non-Anthropic critic seats; accept or overrule each finding in writing. Open a decision page in wiki/decisions/open/ with the frozen design as the provisional default. Do not dispatch.
  2. Ratify and run — DONE 2026-09-06 (S250). Ratified (independent adversarial review qwen/qwen3.7-max, routed non-Anthropic vote moonshotai/kimi-k3 P4, both RATIFY no amendment), dispatched same day (a fresh UTC day), 96 of 96 calls clean, verified by a dedicated analysis/verify.py (the established per-run practice since S034/S086; tools/verify_tierD.py is the superseded S020-era tool and was not extended, matching what S034 and S086 actually did despite this row's older wording). config/models.md records NOT PASSED, EXHAUSTED.
  3. Consequences — DONE 2026-09-06 (S250). FAIL: recorded as approach exhausted in config/models.md. HB-evaluating (W1 family M) does not get written from a PASS licence — when family M is next worked, its entry states plainly that the jury is not, and will not be, calibrated, and that the handbook's guidance rests on workshop findings alone.
  4. Judge translations. Proceeds on provisional labels only, permanently (the calibration route is closed, not paused): re-judge the A4 set, judge every W3 translation filed since S099 that has a frozen log, and file scores on the translation pages. Three blind non-Anthropic jurors, authorship stripped, order-swapped, meaning-preserving micro-paraphrase null (S094's working null), under $2 per session. Ongoing at cadence, one installment per W2-rotation session (same cadence as step 5) — installment 1 done S253: the five A4 primary items and the null pair re-judged ($0.629, RS-20260907-panel-judging-2) — scores hold stable across the 159-session gap (mean absolute shift 0.127 against a ~0.23 retest floor, 6/6 senses) — plus two new W3 items (Chekhov RU, Schwob FR). ~206 of the ~208 candidate files remain; next installment should add a Japanese item (source text not yet fetched) to keep the three screened languages balanced. A pre-run critic caught a real defect worth carrying forward: P2 (google/gemini-3.6-flash) is now degenerate on this instrument (89.8% ceiling, SD > 0.30 on only 2 of 6 senses) — a pre-registered degeneracy check (juror SD > 0.30 on ≥4 of 6 senses) should ride along on every future installment, not just this one.
  5. Regime comparison on prose — the charter's basic experiment, last run S089. Specify R03 (multi-agent: translator, editor, source scholar) as an executable pipeline; run R06 vs R04 vs R03 on two prose passages; judge blind; write the result into HB-pipeline (family P). Then continue at cadence: one evaluation session per rotation.

W3 — Practice: close the current thread, then prose (deliverable 3) — session type T

Tom's direction (2026-09-04): finish the current thread first, then re-centre on narrative prose. Steps 1–4 are the thread and are hard-capped at one session each; a step that does not finish in its session is closed with what it has, and its remainder goes to the family entry's Open section, not to a new session. The thread closed at step 4, S254.

# step state
1 Gulistan باب دوم span E, حکایات ۴۱–۴۸ (ARM-gulistan step 4; register.md questions, b346) done S245
2 span F: re-render حکایات ۱–۱۰ through the closed register, and the craft report → ARM-gulistan closes resolved done S248 — ES-20260905-craft-report-gulistan; ARM-gulistan closed 5/5
3 «Война и мир» I.i.I–II into French — the one cell §7.52 cannot advise on: Paskévitch 1879 (public domain) as the published hand at the 38 loci; the lead renders the chapters into French as the labelled second hand done S251 — T-voina-i-mir-fr-R04-v1; RS-20260906-french-target-voina-i-mir; HB-change-of-language item 12
4 the RS-20260902-rhyme-slot successor: is the spent slot the line-end or the rhyme? One session on analyse.py's distance-to-chime, then the thread is closed — its findings are W1's families B, C, D done S254 — RS-20260908-chime-slot: the spent slot is the rhyme (qafiya slot), not the printed line's last word; on the 14 of 58 Hafez renderings carrying an English radif, §7.49.1's "factor of eleven" corrects to about 5.5, and its by-radif robustness claim (F3) is now unresolved rather than confirmed. Zero-cost reanalysis of RS-20260902's own frozen raw data; no new API call. Thread closed, per Tom's cap
5 choose the long prose work: public domain; narrative prose written after about 1880; the kind of book rarely translated professionally where one can be found (a minor novelist, a serial, a provincial press) rather than a canon monument; the original and a free published English both reachable where possible; 8,000–15,000 words; contamination measured on two candidate spans before choosing (CLAUDE.md §Contamination); declare the spans and the cadence on a new page workshop/translations/<work>/plan.md done S256 — Laza Lazarević, «Вертер» (1881, serialized four parts in Otadžbina), Serbian → EN, 14,293 words, no free English exists (contamination: none, not measured — comparator absent, four search routes documented). Chosen over Slavici's «Popa Tanda» (has a free 1921 comparator but 6,663 words, under the floor — a real dependence_check.py probe on an unseen middle passage returned a 12-token run against a 3-token null floor) and Vizyinos's «Το αμάρτημα της μητρός μου» (no comparator, borderline length). Six spans declared (1,877–2,931 words each), cadence at least one visit every four sessions. First South Slavic language in the corpus. workshop/translations/verter/plan.md
6 spans on cadence (at least one visit every four sessions) under R05 with a binding register; every span's study limb is a question the span raised, answered into a named handbook entry; as entries exist, apply them and keep the followability log — this is the handbook's end-to-end stress test not begun — project closed S257
7 closing craft report; then one canon Japanese piece retranslated under v0.3 and judged by W2's jury not reached — project closed S257

W4 — Evidence base — session type E — on demand

A published translation read whole against its source and catalogued as an anchor (wiki/base/), only when a handbook entry names the gap ("no published hand measured for this problem in this pair"). Never a standalone session otherwise; the shelf is already 23 anchors and 17 sources.

W5 — Typology — on demand

wiki/goodness-senses.md is used by every evaluation and changes only by ratified decision. A motion arises from W2's judging (step 4) or from an entry that cannot name its sense; it is opened as a decision page and ratified by a later session.

The rotation rule

Ledger (read by tools/check_state.py)

The current session is S257.

id workstream status last worked consecutive next step
W1 Handbook v0.3 done S257 0 — (closed; 5 entries active, 9 draft at the close)
W2 Jury calibration, then evaluation done S253 0 — (closed; Tier D exhausted; steps 4–5 not continued)
W3 Practice: close the thread, then prose done S256 0 — (closed; «Вертер» chosen, not begun)
W4 Evidence base done S243 0 — (closed)
W5 Typology done S219 0 — (closed)

Hand-off: set last worked and consecutive for the workstream worked (reset the others' consecutive to 0), advance next step, bump the current session, tick the step table above.