Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260824c-run-placement/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260824c-run-placement
statusfrozen
created2026-08-24
updated2026-08-24
sensesstyle-correspondence
provisionaltrue
internal-judgment-onlytrue
linkswiki/arms/ARM-run-depth.md, workshop/translations/gulistan-bab1e/R47-v1/translation.md, workshop/regimes/R47-three-placements.md, workshop/experiments/E-20260824c-run-placement/critic-response.md, wiki/findings/results/RS-20260823-run-depth.md, wiki/findings/results/RS-20260822b-echo-threshold.md, wiki/findings/results/RS-20260824-eye-or-ear.md, wiki/method-notes.md, config/models.md, tools/rhyme_pairs.py

E-20260824c — the same chime at two distances: ARM-run-depth step 2, rebuilt as an indirect measure

Design v2, frozen before dispatch. v1 was frozen, sent to one adversarial critic round, and rebuilt on its findings; v1 is in git history and every finding is answered in critic-response.md. The translator's log this is built on was frozen and committed at ef22a3df, before v1 existed.

1. The question, and why this is not E-20260823 again

A Persian قطعه holds one rhyme across the ends of consecutive bayts and leaves the hemistichs between them bare: in English, one line to a hemistich, a chime at every second line-end. The published tradition replaces it with couplets. ARM-run-depth step 1 (RS-20260823-run-depth) established that the source's placement can be carried, that carrying it costs repetition rather than invention, and that no published hand has ever held one (0 of 12 cells). It could not establish whether the placement is registered by a reader: the task it bought asked three seats to name the rhyme scheme, and they named it at 1.000 in both conditions. Method note (brb): a direct identification task on model seats is a ceiling, not a measurement, and the remedy it prescribes is RS-20260822b's shape — never ask the seat to identify a chime; ask which of two variants reads better, and infer registration from the difference.

This design is that rebuild. The seat is never asked about rhyme, sound, endings, or form. It is asked which of two passages reads better. The passages differ in one word.

2. The design

For each locus, one four-line English passage in five placements, and within each placement two variants differing only in the last word of line 4:

placement line 2's end line 3's end chime
B0 as written as written none
D2 rewritten to chime with w+ as written at line-ends 2 and 4 — the Persian's placement
D2c rewritten, not chiming as written none — the control for D2
D1 as written rewritten to chime with w+ at line-ends 3 and 4 — adjacent
D1c as written rewritten, not chiming none — the control for D1

Outcome per cell: which variant the seat says reads better. hit = it chose +.

Registered quantities, each the mean over loci of the per-locus rate (equal weight per locus), pooled over both orders and the two reporting seats P1 and P2:

3. Materials

Seven loci, every one of them four lines of the lead's own English rendering «گلستان» باب اول ۲۶–۳۱, written under R47 at one sitting and frozen at ef22a3df before this design existed. The Persian rhyme structure was frozen earlier still, from the Persian alone (workshop/translations/gulistan-bab1e/passages-frozen.md).

locus Persian passage source depth w+ / w− D2 partner D1 partner
L1 V2, p003–p004 2 go / rise so below
L2 V3, p006–p007 2 disarray / confusion day lay
L3 V4, p009–p010 2 go / pass below so
L4 V6, p017–p018 2 tend / serve depend end
L5 V7 bayts 1–2 4 head / skull bled fed
L6 V7 bayts 3–4 4 told / said hold mould
L7 V10, p029–p030 2 (+مطلع) passed / went last cast

Two of the nine four-line units the span supplies were declared NO-ALTERNATIVE by the regime before this design existed (V5, V9) — the hand could not write two equally faithful final words in a family the neighbouring lines could reach. They supply nothing here and are not replaced.

F5, the mechanical check, is discharged before the run and is a precondition of it. All 70 variants were graded over all six line-end pairs with tools/rhyme_pairs.py: every + variant of D1 and D2 carries exactly its intended STRICT pair and no other relation of any kind; every + variant of B0, D1c and D2c carries none; and no − variant carries any relation at all. Recomputed by verify.py.

One instrument limit, found by the critic and not by the tool. v1's L6 chimed the written decree is read with said. tools/rhyme_pairs.py graded it STRICT because CMUdict carries both /rɛd/ and /riːd/ and the rule relates two bearers if any pair of pronunciations relates — but a reader reading a passive present verb does not hear /rɛd/. L6 was rebuilt on hold · mould · told · said, which is unambiguous. A mechanical grader that accepts any dictionary pronunciation can certify a chime a reader will not get**, and that goes on the result page.

Contamination. Measured after the log froze and before v1 was written: tools/dependence_check.py, the lead's whole record English of the span against Eastwick 1852 — 5 shared 7-grams, 0 twelve-grams, longest common run 8 tokens, verdict clean. The declaration on the artifact stays suspected (a measurement locates a run; it does not license an independence claim). No quantity in this design turns on the lead's independence from any published hand: the carriers are the lead's own English and the manipulation is applied to it.

4. Procedure

Stage P — the primary. Prompt, fixed, naming neither rhyme nor sound nor line-endings:

Below are two versions of the same short passage of verse. They differ in one word. Which version reads better? Answer with a single letter, A or B, and nothing else.

Both orders are run for every cell (+ first and − first). Cells: 7 loci × 5 placements × 2 orders × 3 seats = 210 — of which P1 70, P2 70, QR 70.

Stage A — the faithfulness check on the word pair, descriptive, prunes nothing. For each locus: the Persian bayts, the B0 English quatrain with the final word blanked, and three candidates — w+, w−, and a planted word that introduces something the Persian has not got (scream, flames, sell, shear, coffin, bought, rained) — in an order randomised per (locus, seat) with seed 20260824, the map recorded per cell. The seat returns, per candidate, one of ADDS / LOSES / FINE. Cells: 7 loci × 2 seats (P1, P2) = 14.

Seats (config/models.md; slugs are provenance, roles are configuration):

role slug use
P1 openai/gpt-5.6-terra reporting
P2 google/gemini-3.6-flash reporting
QR qwen/qwen3.7-max declared item-set sensitivity check, 70 stage-P calls, reported under its own heading, never pooled with P1/P2, no bar attached

QR's role is exactly what RS-20260824 §4 used it for: if the reporting seats show nothing and QR shows something, the null is a fact about those seats and not about the materials. It is outcome-conditioned selection, declared before dispatch.

Temperature 0, one call per cell, max_tokens 1500 on stage P and 2500 on stage A (note (bnk): the cap covers hidden reasoning as well as the answer). All calls in one foreground process with a thread pool capped at 6 concurrent (note (brf); S210), so at most six calls are ever in flight; the stop-loss is checked before every dispatch and again between placement batches. Raw responses preserved per cell in run.jsonl, which is only ever appended to.

Critic. One round, P1 and P2, on v1 — done, $0.118432, fifteen findings, all accepted, critic-response.md. No second round (note (bqp)).

5. Predictions, registered

  1. R1 Δ_D1 ≥ +0.30. A chime at adjacent line-ends is registered.
  2. R2 Δ_D2 ≥ +0.15, and Δ_D2 < Δ_D1. The Persian's placement is registered, but less.
  3. R3 hit(B0), hit(D1c) and hit(D2c) each lie in [0.20, 0.80] — no reference cell is at floor or ceiling.

The lead's ground for R2 rather than a flat null: §7.26 measured a chime worth +0.542 at a verse line-end, and RS-20260823 showed the seats can identify a distance-2 rhyme perfectly. What is unknown is whether identification and preference come apart.

6. Failure criteria, registered

7. Analysis, fixed before the run

Per-locus rate → mean over loci → Δ by subtraction. Intervals are percentile bootstrap over loci, 10,000 resamples, seed 20260824, and are dispersion statistics, not inferences (A13): seven hand-built loci are not a sample of English quatrains, and every criterion above is a decision rule about this item set. Per-seat and per-locus figures printed. QR printed separately and never pooled with the reporting seats.

Verification. verify.py recomputes every reported number by a path that imports nothing from analyse.py, re-grades all 70 variants from the raw item file, and runs mutation tests that must each be caught.

8. Budget

Declared ceiling $1.20, of which $0.118432 is already spent on the critic round. Pre-flight for what remains: stage P 210 calls ≈ $0.53; stage A 14 calls ≈ $0.08; the F6 repeat allowance ≈ $0.30. Expected total ≈ $0.75, worst planned ≈ $1.03.

Worst case is built from max_tokens, not from an assumed output length (note (abc)), and now includes input tokens (round-1 finding P1-11): 210 × (400 in + 1500 out) + 14 × (700 in + 2500 out), priced at the dearest seat in each stage, is $1.98 — above the ceiling, so the ceiling is the binding constraint: the dispatcher stops and reports at $1.20 − $0.118432 of run spend, and with concurrency capped at 6 the overshoot cannot exceed six calls. Today's UTC headroom at the time of writing: $1.777 after the critic round.

9. What this cannot show

  1. The seats are models. There are no human readers; NEXT.md has carried independent human readers as named-not-built for weeks. Tier D is NOT PASSED and every figure is provisional and internal-judgment-only.
  2. The construct is preference under explicit comparison, not registration in ordinary reading (round-1 finding P1-10). A side-by-side forced choice between two passages differing in one word makes that word salient and may invite deliberate inspection of the line-ends. Salience is identical across all five placements, so it cannot manufacture a difference between them — and D1 and D2 are matched on it exactly, which is why the extent statement is the quantity least exposed. But no sentence anywhere may say a reader hears anything.
  3. D1 is not the tradition's couplet. Nothing here compares the Persian's placement with what Gladwin, Ross, Eastwick or Arnold actually do — RS-20260823 §6 is where that lives.
  4. Seven loci, one hand, one book, one language pair, and two of the span's nine units were dropped by the regime before the design existed.
  5. The hand wrote all five placements and knew the question. The controls remove the carrier rewrite as an explanation; they do not remove the possibility that the hand's D2 rewrites are systematically limper than its D1 rewrites. That is now the leading alternative explanation for any Δ_D1 − Δ_D2 gap, and it is smaller than what it replaced.
  6. One call per cell at temperature 0, which is not determinism — note (bre); F6 is the registered contingency, not a general remedy.
  7. w+ and w− are not perfectly matched words and cannot be: they are what a translator would actually write. F4 and F7 measure the exposure rather than removing it.