Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260823c-member-move/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260823c-member-move
statusfrozen
created2026-08-23
updated2026-08-23
sensesstyle-correspondence
provisionaltrue
internal-judgment-onlytrue
linkswiki/arms/ARM-member-move.md, workshop/translations/gulistan-bab1c/loci-frozen.md, workshop/translations/gulistan-bab1c/collation.md, workshop/translations/gulistan-bab1c/loci_saj.json, workshop/translations/gulistan-bab1c/loci_control.json, workshop/regimes/R45-member-move.md, tools/rhyme_pairs.py, config/models.md, config/budget.md, wiki/findings/results/RS-20260822-synonym-reach.md, framework/v0.2/README.md, wiki/goodness-senses.md

E-20260823c-member-move — does moving the rhyme-bearer put a chime in reach, and is it in reach where the author wrote no chime

Frozen 2026-08-23 before any English for the span existed, before R45 was executed, and before any seat was dispatched. The locus inventory (loci-frozen.md) was frozen first, from the Persian alone. ARM-member-move step 1.

1. Question

framework/v0.2 §7.24 item 1, the handbook's only surviving positive instruction about carrying a source's sound figure, says the productive move at a rhymed locus is structural: change which word bears the rhyme. It rests on two instances and no rate.

At a saj' locus, does the English a plain translator already has on the page contain a full rhyme that the ordinary rendering of the source's own rhyme-bearer does not — and does it contain one just as often at the cola the author left plain?

The second half is what decides what the handbook prints. RS-20260822-synonym-reach withdrew §7.24 item 2 on a yield of +0.042 and showed, in the same run, that English's ambient supply of near-chimes is identical at rhymed and at plain loci. If widening the member set behaves the same way, the move is not finding the author's figure.

2. Materials

«گلستان» باب اول, حکایات ۱۴–۱۹ — Ganjoor copy-text, single Persian witness, collation.md. 27 prose blocks, 1,213 Persian prose words, 42 bayts.

The whole enumerated population of loci, no draw (loci-frozen.md):

class loci cola
PROSE-SAJ / STEM — the primary class 29 58
PROSE-SAJ / AFFIX 13 28
CONTROL-PLAIN — the control 10 20
total 52 106

The control selection rule, written out and then applied against itself (critic round 1 finding 3; critic round 2 finding 2). A CONTROL-PLAIN locus is the first pair, in reading order, of و-coordinated cola inside one prose period, each of them a full predication of at least three Persian words, neither of them a colon of a frozen saj' locus, and whose colon-final words do not rhyme; at most one per period.

Four of the fourteen controls originally listed violate that rule and were removed before dispatch — C03 (a member colon is S12's saj' colon, و طاقتِ بار فاقه نمی‌آرم), C04 and C08 (a two-word member, و بیم and و سودمند), C12 (the second eligible pair in p062, where C11 comes first). Critic round 2 found the first of those by name. Ten remain, and seventeen of the twenty-seven prose periods yield no eligible pair at all — Sa'di rhymes his coordinations here more often than he leaves them plain, which is why this control is small and why P2 is declared underpowered below rather than discovered to be so afterwards.

The lead could not select toward an English outcome it had not seen, which is what the whole frozen-inventory design rests on; what the lead could have selected toward is a source-side correlate, above all colon length, because W2 rises with pool size. That is measured rather than asserted: failure criterion 5.

A fully mechanical control enumeration was built, inspected and rejected in writing rather than quietly not tried. Splitting each prose block on the free-standing coordinator و and taking the first adjacent pair of ≥3-token segments yields 23 candidates, of which 11 are the saj' pairs themselves and several of the rest pair fragments that are not cola at all — کردم against است in p003, است against است in p029. A control made of fragments is not comparable in shape to a saj' locus, which trades one confound for a worse one. The enumeration is preserved as materials/control_mechanical.json so that the rejection can be checked.

3. Design in one line

Two blind seats, each shown one colon at a time and never a sibling, supply (a) a plain literal English rendering of the colon and (b) the plainest English of the colon's rhyme-bearing word. Whether any two English expressions chime is then decided mechanically by tools/rhyme_pairs.py under a rule fixed 2026-08-22 and not altered here. The manipulated variable is the size of the pool the chime may be drawn from; the control is whether the source rhymed at that locus at all.

4. Contamination, declared before the translation and measured after the log freeze

contamination: high, declared in advance. Four public-domain English Gulistans — Gladwin 1806, Ross 1823, Eastwick 1852, Arnold 1899 — are certainly in the lead's training data, and this span contains no Gulistan passage as famous as the بنی آدم bayts, but Nushirvan and the salt (حکایت ۱۹) is anthologised everywhere. The declaration is not a placeholder: the measurement runs after the R45 log is frozen, against Eastwick 1852's text of these six tales, obtained this session as full OCR at $0 (collation.md §A published English witness), using tools/dependence_check.py. Longest common run and shared 7-gram count are both reported, per CLAUDE.md.

The study's validity does not turn on the lead's independence, and this is stated so the standing rule is visibly satisfied rather than quietly bypassed. The availability figures — the primary result — are computed entirely from seat glosses adjudicated mechanically; the lead never supplies a word to them. The R45 log is a separate, labelled hand-level limb whose figures are never summed with the seats', exactly as in RS-20260822-synonym-reach. The lead is not the independent third translator of anything here.

5. Procedure

Stage G — the source-side gate. Persian only, no English anywhere, no hypothesis stated. Each of the 52 loci is put to two seats with its rhyme-bearing words marked in «guillemets», and the seats are asked whether those words end in the same sound in Persian. This exists because the STEM/AFFIX/CONTROL assignment is the lead's ear and the lead's ear on Persian is not evidence. 104 calls.

Stage R — the rendering pool. English only, one colon at a time, siblings never co-visible. Each of the 106 cola is put to two seats, in an order shuffled by declared seed 20260823 so no seat meets siblings adjacently, and each returns two lines:

212 calls. No stage mentions rhyme, chiming, sound, echo, translation quality, or that any hypothesis exists. A seat that never sees a sibling colon cannot chime on purpose — the masking RS-20260821c established and RS-20260822 reused.

Seats. P1 = openai/gpt-5.6-terra, P2 = google/gemini-3.6-flash (config/models.md, logged as provenance). P3 is excluded on cost, P4 by note (bps), P5 by note (bne), GL on prompt length. The same two seats sit on both stages; the calls are stateless and this is not a leak, as E-20260822-synonym-reach's runner repair note records.

6. Adjudication — fixed before dispatch, not alterable inside the run

tools/rhyme_pairs.py, built 2026-08-22, rule unchanged. STRICT is the primary; NEAR is reported beside it and never summed with it.

For a locus L with members m_1…m_k and a seat s, define two pools per member i:

Then, per seat-locus:

Identical bearers never relate (rhyme_pairs.py: a repeated word is not a rhyme), which matters here because two literal renderings of parallel cola share vocabulary.

Aggregates are means over seat-loci — 2 seats × the loci of that class — reported per class (STEM, AFFIX, CONTROL) and per seat.

The analysed set is gate-confirmed, not lead-labelled (critic round 1, finding 2). A locus enters the primary only if both gate seats agree with its frozen class: RHYME from both for a PROSE-SAJ locus, NO from both for a CONTROL-PLAIN one. A split verdict, or any UNSURE, leaves the locus unresolved; unresolved loci are excluded from the primary, counted, and listed by id (critic round 1, finding 5 — with two seats there is no majority, so unanimity is the rule and the alternative is an unstated post-hoc treatment of exactly the contested loci). The lead-labelled figures are printed beside the primary, never instead of it.

W2 is a CEILING, and the result must be read as one (critic round 2, finding 1 — accepted, and the finding is right). Pool2 is a bag of the words a literal rendering puts on the page; a word that sits deep inside an English clause may be impossible to move to the rhyming position under REORDER / RESELECT / RESPLIT without unlicensed rewriting. So W2 over-counts the realisable move and can never under-count it, and the asymmetry that follows is registered here rather than discovered later: F1 firing is decisive — if the rhyme is not even in the pool, no move can reach it — while P1 holding licenses only "available", not "reachable", and the reachable rate is what the R45 hand limb and this arm's step 2 are for. Verifying realisability per colon by blind seat, which is what the critic asked for, is a second 212-call stage and does not fit today's headroom; it is refused in writing (critic-response.md) rather than quietly skipped.

Two registered narrowings are printed beside the full-pool figure, both at $0: W2-tail4, restricting Pool2 to the last four content-word tokens of each colon rendering — the words nearest the clause end, hence the ones a reorder can most plausibly reach — and the size-matched recomputation of failure criterion 5.

The primary STEM set is the intersection of the lead's frozen label and a deterministic morphological procedure (critic round 2, finding 3 — the gate confirms that the marked words rhyme and never checks which kind of rhyme). classify.py strips one ending per word from a closed list fixed before dispatch and asks whether a shared final sound survives; it agrees with the hand on 38 of 42 saj' loci. It over-strips lexical look-alikes (سلطان/پاسبان → سلط/پاسب) and so errs only by moving loci out of STEM; it also caught one hand error in the other direction (S29 بردند/کردند, whose -rd survives the -and). Primary: the 27 loci both call STEM. The hand-only (29) and procedure-only (29) figures are printed beside it, and the verdict is reported as holding only if it holds on all three.

Dictionary coverage is checked over every Pool2 token, not only over Pool0 (critic round 1, finding 4). An UNKNOWN token cannot form a mechanically detected rhyme, so an uneven UNKNOWN rate across classes would bias the contrast in the direction of whichever class has more of them; the rate is therefore reported per class and a recomputation excluding every locus containing an UNKNOWN Pool2 token is printed beside the primary.

7. Registered predictions and criteria

Fixed before dispatch. P2 is the one that decides what the handbook prints.

id prediction bar
P1 the reach claim — widening the pool puts a chime in reach at STEM loci W2(STEM) − W0(STEM) ≥ +0.25
F1 the withdrawal criterion — the same bar RS-20260822 used on §7.24 item 2 if W2(STEM) − W0(STEM) < +0.10, §7.24 item 1 is withdrawn by name
P2 the specificity claim — the widening is about the source, not about English. A difference in differences, per critic round 1 finding 1: an absolute W2(STEM) − W2(CONTROL) can clear a bar purely because STEM loci start higher at W0, and can hide a real widening effect behind a baseline gap [W2−W0](STEM) − [W2−W0](CONTROL) ≥ +0.10
P3 the hand claim — a hand under licence beats the blind pool, as it did in RS-20260822 lead CHIMED-MOVED rate at STEM loci ≥ W2(STEM) − W0(STEM)
P4 the price claim — the move is not free among the lead's CHIMED-MOVED loci, the share logged at no price ≤ 0.50

P3 and P4 are computed from the lead's own frozen R45 log and are internal-judgment-only; they measure a hand, and are never summed with the seat figures.

What each outcome makes the handbook say, written now so the writing-up cannot drift:

8. Failure criteria — when the run is reported broken rather than answered

  1. Gate attrition. The primary runs on gate-confirmed loci (§6). If fewer than 12 STEM or fewer than 6 CONTROL loci survive the gate, the run is reported as underpowered for P2 and the difference in differences is printed with its n and without a verdict. P2 is declared underpowered from the start: 10 control loci × 2 seats is 20 seat-loci, which can see a large difference in differences and cannot see a small one, and the result page says so in its headline rather than in its limits.
  2. Non-return. If more than 15% of stage-R cola fail to yield both lines after two attempts, the affected loci are dropped and the drop is reported by class.
  3. Dictionary coverage. The UNKNOWN rate is reported per class over all Pool2 tokens (§6). If it exceeds 10% in any class, the exclusion recomputation is the one quoted in the headline and the raw figure is the sensitivity, not the other way round.
  4. Degenerate pools. If the median Pool2 size exceeds 12 tokens per colon, the widening is reported as having tested "any word in a long paraphrase" rather than "any word of the colon", and a size-capped recomputation at the 8 tokens nearest the colon end is printed beside it.
  5. Class-imbalanced pools. Because W2 rises with pool size, a Pool2 size difference between STEM and CONTROL would produce a difference in differences with no source in the source. Persian colon length and mean Pool2 size are therefore reported per class, and a size-matched recomputation — both classes restricted to loci whose mean Pool2 size lies in the two classes' common support — is printed beside the primary. This is the registered answer to critic round 1 finding 3 and it is a measurement, not an assurance.

9. Budget — pre-flight, built from max_tokens and not from an expected length

Note (abc): the worst case is what the cap permits. max_tokens 1600, reasoning cap 700, per the S213 finding that 500/200 killed 14 of 20 calls and 1600/700 fixed it.

stage calls worst-case
critic, 2 rounds 2 $0.132079 actual
G — 52 loci × 2 seats 104 $0.70
R — 106 cola × 2 seats 212 $1.88
retries (15% at double cap) ~47 $0.59
ceiling declared 316 + retries $3.40

Headroom on UTC day 2026-08-23 at the time of freezing: $4.157546950, of which $0.132079 is already spent on the two critic rounds.

Declared de-scope, registered rather than improvised. If the actual spend after stage G exceeds $0.60, stage R drops the 13 AFFIX loci (26 cola, 52 calls) before dispatch. That touches neither P1, F1 nor P2, whose classes are STEM and CONTROL.

10. Verification

A verify.py that recomputes every number the result page reports from run.jsonl and the frozen loci files, independently of analyse.py, plus mutation checks that must be caught: a flipped locus class, a NEAR counted as STRICT, and a stopword left in a Pool2.