Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260828-purchased-figure/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260828-purchased-figure
statusfrozen
created2026-08-28
updated2026-08-28
sensesaccuracy, style-correspondence
provisionaltrue
trackT1
linkswiki/arms/ARM-gulistan.md, workshop/translations/gulistan-bab2/R05-v1/translation.md, workshop/translations/gulistan-bab2/collation-chapter.md, wiki/base/anchors/A-gulistan-hands/README.md, wiki/findings/results/RS-20260822c-persian-hands.md, config/models.md, framework/v0.2/README.md

E-20260828 — where a rhyme gets paid for

v2, frozen after the pre-run critic and before any judging call. ARM-gulistan step 1's study limb. Translation limb: T-gulistan-bab2-R05-v1 span B, frozen and committed at b4daef0f with its log before v1 of this design was written and before any English of the chapter was opened (charter A4). The critic's findings and their disposition are §10; v1 is in the git history at b4daef0f's successor commit and is not reproduced here.

1. The wire

The translating produced the question. Rendering 39 of Sa'di's bayts under a rule that forbids buying a figure with a word the source has not got (V2), the translator refused six rhymes, and every one of the six was refusable for the same reason: the English rhyme was one small added word away, and the word would have stood at the end of the line (D24) — astray, at all, slain, door, gown, hall. Nothing was refused for want of a rhyme; the rhymes were there.

That is a claim about where the cost of a formal constraint falls, and it is testable on a hand that took the rhymes. Eastwick 1852 rhymes 38 of 42 bayts in this book (RS-20260822c, on باب اول), and the rhyme coding below finds 17 of his 20 sampled line ends chiming with another line end in the same stanza, against 0 of 20 for the unrhymed control.

PR-PRIMARY, in one sentence: in Eastwick's 1852 verse rendering of these ten tales, the word at the end of a line is more often a word with no counterpart in the Persian than a word from the middle of the same line is.

2. What this is not

3. Materials

Persian «گلستان» باب دوم حکایات ۱۱–۲۰, the 39 bayts of blocks b073–b155, from source-ganjoor-bab2-whole.txt (single witness, declared in collation-chapter.md)
EAS Eastwick 1852 (2nd ed. 1880), archive.org gulistanorrosega00sadiuoft, chapter II stories XI–XX — 24 verse blocks, 76 lines, 70 eligible. Blocks found mechanically by Eastwick's own printed labels; OCR repair and prose truncation done by hand and declared in materials/verse_blocks.py; raw OCR kept in materials/eas_verse_raw.json
UNR the control the critic required. The same 39 bayts rendered as two lines of unrhymed English verse each by x-ai/grok-4.5 — told not to rhyme and told nothing whatever about fidelity — make_unr.py, 39 calls, $0.150823200, 0 void. Deliberately not one of the three judging seats, so no model judges its own output. 78 lines, 73 eligible
LEAD T-gulistan-bab2-R05-v1, span B — 79 lines, 68 eligible, frozen at b4daef0f. Descriptive only
excluded Arnold 1899. His scan wraps nearly every verse line onto two OCR lines and his italics survive only as stray marks. The line end is this experiment's manipulation, so recovering Arnold's line ends would put lead reconstruction inside the manipulation. Gladwin 1806 and Ross 1823 print Sa'di's verse as prose and have no line ends at all

Rhyme coding, mechanical, from tools/rhyme_pairs.py and its CMU dictionary, computed on the sampled lines before any judging call: a line end is coded by its best relation to any other line end in the same block.

hand STRICT NEAR NONE
EAS 17 0 3
UNR 0 2 18
LEAD 2 2 8

UNR is verifiably unrhymed and EAS is verifiably rhymed; the arm contrast is what it says it is. LEAD's four are D22's and D23's — the rhymes that fell out of the plain sentence.

4. Arms — positions, not texts

Every item is one English word set against the Persian bayt(s) its block renders, and nothing else. The seat cannot see the arm, the hand, the line, or the hypothesis.

arm what the word is n
END the final word of an English verse line 20 EAS · 20 UNR · 12 LEAD
MID the content word nearest the midpoint of the same line — paired, not independently sampled 20 · 20 · 12
DECOY a content word from a different block of the same hand, set against this block's Persian 8 per hand
CTRL-ANCH a known-answer item: an English word that renders a thing the bayt names outright (camel for بُختی, stone for سنگ, muezzin for مؤذّن …) 12

Selection is mechanical and seeded (build_items.py, seed 20260828). A line is eligible only if its final token and at least one interior token pass the same content filter, and an eligible line contributes one END and one MID item — so the two arms are drawn from an identical line set and cannot differ in lexical class by construction. 140 items, one shuffled order.

5. Procedure

Three seats, per config/models.md — P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, QR qwen/qwen3.7-max (P3 is the UNR translator and does not judge; P4 and P5 are out on notes (bps) and (bne)). One call per item per seat, no batching, temperature 0. Each seat is asked, for one Persian bayt and one English word:

Does the English word render something that is actually present in this Persian — a word, an image, an action, a thing named, a quality, a relation? … If you cannot read the Persian well enough to say, answer UNSURE — that is a real answer and it is better than a guess.

Item verdict = majority of the three seats; a three-way split or any tie is UNSURE.

Cap probe first, note (brt): the six longest prompts, three seats, 18 calls. Acceptance rule, registered (critic A7): all 18 bodies must parse and none may carry finish_reason == "length". If either fails, the cap is doubled and the probe re-run once; a second failure stops the run before the main block.

6. Registered primaries and predictions

Both two-sided, alpha 0.05, Holm across the two (critic A4). Significance by exact permutation over the arm labels within hand, 20,000 relabelings, seed 20260828, reported with the risk difference and its 95% interval. The test is associational, not causal (§2).

Reported, not registered as a primary: the LEAD gap, and the overall ADDED rate by hand. The LEAD gap is declared partly entailed — V2 forbade the lead to supply an unanchored word anywhere, so a small gap there is policy, not evidence (critic A3). It is reported because it is the wire back to the translation limb, and it is interpreted only as a check that the policy was kept.

7. Gates — what withholds the primaries

A gate that fires is the result. No primary is reported past a fired gate, and no number from a withheld primary may be cited anywhere in this repository.

8. Budget

Ceiling raised from $1.60 to $2.40, and the reason is the critic: two BLOCKING findings required a third translation arm (UNR, $0.150823200 spent) and forty more judging items. Spent so far: critic $0.093254750 (C2 re-dispatched once at a larger cap after its first body truncated), UNR $0.150823200. Remaining plan: 18 probe calls, then 140 × 3 = 420 judging calls. Worst case built from the cap actually set and not from an assumed output length, note (abc). Headroom on the day at design time: $3.769373050.

9. Limitations, written before the run

  1. Not causal, and one published hand. PR1 is a statement about Eastwick 1852's rendering of باب دوم حکایات ۱۱–۲۰ and nothing wider. Arnold's exclusion (§3) is what costs the generalisation; span C can pay it back from page images.
  2. Alignment is at block level, not hemistich level. A word absent from its own hemistich but present elsewhere in the same block codes ANCHORED. This makes ADDED a stronger claim and depresses all arms equally.
  3. The UNR control is a model, not a person. It is the right control for line-end convention without rhyme, and it is not evidence about what human unrhymed verse translators do.
  4. The design is not disinterested. The lead produced two of the three arms and the hypothesis.
  5. Single-witness copy-text, inherited and declared.
  6. G1 is 12 items. A seat that reads Persian badly but is right about camel and stone will pass it. It is a floor, not a certificate.

10. The pre-run critic, and what was done with it

Two seats, both adversarial, both given the frozen v1 design, the item builder verbatim, eight items as the seats would see them, and D24. C1 openai/gpt-5.6-terra returned NEEDS-REDESIGN, 7 findings, 2 BLOCKING. C2 google/gemini-3.6-flash returned NEEDS-REDESIGN, 5 findings, 2 BLOCKING. Raw bodies: raw/critic.json. Eleven of the twelve findings are accepted; one is overruled.

# seat severity finding disposition
A1 C1 BLOCKING END is not a verified rhyme position, and line-final words differ from interior words in closure, stress and information structure — MID controls none of it accepted. Line ends are now rhyme-coded mechanically (§3) and a third arm, UNR, supplies unrhymed English verse line ends (§3, PR2)
A2 C1 BLOCKING PR2 confounded by hand and material; END and MID independently sampled rather than paired accepted. Arms are now paired within line, from an identical eligible-line set
A3 C2 BLOCKING ×2 The LEAD gap is algebraically near-entailed by V2, so the old PR2 restated PR1 accepted. LEAD is struck from the primaries and reported as declared-entailed; UNR takes its place
A4 C1 MAJOR No alpha, no multiplicity, and permutation is associational not causal accepted. Alpha 0.05, Holm, risk differences with intervals, and §2's causal disclaimer
A5 C1 MAJOR G2 (item-level UNSURE) cannot catch seats that are confidently wrong accepted. CTRL-ANCH, twelve known-answer items, is now G1
A6 C2 BLOCKING The decoy exclusion used substring matching, so short words were rejected wherever their letters occurred inside longer ones accepted. Token-set membership
A7 C1 MAJOR The cap probe has no acceptance rule and cannot withhold anything accepted. §5
A8 C1 MAJOR Decoys are not ADDED by construction, so the old 0.60 threshold has no validated meaning, and it fires only at ≤14/24 accepted. G2 reframed as a ceiling diagnostic, threshold and discreteness stated
A9 C1/C2 MAJOR/MINOR Even if everything comes out as predicted, only a narrow associational claim is licensed accepted. §1's PR-PRIMARY and §2 rewritten to the narrow claim
A10 C2 MAJOR The content filter treated END and MID asymmetrically accepted, by the same change as A2
A11 C2 MAJOR Draw decoys from outside باب دوم, because these ten dervish tales share moral vocabulary OVERRULED. C1's A8 says the opposite about the same set — that decoys already differ too much from END/MID candidates — and an out-of-corpus decoy would differ further in register and topic, which is the defect C1 names. The substance of C2's worry is met by G1, which does the calibration with known answers instead of with decoys, and by G2's lowered, explicitly diagnostic threshold
A12 C1 BLOCKING (the fix) Obtain translations under randomised rhyme-required / rhyme-forbidden instructions from the same translators noted as the successor design, not adopted. It cannot be done to a hand that died in 1883, and UNR is the nearest thing this run can build. Recorded for span C