Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260830b-rhyme-family/design-v2.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260830b-rhyme-family-v2
statusfrozen
created2026-08-30
updated2026-08-30
sensesaccuracy, style-correspondence
linkswiki/arms/ARM-rhyme-family.md, workshop/experiments/E-20260830b-rhyme-family/design.md, workshop/experiments/E-20260830b-rhyme-family/critic-response.md, workshop/regimes/R57-leaf-contract.md, workshop/translations/hafez-bekonad/R57-v1/translation.md, wiki/base/anchors/A-leaf-hafez/README.md, wiki/findings/results/RS-20260830-leaf-contract.md, runs/RS-20260829-radif-hands/census.json, config/models.md

E-20260830b design v2 — can one English rhyme hold the senses a ghazal puts at its rhyme?

This file supersedes design.md and is what was dispatched. v1 was frozen, sent to two independent adversarial critics, and both returned NEEDS-REDESIGN — 27 findings, 18 of them BLOCKING, converging independently on the same five structural faults. Eleven amendments were accepted and three remedies overruled with written reasons: critic-response.md. Frozen before any stage-F or stage-S call.

1. Where the question comes from

T-hafez-bekonad-R57-v1 (S233) rendered Hafez غزل ۱۸۷ whole under Walter Leaf's printed 1898 contract. Its translator's log, frozen before any evaluation existed, registered one prediction that has never been tested:

the constraint that actually bites in a monorhymed English ghazal is not the metre and not the inversion but whether the source's rhyme-bearing senses happen to lie inside a single English rhyme.

The log's §D2 records that -ending happened to hold an English word for almost every sense that poem placed at its rhyme, and §D4 that the poem's thirteen losses were content words squeezed by the metre, not one of them caused by the rhyme. Both are one hand's observations on one poem.

2. The subject-rule sentence (wiki/tracks.md §The subject rule)

What this unit teaches about translating literature: whether the senses a source poem places at its rhyme can be covered by a single English rhyme — something a translator can find out from the source before writing a line — predicts how much of that rhyme a published translator delivers, and whether it predicts which poems he agreed to translate at all.

3. Constructs, named honestly

For a ghazal with N rhyming positions the Persian places N senses at the line-end.

What A is not. A is measured by asking competent readers of both languages to find such a rhyme. It is therefore what a competent reader can find, which is the translator's real situation — not a property of the English dictionary and not a census of English (critic C1-1, overruled remedy O2). The result page may not call it either.

Q1 — does A, computed from the Persian alone, predict R?

Q2 — does A predict which ghazals Leaf agreed to translate? (critic C2-2, accepted as A6.)

v1 said a positive Q1 would show "the difficulty is in the language pair and not in the translator". That sentence is struck: one translator cannot support it (C1-4). Whatever Q1 returns is a statement about Leaf's twenty-eight odes.

4. Materials

Persian. Ganjoor's Qazvini–Ghani Divan of Hafez, fetched 2026-08-30, raw JSON preserved at runs/RS-20260830b-rhyme-family/ganjoor/. Public domain. Rhyming positions, radif and qāfiya by the frozen census rule of runs/RS-20260829-radif-hands/census.py, imported verbatim into build.py: positions = hemistich 1 plus every second hemistich; radif = longest common trailing word sequence over the positions, minus one word reserved for the qāfiya.

English. Walter Leaf, Versions from Hafiz: An Essay in Persian Metre (1898), public domain, Internet Archive OCR; the same 28 odes and the same analyse.bearers extraction as RS-20260830-leaf-contract, imported unchanged.

The Leaf↔Ganjoor match is runs/RS-20260829-radif-hands/leaf_match.json (S231).

Unselected comparison pool for Q2, drawn before any coverage number exists. From the 495-ghazal census: radif-bearing, 7–9 bayts, neither Leaf's 28 nor already rendered by this project (sh124, sh118, sh4, sh131, sh187) — 215 ghazals — drawn at even index intervals by build.py, in this frozen order:

sh2, sh30, sh45, sh60, sh76, sh105, sh123, sh147, sh165, sh176, sh191, sh209, sh221, sh237, sh263, sh275, sh339, sh357, sh377, sh402

Any budget-forced truncation takes a prefix of that order and never a selection by A (C2-12).

5. Strata

The primary is all 28 odes. v1's PRIM cut on |N − M| ≤ 1, where M is Leaf's own position count — a property of the translation, reflecting his omissions (C1-5, C2-5). It survives only as a declared exploratory sensitivity analysis on the 19 odes IV, V, VI, VII, VIII, IX, X, XI, XII, XIII, XV, XIX, XX, XXII, XXIII, XXIV, XXV, XXVI, XXVII, and the result page must label it so.

6. Set measurement, and what it does and does not mean

Fifteen of the 28 odes have N ≠ M, so no honest position-by-position pairing exists. Both stages therefore work on sets: a sense displaced from one couplet to another is not counted as a loss. The construct being measured is coverage of the source's rhyme senses by the translator's monorhyme, not positional carriage, and §11 forbids the result page from sliding between the two (C2-9).

7. The bought stages — rater-disjoint by construction

The single fault both critics named first is that v1 drew A and R from the same seats (C1-2, C2-1). The two sides of the correlation now come from disjoint seats at disjoint labs.

stage seats lab
F (produces A) P1 openai/gpt-5.6-terra, P3 x-ai/grok-4.5 OpenAI, xAI
S (produces R) P2 google/gemini-3.6-flash, PR qwen/qwen3.7-max Google, Alibaba

PR is the first reserve in config/models.md, promoted for this run for rater-disjointness and for no other reason; that is provenance, not a change of panel composition. P4 and P5 are out under notes (bps) and (bne). temperature 0, max_tokens 2500, dispatch order shuffled against printed seed 20260830. No seat is told the hypothesis, the translator, the poet, the century, or that another stage exists.

7.1 Stage F1 — senses only, rhyme never mentioned

One call per ghazal per stage-F seat. The seat sees the ghazal's rhyming hemistichs in Persian script with the qāfiya of each printed after it, and the radif if there is one. It gives the contextual sense of each qāfiya in one English word or short phrase, or answers UNGLOSSABLE where the marked unit is not a meaningful word on its own (C1-7, C2-11). The word "rhyme" appears only in the neutral description of what a qāfiya is; nothing about English, verse or translation appears at all.

7.2 Stage F2 — the rhyme, on frozen senses

A second call to the same seat, carrying its own F1 sense list verbatim and the instruction that the senses are fixed and may not be revised. It proposes one English rhyme — naming the rhyme and listing its words — and says for each sense whether that rhyme supplies a word carrying it, naming the word when it does. Sealing F1 from F2 is what stops a seat bending a gloss toward something that rhymes (C1-1, C2-7).

Mechanical validity screen. Every named word is looked up in CMUdict via tools/rhyme_pairs.py, and rhyme is judged by RS-20260830-leaf-contract's analyse.rime_nor, which carries Leaf's own non-rhotic licence. A named word that does not rhyme with the modal rime of that seat's own named set is recoded NO. A word absent from CMUdict is counted UNSCREENABLE, and every screened figure is reported twice — once counting those NO, once dropping them (C1-6, C2-8). The screen can only lower A.

A(ghazal) = mean of the two stage-F seats' screened coverage. Declared robustness estimator, reported beside it and never substituted for it: the max of the two seats.

Stage F runs on 48 ghazals: Leaf's 28 and the 20 unselected.

7.3 Stage S — realised coverage

One call per item per stage-S seat. The seat sees the N Persian rhyming hemistichs with the qāfiya marked, and separately the M line-endings of an anonymous English verse rendering. For each Persian qāfiya it answers YES (naming the English word), NO, or UNGLOSSABLE. The prompt tells it to be strict and that a word merely in the same area of meaning is NO.

The screen that makes R monorhyme carriage rather than set cover (C1-3, C2-9). A YES is kept only if the named English word (i) is one of that rendering's actual rhyme-bearing words and (ii) rhymes, under rime_nor, with the modal rime of that rendering's own bearer set. Otherwise it is recoded NO. R_rhyme is the primary; R_set, unscreened, is reported beside it.

UNGLOSSABLE positions, wherever they arise, are dropped from numerator and denominator of both A and R identically.

7.4 Controls

The eight control pairings are printed by stages.py --print-controls into the run record before dispatch (C1-14).

8. The translation limb — an illustration, and it is not evidence

Both critics established that a limb the lead selects, predicts and executes cannot falsify anything (C1-10, C1-11, C2-3). v1's registered P4 is withdrawn.

What remains is worth doing and is honest about what it is: the study limb's measure selects a poem, and translating that poem generates an enumeration — bayt by bayt, in a frozen log — of what a scarce rhyme actually costs, which no correlation can show. That is the wire, and it is generates, not tests (continue-prompt.md §4).

Selection. Stage F runs on the twenty unselected ghazals; the poem translated is the one with the lowest A. The selection script prints the chosen slug and nothing else — no coverage values, no ranking, no margin. Every stage-F body for the chosen ghazal stays sealed until the translator's log is frozen. No published English of it is opened before the freeze; payne_odes.json is in this repository and is not consulted for the chosen slug at any point.

Regime: R57, the FREE arm — monorhyme at every rhyming position, the radif mandatory, one English couplet per Persian bayt, the metre released. The same arm of the same regime rendered غزل ۱۸۷, which makes the two poems comparable in everything but A.

Its blind stage-S score is reported descriptively, is internal-judgment-only as craft evidence, and is never support for Q1.

9. Predictions, gates, failure criteria — registered

P1, primary, n = 28. Spearman ρ between A (stage-F seats) and R_rhyme (stage-S seats) is positive. Monte Carlo permutation test — random.Random(20260830), 100,000 relabelings, one-sided, (1 + #{ρ* ≥ ρ}) / (1 + 100000), ties by midrank — at α = 0.05 (C1-13).

Power, registered before the run (C1-12, C2-6): n = 28 one-sided at α = 0.05 has ~80% power at ρ = 0.45, which is the minimum effect of interest. A non-significant P1 may not be reported as evidence against a moderate effect, and the result page must say so.

P2, secondary — Q2. Leaf's 28 have higher A than the 20 unselected ghazals. Mann–Whitney U, two-sided, exact where feasible.

P3, mechanical, no API cost. A correlates positively with the per-ode rhyme-pass rate adjudicated blind at S233 and negatively with Leaf's per-ode rhyme-word repetition rate. Labelled convergent evidence for general rhymeability, not for sense coverage (C1-15).

Withholding criteria — any one withholds P1, and the run says which.

What would refute the conjecture. P1 returning ρ ≤ 0 with W1–W5 clear. The translation limb takes no part in this rule (C1-11).

10. Spend

Declared ceiling $2.60; $0.112091900 already spent on the two critic seats. UTC-day headroom at session start $3.579539500 (S233 spent $1.420460500 of $5.00).

Planned calls: F1 48 × 2 = 96, F2 48 × 2 = 96, S 38 items × 2 = 76 — 268. Worst case from max_tokens alone, per note (abc), is 268 × 2500 output at the dearest seat rate ($6.00/M) = $4.02, above the ceiling, so the run is probed per note (bsh): the first 8 calls are dispatched alone, billed cost read from usage.include, and the rest goes only if the extrapolated total sits under the ceiling. If it does not, the candidate pool truncates to a prefix of the frozen draw order — twelve, then eight — and never by A.

The translation limb, the census, the Ganjoor fetch, all extraction, all screening and the contamination measurement are the lead's own and are not ledgered (charter §3, A4).

11. What this run cannot do