Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: wiki/findings/results/RS-20260902-rhyme-slot.md · rendered 2026-09-09

Page metadata (front matter)
typeresult
idRS-20260902-rhyme-slot
statusfrozen
created2026-09-02
updated2026-09-02
provisionaltrue
sensesaccuracy, style-correspondence
linksworkshop/experiments/E-20260902-rhyme-slot/design.md, workshop/experiments/E-20260902-rhyme-slot/critic-response.md, workshop/experiments/E-20260902-rhyme-slot/amendment-v2.md, workshop/experiments/E-20260902-rhyme-slot/amendment-v2-1-probe.md, wiki/arms/ARM-rhyme-family.md, wiki/findings/results/RS-20260830b-rhyme-family.md, workshop/regimes/R59-rhyme-clause-alone.md, workshop/translations/hafez-boro/R59-MONO-v1/translation.md, workshop/translations/hafez-boro/R59-PLAIN-v1/translation.md, framework/v0.2/README.md, config/budget.md, config/models.md

RS-20260902 — an English rhyme costs the line-end slot by a factor of eleven, and the pre-rhyme slot is not a rhyme effect at all

ARM-rhyme-family step 2 (T2), and the arm closes. E-20260902-rhyme-slot design v1, amendment v2 after two NEEDS-REDESIGN verdicts (19 findings, 5 BLOCKING), amendment v2.1 after the seat probe. 149 API bodies, $2.057401350 of a declared $2.80. Verifier 35 checks, 0 failures, three mutation tests, all three caught. Every figure below is recomputed from the raw bodies by runs/RS-20260902-rhyme-slot/verify.py, which re-derives the line geometry from its own code and imports nothing from the analyser.

1. The question, and what step 1 left open

RS-20260830b-rhyme-family found that Leaf's English rhyme words almost never carry the senses Hafez put at his rhyme — pooled 0.0384, nineteen of twenty-eight odes at zero — while the same senses were found eight times as often somewhere in the last words of the same lines. It had nothing to compare that 0.0384 against, and its §8 asked: is the pre-rhyme slot where the rhyme sense goes in English, in every hand, or only in these two?

Both pre-run critics rejected the first design for asking it the wrong way, and the second question they forced is the one this page answers:

When a translator carries a Persian hemistich-final sense into English, does it land on the line's last word — and how much less often when that English line has to rhyme?

2. What the critics changed, because the amended design is not the registered one

Two seats, disjoint labs, dispatched before any other call: C1 openai/gpt-5.6-terra, C2 x-ai/grok-4.5, $0.104133000. Both NEEDS-REDESIGN. The response is in critic-response.md; three findings governed the rebuild.

C1-10 also restored the 14 English-radif renderings that v1 had excluded, as a stratum; the sample is larger than the frozen design's because of it.

3. Materials

56 published renderings — Walter Leaf 1898 ×28 and John Payne 1901 ×28 — of 54 distinct Hafez ghazals, matched to the Ganjoor text (Leaf at S231, Payne at S236). 747 sense positions: 371 from rhyming hemistichs, 376 from free ones. Stage G glossed each ghazal's hemistich-final words from the Persian alone (P1 openai/gpt-5.6-terra); stage L located each gloss in the whole English rendering, with the poem's glosses in a seeded shuffle, so the seat could not tell which sense came from a rhyming hemistich (P3 x-ai/grok-4.5, low reasoning effort — see §8.1). Disjoint labs. 368 senses located and verified.

4. The primary: no gate fired, and the effect is large

qāfiya sense ordinary hemistich-end sense
landed in a rhyming English line 0.3712 (n=132) 0.0244 (n=41)
landed in an open English line 0.3333 (n=54) 0.2695 (n=141)

Cell entries are the share of located carriers standing on the line's last word.

P1 = 0.2695 − 0.0244 = +0.2451, 95% cluster bootstrap interval [0.1735, 0.3129], resampling whole renderings, 20,000 draws. F1 does not fire.

An ordinary Persian line-final sense reaches the end of an English line 27% of the time when that line is free, and 2.4% of the time when it has to rhyme — a factor of eleven, with the Persian word class held constant. Step 1's 0.0384 now has the thing it lacked: a baseline. The two figures are the same order and were measured on different senses by different instruments; what is new is not the small number but the large one beside it.

Both hands, separately (F2 does not fire): Leaf 0.2115 (0.2115 at n=52 against 0.0000 at n=15); Payne 0.2649 (0.3034 at n=89 against 0.0385 at n=26). Both strata (F3 does not fire): renderings that carry an English radif +0.2571, renderings that do not +0.2413.

Selection is not doing it. The two Persian conditions were located at nearly the same rate — 0.5013 against 0.4840, a gap of 0.017 against W1b's 0.10 — and an adversarial bound that assigns all 194 unlocated free-hemistich senses to the open cell as not-at-end still leaves +0.0890. (That bound replaces the registered ALLITEM, which is not computable: an unlocated sense has no landing line and therefore no cell. Constructed after seeing the data, and it can only weaken the finding.)

5. The registered gates, all of them

gate value fires?
W1 location floor ≥ 0.30 0.5013 / 0.4840 by Persian condition no. The cell-level version is not computable as written — a cell is defined by where a sense landed — and this is a defect in the gate's own wording, stated rather than glossed over
W1b differential location ≤ 0.10 0.0173 no
W2 verification ≤ 20% 2.90% (11 of 379 named words absent from their named line) no
W3 parse ≤ 15% 0, after three stage-G re-buys (§8.2) no
W4 positional bias ≥ 4 of 6 4 of 6, on a re-executed control (§6) no, at the threshold exactly
W5 line length ≤ 2 tokens 11.24 against 11.18 — 0.06 no; no length adjustment was needed
W6 gloss parity ≤ 2 words 2.07 against 1.92 — 0.15 no

6. The second registered prediction FAILS, and that is the more interesting half

(The design calls it P2; it is a prediction, not the panel seat of the same name.)

Step 1 §3(c) — and framework §7.44's own headline — say that a displaced rhyme sense sits one word to the left of the rhyme. That was measured on one poem, the lead's own. Registered here against two comparators, and it clears one and fails the other.

among carriers not on the qāfiya slot share exactly one word left uniform-placement null
in rhyming lines (n=118) 0.1780 0.1045
in open lines (n=139) 0.1727 0.1014

The concentration is real — about 1.7× what uniform placement over the line's remaining positions would give — but it is the same in lines that must rhyme and lines that need not. The prediction required the rhyming share to exceed both comparators and it exceeds only the null.

The penultimate slot is where English puts a displaced sense-word generally. It is not something the rhyme does. The distributions confirm it: capped at six, rhyming lines run {0: 38, 1: 21, 2: 16, 3: 12, 4: 9, 5: 17, ≥6: 43} and open lines {0: 56, 1: 24, 2: 20, 3: 14, 4: 13, 5: 9, ≥6: 59} — the same shape.

So §7.44's headline is half right and half wrong, and the correction is written into §7.49. The first clause — the sense is not at the rhyme — is confirmed and strengthened on two hands and 964 positions. The second clause — it sits one word to its left — describes English, not rhyme, and was generalised from a single poem.

7. Three controls, and one of them was void before it was repaired

X-POS, the cross-poem control (bought on C2-8, because step 1 §6's mismatch prior covers presence and not position). Four items pairing one ghazal's glosses with another's English, 63 sense-items, 11 located — 0.175 against the true items' 0.492. The instrument is not simply naming whatever content word is available: put the senses in the wrong poem and it mostly answers NONE.

POSBIAS, the positional positive control — the first execution was VOID, not failed. Three of its six items marked a bare punctuation token, because Leaf's OCR spaces ; and ? off as separate words and the marker took the line's last token; the paraphrasing seat duly glossed the semicolon ("a pause between clauses"). Scored as it stood it returns 1 of 6 and would have withheld the whole run on a bug of the project's own making. It was rebuilt — last token containing a letter, not a function word, and occurring exactly once in the rendering — and re-executed: 4 of 6, at W4's threshold exactly. The two misses are real: the seat named unite for wed four words from the end, and returned NONE on a poor paraphrase ("are spoken" for flow). Both executions are on the record and the repair is declared; a control repaired after it ran is weaker than one that was right the first time, and this page does not pretend otherwise.

MIGRATION, and it vindicates the critics. Across Payne, whose renderings are bayt-for-bayt by construction, 57.1% of located carriers stand outside their own bayt's two English lines. Any design assuming hemistich-to-line alignment would have been wrong more often than right — which is exactly what C1-2 and C2-1 said, and why the amended predictor is a property of the landing line instead.

8. Instrument findings, both of which cost money

8.1 A larger cap made the answer worse — method note (bst)

On the same item, same prompt, same temperature, P2 google/gemini-3.6-flash located 8 of 14 senses at cap 2500 and 4 of 14 at cap 6000, losing three it had already found and spending 5,833 completion tokens to do it. This is not truncation. Note (bsf) has covered caps too small for a seat's hidden reasoning; here the seat had more room, used it, and got worse. A cap is a parameter of the answer, not only a pre-flight measurement. P2 was dropped from this shape and the locating seat became P3 x-ai/grok-4.5 at low reasoning effort — 12 of 14, 1,659 tokens, $0.011 — while P1 truncated at 2500 for $0.031 and stayed on glossing. Full probe: amendment-v2-1-probe.md.

8.2 Three re-buys, and note (brx) did not apply

g_sh56 and g_sh9 returned empty bodies after spending the whole 2500-token cap inside their own reasoning; g_sh14's answer list stopped mid-item. Nothing could be re-parsed, so all three were re-bought at cap 6000, prompt, seat, temperature and item identical. The re-buys cost $0.062950 and the three dead bodies had already cost $0.092926 — the abandoned calls were the more expensive half, because a body that spends its whole cap and returns nothing bills for the whole cap. The cap-2500 bodies are kept under g_sh*_cap2500.

9. The lead's own pair — description only, and the breach is prior

T-hafez-boro-R59-MONO-v1 and T-hafez-boro-R59-PLAIN-v1 render Hafez غزل ۳۵ whole, twice, under the new regime R59, whose two arms differ in exactly one clause: whether the English line-ends must chime. Both logs were frozen before this design existed; contamination against Payne's whole stored corpus is 0 shared 7-grams, longest run 4 tokens, clean. Payne's ode 39 renders the same ghazal and was added to the sample before anything was glossed.

senses located at the line end (qāfiya senses) at the line end (ordinary senses) mean distance from the line end
lead, monorhymed 8 0.333 (n=3) 0.400 (n=5) 4.13
lead, unrhymed 9 0.400 (n=5) 1.000 (n=4) 1.89
Payne 1901, same ghazal 6 0.000 (n=2) 0.750 (n=4) 4.33

It points the same way as the published hands and it is not evidence: n is tiny, and the lead conceived the outcome variable before writing either arm (R59 §Known limitation). What it does add is a record from inside the hand — the MONO log's own count, frozen in advance, says four of eight rhyme-position senses reached the rhyme word and four were displaced by one to four words, and its §D names the three non-positional prices the rhyme charged: a dulled chime, one inversion, and two verbs made periphrastic.

The instrument agrees that the two arms differ in the intended way, on a figure nothing in the design tuned: 87.5% of the monorhymed arm's expected rhyming lines are detected as rhyming, against 12.5% of the unrhymed arm's.

10. Limits

  1. Observational. Nothing here was randomly assigned. Both critics asked for matched unrhymed renderings by independent hands under assigned instructions; the project cannot reach independent translators, and the one hand it can commission knew the measure. §9 is that experiment at n = 1 with a compromised hand, and it is labelled as such.
  2. The rhyme detector misses about a quarter of the lines it should catch — 73.95% of expected rhyming lines are classed rhyming, the rest lost to OCR damage and CMUdict's coverage of Victorian diction. Misclassified rhyming lines land in the open cell, so the effect is attenuated: +0.2451 is a floor.
  3. TERMPOS is a partial alternative explanation and does not reach. Rhyming lines end on a function word 9.25% of the time against open lines' 5.13% — a real difference in the right direction, four points against an effect of twenty-four.
  4. Senses migrate (§7): a located carrier is often not in the line that renders its own hemistich, so this page can say where an English sense-word sits given its line's rhyme obligation and may not say the rhyme pushed this sense out of this line.
  5. The gloss seat can see which Persian words are qāfiya. W6 finds no gloss-length difference (2.07 against 1.92 words), and the primary uses only non-qāfiya senses, where the objection does not reach. It stands against §4's qāfiya row.
  6. W4 passed at exactly its threshold, on a control that had to be rebuilt (§7).
  7. Payne is OCR and Leaf is worse in places; damage falls on both line types and cannot create the contrast, but it costs coverage.
  8. One seat on each side, no sense Tier-D calibrated. Descriptive lexical adjudications — the shape S015 found the panel usable on in the failing direction. provisional: true.
  9. Not comparable to step 1's 0.0384 as a number. That figure required the named English word to be among the last two of its line and to rhyme with the rendering's modal rime, from an eight-word window; this one asks a differently-framed question of a different seat on a whole poem. §4's small cell agrees with it in magnitude; it does not replace it.

11. Cost

stage calls actual
C pre-run critics, C1 + C2, cap 12000 — both NEEDS-REDESIGN, 19 findings 2 $0.104133000
seat and cap probe, on the lead's own poem only (note (bst)) 5 $0.094716750
G gloss, P1, cap 2500 — kept bodies, including the three re-buys at 6000 53 $0.580438000
G gloss — the three abandoned cap-2500 bodies, paid for and unusable 3 $0.092926000
L locate, P3 low effort, cap 2500 58 $1.078872400
POSBIAS (both executions) and X-POS 28 $0.106315200
per-request sum, and the ledgered figure 149 $2.057401350

Key usage 172.327928648 → 174.385329998, delta 2.057401350 against a per-request sum of 2.057401350 — exact to 1e-9, the second exact reconciliation in a row.