Repository path: wiki/findings/results/RS-20260902-rhyme-slot.md · rendered 2026-09-09
Page metadata (front matter)
| type | result |
|---|---|
| id | RS-20260902-rhyme-slot |
| status | frozen |
| created | 2026-09-02 |
| updated | 2026-09-02 |
| provisional | true |
| senses | accuracy, style-correspondence |
| links | workshop/experiments/E-20260902-rhyme-slot/design.md, workshop/experiments/E-20260902-rhyme-slot/critic-response.md, workshop/experiments/E-20260902-rhyme-slot/amendment-v2.md, workshop/experiments/E-20260902-rhyme-slot/amendment-v2-1-probe.md, wiki/arms/ARM-rhyme-family.md, wiki/findings/results/RS-20260830b-rhyme-family.md, workshop/regimes/R59-rhyme-clause-alone.md, workshop/translations/hafez-boro/R59-MONO-v1/translation.md, workshop/translations/hafez-boro/R59-PLAIN-v1/translation.md, framework/v0.2/README.md, config/budget.md, config/models.md |
RS-20260902 — an English rhyme costs the line-end slot by a factor of eleven, and the pre-rhyme slot is not a rhyme effect at all
ARM-rhyme-family step 2 (T2), and the arm closes. E-20260902-rhyme-slot design v1,
amendment v2 after two NEEDS-REDESIGN verdicts (19 findings, 5 BLOCKING), amendment v2.1 after
the seat probe. 149 API bodies, $2.057401350 of a declared $2.80. Verifier 35 checks, 0
failures, three mutation tests, all three caught. Every figure below is recomputed from the raw
bodies by runs/RS-20260902-rhyme-slot/verify.py, which re-derives the line geometry from its own
code and imports nothing from the analyser.
1. The question, and what step 1 left open
RS-20260830b-rhyme-family found that Leaf's English rhyme words almost never carry the senses
Hafez put at his rhyme — pooled 0.0384, nineteen of twenty-eight odes at zero — while the same senses
were found eight times as often somewhere in the last words of the same lines. It had nothing to
compare that 0.0384 against, and its §8 asked: is the pre-rhyme slot where the rhyme sense goes in
English, in every hand, or only in these two?
Both pre-run critics rejected the first design for asking it the wrong way, and the second question they forced is the one this page answers:
When a translator carries a Persian hemistich-final sense into English, does it land on the line's last word — and how much less often when that English line has to rhyme?
2. What the critics changed, because the amended design is not the registered one
Two seats, disjoint labs, dispatched before any other call: C1 openai/gpt-5.6-terra, C2
x-ai/grok-4.5, $0.104133000. Both NEEDS-REDESIGN. The response is in critic-response.md;
three findings governed the rebuild.
C1-1 andC2-2, independently: the predictor was Persian, not English. v1 compared the ghazal's rhyming hemistichs against its free ones — but a Hafez qāfiya is stock lexicon and here fuses with the copula, so the contrast could have come out positive with no English rhyme involved. The predictor is now the observable rhyme status of the English line the sense landed in, and the primary uses only senses from ordinary, non-qāfiya Persian hemistich-ends, so Persian word class is held constant.C2-3: v1's primary was close to a re-measurement. On a rhyming line, at the line end means the rhyme word carries the sense — which is step 1's own figure. The estimand was changed.C1-6: the permutation null was invalid, because the labels were never randomly assigned. Withdrawn in full. Intervals are cluster bootstrap over whole renderings, 20,000 draws, and every contrast on this page is observational.
C1-10 also restored the 14 English-radif renderings that v1 had excluded, as a stratum; the sample
is larger than the frozen design's because of it.
3. Materials
56 published renderings — Walter Leaf 1898 ×28 and John Payne 1901 ×28 — of 54 distinct Hafez
ghazals, matched to the Ganjoor text (Leaf at S231, Payne at S236). 747 sense positions: 371 from
rhyming hemistichs, 376 from free ones. Stage G glossed each ghazal's hemistich-final words from
the Persian alone (P1 openai/gpt-5.6-terra); stage L located each gloss in the whole English
rendering, with the poem's glosses in a seeded shuffle, so the seat could not tell which sense
came from a rhyming hemistich (P3 x-ai/grok-4.5, low reasoning effort — see §8.1). Disjoint labs.
368 senses located and verified.
4. The primary: no gate fired, and the effect is large
| qāfiya sense | ordinary hemistich-end sense | |
|---|---|---|
| landed in a rhyming English line | 0.3712 (n=132) | 0.0244 (n=41) |
| landed in an open English line | 0.3333 (n=54) | 0.2695 (n=141) |
Cell entries are the share of located carriers standing on the line's last word.
P1= 0.2695 − 0.0244 = +0.2451, 95% cluster bootstrap interval [0.1735, 0.3129], resampling whole renderings, 20,000 draws.F1does not fire.
An ordinary Persian line-final sense reaches the end of an English line 27% of the time when that line is free, and 2.4% of the time when it has to rhyme — a factor of eleven, with the Persian word class held constant. Step 1's 0.0384 now has the thing it lacked: a baseline. The two figures are the same order and were measured on different senses by different instruments; what is new is not the small number but the large one beside it.
Both hands, separately (F2 does not fire): Leaf 0.2115 (0.2115 at n=52 against 0.0000 at
n=15); Payne 0.2649 (0.3034 at n=89 against 0.0385 at n=26). Both strata (F3 does not
fire): renderings that carry an English radif +0.2571, renderings that do not +0.2413.
Selection is not doing it. The two Persian conditions were located at nearly the same rate —
0.5013 against 0.4840, a gap of 0.017 against W1b's 0.10 — and an adversarial bound that
assigns all 194 unlocated free-hemistich senses to the open cell as not-at-end still leaves
+0.0890. (That bound replaces the registered ALLITEM, which is not computable: an unlocated
sense has no landing line and therefore no cell. Constructed after seeing the data, and it can only
weaken the finding.)
5. The registered gates, all of them
| gate | value | fires? |
|---|---|---|
W1 location floor ≥ 0.30 |
0.5013 / 0.4840 by Persian condition | no. The cell-level version is not computable as written — a cell is defined by where a sense landed — and this is a defect in the gate's own wording, stated rather than glossed over |
W1b differential location ≤ 0.10 |
0.0173 | no |
W2 verification ≤ 20% |
2.90% (11 of 379 named words absent from their named line) | no |
W3 parse ≤ 15% |
0, after three stage-G re-buys (§8.2) | no |
W4 positional bias ≥ 4 of 6 |
4 of 6, on a re-executed control (§6) | no, at the threshold exactly |
W5 line length ≤ 2 tokens |
11.24 against 11.18 — 0.06 | no; no length adjustment was needed |
W6 gloss parity ≤ 2 words |
2.07 against 1.92 — 0.15 | no |
6. The second registered prediction FAILS, and that is the more interesting half
(The design calls it P2; it is a prediction, not the panel seat of the same name.)
Step 1 §3(c) — and framework §7.44's own headline — say that a displaced rhyme sense sits one word to the left of the rhyme. That was measured on one poem, the lead's own. Registered here against two comparators, and it clears one and fails the other.
| among carriers not on the qāfiya slot | share exactly one word left | uniform-placement null |
|---|---|---|
| in rhyming lines (n=118) | 0.1780 | 0.1045 |
| in open lines (n=139) | 0.1727 | 0.1014 |
The concentration is real — about 1.7× what uniform placement over the line's remaining positions would give — but it is the same in lines that must rhyme and lines that need not. The prediction required the rhyming share to exceed both comparators and it exceeds only the null.
The penultimate slot is where English puts a displaced sense-word generally. It is not something the rhyme does. The distributions confirm it: capped at six, rhyming lines run {0: 38, 1: 21, 2: 16, 3: 12, 4: 9, 5: 17, ≥6: 43} and open lines {0: 56, 1: 24, 2: 20, 3: 14, 4: 13, 5: 9, ≥6: 59} — the same shape.
So §7.44's headline is half right and half wrong, and the correction is written into §7.49. The first clause — the sense is not at the rhyme — is confirmed and strengthened on two hands and 964 positions. The second clause — it sits one word to its left — describes English, not rhyme, and was generalised from a single poem.
7. Three controls, and one of them was void before it was repaired
X-POS, the cross-poem control (bought on C2-8, because step 1 §6's mismatch prior covers
presence and not position). Four items pairing one ghazal's glosses with another's English, 63
sense-items, 11 located — 0.175 against the true items' 0.492. The instrument is not simply naming
whatever content word is available: put the senses in the wrong poem and it mostly answers NONE.
POSBIAS, the positional positive control — the first execution was VOID, not failed. Three of
its six items marked a bare punctuation token, because Leaf's OCR spaces ; and ? off as
separate words and the marker took the line's last token; the paraphrasing seat duly glossed the
semicolon ("a pause between clauses"). Scored as it stood it returns 1 of 6 and would have
withheld the whole run on a bug of the project's own making. It was rebuilt — last token containing
a letter, not a function word, and occurring exactly once in the rendering — and re-executed:
4 of 6, at W4's threshold exactly. The two misses are real: the seat named unite for wed
four words from the end, and returned NONE on a poor paraphrase ("are spoken" for flow). Both
executions are on the record and the repair is declared; a control repaired after it ran is weaker
than one that was right the first time, and this page does not pretend otherwise.
MIGRATION, and it vindicates the critics. Across Payne, whose renderings are bayt-for-bayt by
construction, 57.1% of located carriers stand outside their own bayt's two English lines. Any
design assuming hemistich-to-line alignment would have been wrong more often than right — which is
exactly what C1-2 and C2-1 said, and why the amended predictor is a property of the landing line
instead.
8. Instrument findings, both of which cost money
8.1 A larger cap made the answer worse — method note (bst)
On the same item, same prompt, same temperature, P2 google/gemini-3.6-flash located 8 of 14
senses at cap 2500 and 4 of 14 at cap 6000, losing three it had already found and spending 5,833
completion tokens to do it. This is not truncation. Note (bsf) has covered caps too small for a
seat's hidden reasoning; here the seat had more room, used it, and got worse. A cap is a
parameter of the answer, not only a pre-flight measurement. P2 was dropped from this shape and
the locating seat became P3 x-ai/grok-4.5 at low reasoning effort — 12 of 14, 1,659 tokens,
$0.011 — while P1 truncated at 2500 for $0.031 and stayed on glossing. Full probe:
amendment-v2-1-probe.md.
8.2 Three re-buys, and note (brx) did not apply
g_sh56 and g_sh9 returned empty bodies after spending the whole 2500-token cap inside their
own reasoning; g_sh14's answer list stopped mid-item. Nothing could be re-parsed, so all three were
re-bought at cap 6000, prompt, seat, temperature and item identical. The re-buys cost $0.062950
and the three dead bodies had already cost $0.092926 — the abandoned calls were the more expensive
half, because a body that spends its whole cap and returns nothing bills for the whole cap. The
cap-2500 bodies are kept under g_sh*_cap2500.
9. The lead's own pair — description only, and the breach is prior
T-hafez-boro-R59-MONO-v1 and T-hafez-boro-R59-PLAIN-v1 render Hafez غزل ۳۵ whole, twice, under
the new regime R59, whose two arms differ in exactly one clause: whether the English line-ends
must chime. Both logs were frozen before this design existed; contamination against Payne's whole
stored corpus is 0 shared 7-grams, longest run 4 tokens, clean. Payne's ode 39 renders the same
ghazal and was added to the sample before anything was glossed.
| senses located | at the line end (qāfiya senses) | at the line end (ordinary senses) | mean distance from the line end | |
|---|---|---|---|---|
| lead, monorhymed | 8 | 0.333 (n=3) | 0.400 (n=5) | 4.13 |
| lead, unrhymed | 9 | 0.400 (n=5) | 1.000 (n=4) | 1.89 |
| Payne 1901, same ghazal | 6 | 0.000 (n=2) | 0.750 (n=4) | 4.33 |
It points the same way as the published hands and it is not evidence: n is tiny, and the lead
conceived the outcome variable before writing either arm (R59 §Known limitation). What it does add
is a record from inside the hand — the MONO log's own count, frozen in advance, says four of eight
rhyme-position senses reached the rhyme word and four were displaced by one to four words, and its
§D names the three non-positional prices the rhyme charged: a dulled chime, one inversion, and two
verbs made periphrastic.
The instrument agrees that the two arms differ in the intended way, on a figure nothing in the design tuned: 87.5% of the monorhymed arm's expected rhyming lines are detected as rhyming, against 12.5% of the unrhymed arm's.
10. Limits
- Observational. Nothing here was randomly assigned. Both critics asked for matched unrhymed renderings by independent hands under assigned instructions; the project cannot reach independent translators, and the one hand it can commission knew the measure. §9 is that experiment at n = 1 with a compromised hand, and it is labelled as such.
- The rhyme detector misses about a quarter of the lines it should catch — 73.95% of expected rhyming lines are classed rhyming, the rest lost to OCR damage and CMUdict's coverage of Victorian diction. Misclassified rhyming lines land in the open cell, so the effect is attenuated: +0.2451 is a floor.
TERMPOSis a partial alternative explanation and does not reach. Rhyming lines end on a function word 9.25% of the time against open lines' 5.13% — a real difference in the right direction, four points against an effect of twenty-four.- Senses migrate (§7): a located carrier is often not in the line that renders its own hemistich, so this page can say where an English sense-word sits given its line's rhyme obligation and may not say the rhyme pushed this sense out of this line.
- The gloss seat can see which Persian words are qāfiya.
W6finds no gloss-length difference (2.07 against 1.92 words), and the primary uses only non-qāfiya senses, where the objection does not reach. It stands against §4's qāfiya row. W4passed at exactly its threshold, on a control that had to be rebuilt (§7).- Payne is OCR and Leaf is worse in places; damage falls on both line types and cannot create the contrast, but it costs coverage.
- One seat on each side, no sense Tier-D calibrated. Descriptive lexical adjudications — the
shape S015 found the panel usable on in the failing direction.
provisional: true. - Not comparable to step 1's 0.0384 as a number. That figure required the named English word to be among the last two of its line and to rhyme with the rendering's modal rime, from an eight-word window; this one asks a differently-framed question of a different seat on a whole poem. §4's small cell agrees with it in magnitude; it does not replace it.
11. Cost
| stage | calls | actual |
|---|---|---|
C pre-run critics, C1 + C2, cap 12000 — both NEEDS-REDESIGN, 19 findings |
2 | $0.104133000 |
| seat and cap probe, on the lead's own poem only (note (bst)) | 5 | $0.094716750 |
G gloss, P1, cap 2500 — kept bodies, including the three re-buys at 6000 |
53 | $0.580438000 |
G gloss — the three abandoned cap-2500 bodies, paid for and unusable |
3 | $0.092926000 |
L locate, P3 low effort, cap 2500 |
58 | $1.078872400 |
POSBIAS (both executions) and X-POS |
28 | $0.106315200 |
| per-request sum, and the ledgered figure | 149 | $2.057401350 |
Key usage 172.327928648 → 174.385329998, delta 2.057401350 against a per-request sum of 2.057401350 — exact to 1e-9, the second exact reconciliation in a row.