Repository path: workshop/experiments/E-20260830b-rhyme-family/design-v2.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260830b-rhyme-family-v2 |
| status | frozen |
| created | 2026-08-30 |
| updated | 2026-08-30 |
| senses | accuracy, style-correspondence |
| links | wiki/arms/ARM-rhyme-family.md, workshop/experiments/E-20260830b-rhyme-family/design.md, workshop/experiments/E-20260830b-rhyme-family/critic-response.md, workshop/regimes/R57-leaf-contract.md, workshop/translations/hafez-bekonad/R57-v1/translation.md, wiki/base/anchors/A-leaf-hafez/README.md, wiki/findings/results/RS-20260830-leaf-contract.md, runs/RS-20260829-radif-hands/census.json, config/models.md |
E-20260830b design v2 — can one English rhyme hold the senses a ghazal puts at its rhyme?
This file supersedes design.md and is what was dispatched. v1 was frozen, sent to two
independent adversarial critics, and both returned NEEDS-REDESIGN — 27 findings, 18 of them
BLOCKING, converging independently on the same five structural faults. Eleven amendments were
accepted and three remedies overruled with written reasons: critic-response.md. Frozen before
any stage-F or stage-S call.
1. Where the question comes from
T-hafez-bekonad-R57-v1 (S233) rendered Hafez غزل ۱۸۷ whole under Walter Leaf's printed 1898
contract. Its translator's log, frozen before any evaluation existed, registered one prediction
that has never been tested:
the constraint that actually bites in a monorhymed English ghazal is not the metre and not the inversion but whether the source's rhyme-bearing senses happen to lie inside a single English rhyme.
The log's §D2 records that -ending happened to hold an English word for almost every sense that
poem placed at its rhyme, and §D4 that the poem's thirteen losses were content words squeezed by
the metre, not one of them caused by the rhyme. Both are one hand's observations on one poem.
2. The subject-rule sentence (wiki/tracks.md §The subject rule)
What this unit teaches about translating literature: whether the senses a source poem places at its rhyme can be covered by a single English rhyme — something a translator can find out from the source before writing a line — predicts how much of that rhyme a published translator delivers, and whether it predicts which poems he agreed to translate at all.
3. Constructs, named honestly
For a ghazal with N rhyming positions the Persian places N senses at the line-end.
A, available coverage — from the Persian alone, blind to every English rendering: the fraction of those N senses for which one single English rhyme supplies a word.R, realised coverage — the fraction of those N senses for which the published translator's actual monorhyme words supply one.
What A is not. A is measured by asking competent readers of both languages to find such a
rhyme. It is therefore what a competent reader can find, which is the translator's real
situation — not a property of the English dictionary and not a census of English (critic
C1-1, overruled remedy O2). The result page may not call it either.
Q1 — does A, computed from the Persian alone, predict R?
Q2 — does A predict which ghazals Leaf agreed to translate? (critic C2-2, accepted as A6.)
v1 said a positive Q1 would show "the difficulty is in the language pair and not in the
translator". That sentence is struck: one translator cannot support it (C1-4). Whatever Q1
returns is a statement about Leaf's twenty-eight odes.
4. Materials
Persian. Ganjoor's Qazvini–Ghani Divan of Hafez, fetched 2026-08-30, raw JSON preserved at
runs/RS-20260830b-rhyme-family/ganjoor/. Public domain. Rhyming positions, radif and qāfiya by the
frozen census rule of runs/RS-20260829-radif-hands/census.py, imported verbatim into
build.py: positions = hemistich 1 plus every second hemistich; radif = longest common trailing
word sequence over the positions, minus one word reserved for the qāfiya.
English. Walter Leaf, Versions from Hafiz: An Essay in Persian Metre (1898), public domain,
Internet Archive OCR; the same 28 odes and the same analyse.bearers extraction as
RS-20260830-leaf-contract, imported unchanged.
The Leaf↔Ganjoor match is runs/RS-20260829-radif-hands/leaf_match.json (S231).
Unselected comparison pool for Q2, drawn before any coverage number exists. From the
495-ghazal census: radif-bearing, 7–9 bayts, neither Leaf's 28 nor already rendered by this project
(sh124, sh118, sh4, sh131, sh187) — 215 ghazals — drawn at even index intervals by
build.py, in this frozen order:
sh2, sh30, sh45, sh60, sh76, sh105, sh123, sh147, sh165, sh176, sh191, sh209, sh221, sh237,
sh263, sh275, sh339, sh357, sh377, sh402
Any budget-forced truncation takes a prefix of that order and never a selection by A
(C2-12).
5. Strata
The primary is all 28 odes. v1's PRIM cut on |N − M| ≤ 1, where M is Leaf's own position
count — a property of the translation, reflecting his omissions (C1-5, C2-5). It survives only
as a declared exploratory sensitivity analysis on the 19 odes IV, V, VI, VII, VIII, IX, X, XI,
XII, XIII, XV, XIX, XX, XXII, XXIII, XXIV, XXV, XXVI, XXVII, and the result page must label it so.
6. Set measurement, and what it does and does not mean
Fifteen of the 28 odes have N ≠ M, so no honest position-by-position pairing exists. Both stages
therefore work on sets: a sense displaced from one couplet to another is not counted as a loss. The
construct being measured is coverage of the source's rhyme senses by the translator's monorhyme,
not positional carriage, and §11 forbids the result page from sliding between the two (C2-9).
7. The bought stages — rater-disjoint by construction
The single fault both critics named first is that v1 drew A and R from the same seats
(C1-2, C2-1). The two sides of the correlation now come from disjoint seats at disjoint labs.
| stage | seats | lab |
|---|---|---|
F (produces A) |
P1 openai/gpt-5.6-terra, P3 x-ai/grok-4.5 |
OpenAI, xAI |
S (produces R) |
P2 google/gemini-3.6-flash, PR qwen/qwen3.7-max |
Google, Alibaba |
PR is the first reserve in config/models.md, promoted for this run for rater-disjointness
and for no other reason; that is provenance, not a change of panel composition. P4 and P5 are
out under notes (bps) and (bne). temperature 0, max_tokens 2500, dispatch order shuffled
against printed seed 20260830. No seat is told the hypothesis, the translator, the poet, the
century, or that another stage exists.
7.1 Stage F1 — senses only, rhyme never mentioned
One call per ghazal per stage-F seat. The seat sees the ghazal's rhyming hemistichs in Persian
script with the qāfiya of each printed after it, and the radif if there is one. It gives the
contextual sense of each qāfiya in one English word or short phrase, or answers UNGLOSSABLE where
the marked unit is not a meaningful word on its own (C1-7, C2-11). The word "rhyme" appears
only in the neutral description of what a qāfiya is; nothing about English, verse or translation
appears at all.
7.2 Stage F2 — the rhyme, on frozen senses
A second call to the same seat, carrying its own F1 sense list verbatim and the instruction
that the senses are fixed and may not be revised. It proposes one English rhyme — naming the rhyme
and listing its words — and says for each sense whether that rhyme supplies a word carrying it,
naming the word when it does. Sealing F1 from F2 is what stops a seat bending a gloss toward
something that rhymes (C1-1, C2-7).
Mechanical validity screen. Every named word is looked up in CMUdict via
tools/rhyme_pairs.py, and rhyme is judged by RS-20260830-leaf-contract's analyse.rime_nor,
which carries Leaf's own non-rhotic licence. A named word that does not rhyme with the modal rime of
that seat's own named set is recoded NO. A word absent from CMUdict is counted UNSCREENABLE, and
every screened figure is reported twice — once counting those NO, once dropping them (C1-6,
C2-8). The screen can only lower A.
A(ghazal) = mean of the two stage-F seats' screened coverage. Declared robustness estimator,
reported beside it and never substituted for it: the max of the two seats.
Stage F runs on 48 ghazals: Leaf's 28 and the 20 unselected.
7.3 Stage S — realised coverage
One call per item per stage-S seat. The seat sees the N Persian rhyming hemistichs with the qāfiya
marked, and separately the M line-endings of an anonymous English verse rendering. For each Persian
qāfiya it answers YES (naming the English word), NO, or UNGLOSSABLE. The prompt tells it to be
strict and that a word merely in the same area of meaning is NO.
The screen that makes R monorhyme carriage rather than set cover (C1-3, C2-9). A YES is
kept only if the named English word (i) is one of that rendering's actual rhyme-bearing words and
(ii) rhymes, under rime_nor, with the modal rime of that rendering's own bearer set. Otherwise it
is recoded NO. R_rhyme is the primary; R_set, unscreened, is reported beside it.
UNGLOSSABLE positions, wherever they arise, are dropped from numerator and denominator of both
A and R identically.
7.4 Controls
X-SHUF, 4 items — carries the gate. An ode's Persian shown against a list built by taking one bearer word from each of four different odes, destroying poem coherence and any single rhyme family. Gate: pooledR_true− pooledR_SHUF≥ 0.20, else the instrument is not measuring sense correspondence andQ1is withheld.X-MIS, 4 items — descriptive only. The in-field mismatch, one ode's Persian against another ode's English. Hafez's rhyme senses are stock and Leaf's Victorian rhyme vocabulary is stock, so a small gap here is expected and is not a gate (C2-4).POS—T-hafez-bekonad-R57-v1'sFREEarm, whose own frozen log claims one rhyme carried all eight senses. It should score high. A failure is an instrument limit, not a withholding.
The eight control pairings are printed by stages.py --print-controls into the run record before
dispatch (C1-14).
8. The translation limb — an illustration, and it is not evidence
Both critics established that a limb the lead selects, predicts and executes cannot falsify anything
(C1-10, C1-11, C2-3). v1's registered P4 is withdrawn.
What remains is worth doing and is honest about what it is: the study limb's measure selects a
poem, and translating that poem generates an enumeration — bayt by bayt, in a frozen log — of
what a scarce rhyme actually costs, which no correlation can show. That is the wire, and it is
generates, not tests (continue-prompt.md §4).
Selection. Stage F runs on the twenty unselected ghazals; the poem translated is the one with
the lowest A. The selection script prints the chosen slug and nothing else — no coverage
values, no ranking, no margin. Every stage-F body for the chosen ghazal stays sealed until the
translator's log is frozen. No published English of it is opened before the freeze;
payne_odes.json is in this repository and is not consulted for the chosen slug at any point.
Regime: R57, the FREE arm — monorhyme at every rhyming position, the radif mandatory, one
English couplet per Persian bayt, the metre released. The same arm of the same regime rendered
غزل ۱۸۷, which makes the two poems comparable in everything but A.
Its blind stage-S score is reported descriptively, is internal-judgment-only as craft
evidence, and is never support for Q1.
9. Predictions, gates, failure criteria — registered
P1, primary, n = 28. Spearman ρ between A (stage-F seats) and R_rhyme (stage-S seats) is
positive. Monte Carlo permutation test — random.Random(20260830), 100,000 relabelings, one-sided,
(1 + #{ρ* ≥ ρ}) / (1 + 100000), ties by midrank — at α = 0.05 (C1-13).
Power, registered before the run (C1-12, C2-6): n = 28 one-sided at α = 0.05 has ~80% power
at ρ = 0.45, which is the minimum effect of interest. A non-significant P1 may not be
reported as evidence against a moderate effect, and the result page must say so.
P2, secondary — Q2. Leaf's 28 have higher A than the 20 unselected ghazals.
Mann–Whitney U, two-sided, exact where feasible.
P3, mechanical, no API cost. A correlates positively with the per-ode rhyme-pass rate
adjudicated blind at S233 and negatively with Leaf's per-ode rhyme-word repetition rate. Labelled
convergent evidence for general rhymeability, not for sense coverage (C1-15).
Withholding criteria — any one withholds P1, and the run says which.
W1— IQR ofAacross the 28 below 0.10: a predictor with no variance.W2— theX-SHUFgate fails its 0.20 margin.W3— more than 30% of calls in either stage fail to parse after one re-parse.W4— the two stage-S seats' per-odeR_rhymediffer by more than 0.30 on average.W5, the placebo gate (C1-9,C2-10). Four quantities that are not rhyme-family availability — Persian position count N, mean qāfiya character length, the S231 match ratio, and Leaf's position ratio M/N — are correlated withR_rhymeby the same statistic. If any of them reaches or exceedsA's ρ, the mechanism claim is withheld and the run reports the association as unexplained.
What would refute the conjecture. P1 returning ρ ≤ 0 with W1–W5 clear. The translation
limb takes no part in this rule (C1-11).
10. Spend
Declared ceiling $2.60; $0.112091900 already spent on the two critic seats. UTC-day headroom at session start $3.579539500 (S233 spent $1.420460500 of $5.00).
Planned calls: F1 48 × 2 = 96, F2 48 × 2 = 96, S 38 items × 2 = 76 — 268. Worst case from
max_tokens alone, per note (abc), is 268 × 2500 output at the dearest seat rate ($6.00/M) =
$4.02, above the ceiling, so the run is probed per note (bsh): the first 8 calls are
dispatched alone, billed cost read from usage.include, and the rest goes only if the extrapolated
total sits under the ceiling. If it does not, the candidate pool truncates to a prefix of the
frozen draw order — twelve, then eight — and never by A.
The translation limb, the census, the Ganjoor fetch, all extraction, all screening and the contamination measurement are the lead's own and are not ledgered (charter §3, A4).
11. What this run cannot do
- One hand, and a hand who chose his own poems. Everything about
Q1is conditional on Leaf's menu (C2-2). Payne 1901, who monorhymed all 573, is step 2 of the arm; his text in this repository is OCR of visibly poor quality and unmatched to Ganjoor, and forcing him into this run would add noise, not power (overrule O1). - Four model seats, no sense Tier-D calibrated. The judgments here are descriptive lexical adjudications, the shape S015 found the panel usable on in the failing direction. Not calibration.
Ais a lower bound found by readers, not a census of English (§3).- Set coverage is not positional carriage (§6). No sentence may claim the second from the first.
- CMUdict is modern and American; Leaf is 1898 and British. Mitigated by
rime_norand by double reporting, not solved. - n = 28 detects ρ ≈ 0.45. A null here is not a refutation of a moderate effect.