Repository path: workshop/experiments/E-20260830b-rhyme-family/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260830b-rhyme-family |
| status | superseded |
| created | 2026-08-30 |
| updated | 2026-08-30 |
| senses | accuracy, style-correspondence |
| links | wiki/arms/ARM-rhyme-family.md, wiki/tracks.md, workshop/regimes/R57-leaf-contract.md, workshop/translations/hafez-bekonad/R57-v1/translation.md, wiki/base/anchors/A-leaf-hafez/README.md, wiki/findings/results/RS-20260830-leaf-contract.md, runs/RS-20260829-radif-hands/census.json, config/models.md |
E-20260830b — is a monorhyme carryable because of the poem, or because of the English dictionary?
Frozen before any call. Nothing below is amended after a datum is read; amendments made in
reply to the pre-run critic are marked and dated in critic-response.md, and the design file is
re-frozen before dispatch.
1. Where the question comes from
T-hafez-bekonad-R57-v1 (S233, 2026-08-30) rendered Hafez غزل ۱۸۷ whole under Walter Leaf's own
printed 1898 contract. Its translator's log, frozen before any evaluation existed, registered one
prediction, and it has not been tested:
This is the log's one prediction, registered here before the audit ran: the constraint that actually bites in a monorhymed English ghazal is not the metre and not the inversion but whether the source's rhyme-bearing senses happen to lie inside a single English rhyme. Leaf, translating twenty-eight odes, met that condition twenty-eight times or gave something up.
The log's own §D2 records why that poem yielded: -ending happens to hold mending, tending,
defending, ending, sending, ascending, befriending, amending — an English word for almost every
sense the Persian put at its rhyme. §D4 then records that the poem's thirteen losses were content
words squeezed out by the metre and that not one of them is caused by the rhyme or by the radif.
Both observations are the lead's own, on one poem, and neither is evidence.
The claim is falsifiable and it is about the two languages, not about this project: whether a fixed monorhyme can be carried out of Persian into English is decided in part by an accident of the English lexicon — whether one English rhyme family happens to hold words for the senses the source placed at its rhyme. If it is true, a translator can run the test before writing a line, and it belongs in the handbook next to §7.36 and §7.40, which already tell a translator what he can know from the source alone.
2. The subject-rule sentence (wiki/tracks.md §The subject rule)
What this unit teaches about translating literature: whether the senses a source poem places at its rhyme can be covered by a single English rhyme — a fact about the two lexicons, readable off the source before any English exists — predicts how much of that rhyme a published translator actually delivers.
It is not a question about this project's instruments. The rhyme checker and the census rule are mechanical and are controlled inside the run; they are not the subject.
3. Question, and the two coverages
For a ghazal with N rhyming positions, the Persian places N rhyme-bearing senses at the line-end.
- Available coverage
A— computed from the Persian alone, blind to every English rendering: the largest fraction of those N senses for which one single English rhyme can supply a word. - Realised coverage
R— computed on a published English rendering of the same ghazal: the fraction of those N senses for which the translator's actual English rhyme words supply one.
Q1— doesA, computed from the Persian alone, predictR?
A is a property of the poem and the English dictionary. R is what one man did with it. If A
predicts R, the difficulty is in the language pair and not in the translator.
4. Materials
Persian. Ganjoor's Qazvini–Ghani text of the Divan of Hafez, fetched 2026-08-30, raw JSON
preserved under runs/RS-20260830b-rhyme-family/ganjoor/. Public domain (Hafez, d. c. 1390).
Rhyming positions and qāfiya are taken by the frozen census rule of
runs/RS-20260829-radif-hands/census.py, imported verbatim into build.py and not modified:
positions = hemistich 1 plus every second hemistich; the radif is the longest common trailing word
sequence over the positions, minus one word reserved for the qāfiya.
English. Walter Leaf, Versions from Hafiz: An Essay in Persian Metre (London: Grant Richards,
1898), public domain, Internet Archive OCR, the same 28 odes and the same extraction
RS-20260830-leaf-contract used; the rhyme-bearing word per position is analyse.bearers,
imported unchanged.
The Leaf↔Ganjoor match is runs/RS-20260829-radif-hands/leaf_match.json (S231), built by
consonant-skeleton similarity on the opening hemistich. It carries a ratio per ode.
The two lead renderings, entered as extra items and never in the correlation:
T-hafez-bekonad-R57-v1, theFREEarm — غزل ۱۸۷, highAby the log's own account.- the translation limb of this session (§8), on the lowest-
Acandidate.
5. Strata, frozen now
Leaf does not render every bayt of every ghazal, and the OCR does not always split his lines where he did. Comparing the two position counts:
PRIM, the primary stratum — the 19 odes with|N − M| ≤ 1and matchratio ≥ 0.70, where N is the Persian position count and M is Leaf's: IV, V, VI, VII, VIII, IX, X, XI, XII, XIII, XV, XIX, XX, XXII, XXIII, XXIV, XXV, XXVI, XXVII.ALL, all 28, reported as a secondary and never as the primary.
The nine odes outside PRIM are I, II, III, XIV, XVI, XVII, XVIII, XXI, XXVIII. This partition is
fixed before any A or R exists and is a function of the printed page and the S231 match only.
6. The measurement is alignment-free, and that is deliberate
Fifteen of the 28 odes have N ≠ M, so a position-by-position pairing of a Persian rhyme word with
an English one cannot be built honestly. It is also not the right question: the conjecture is about
whether the English rhyme as a set covers the Persian rhyme senses as a set. Both stages
therefore work on sets, and displacement of a sense from one couplet to another is not counted as a
loss.
7. The two bought stages
Seats are P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5
(config/models.md; P4 and P5 are out under notes (bps) and (bne)). Three seats, one call per
seat per item, temperature 0, max_tokens 2500, dispatch order shuffled against printed seed
20260830. No seat is told the hypothesis, the translator, the poet, the century, or that a
second stage exists.
7.1 Stage F — available coverage A, from the Persian alone
One call per ghazal. The seat sees the ghazal's rhyming hemistichs in Persian script with the qāfiya word of each marked, and the radif if there is one. No English rendering of any kind is shown.
The seat must (i) give the core sense of each qāfiya word in its line, in one English word or short phrase; (ii) propose one English rhyme — naming the rhyme sound and listing its words; (iii) for each position say whether that rhyme supplies a word carrying the sense it just gave, and name the word when it does.
Mechanical validity screen, applied to every seat answer before any A is computed. Every word
the seat names is looked up in CMUdict (tools/rhyme_pairs.py). A named word that does not
strictly rhyme with the modal rime of that seat's own named set is recoded to NO. A named word
absent from CMUdict is recoded to NO. The screen can only lower A, never raise it, and it is
applied identically to every item.
A(ode) = mean of the three seats' screened coverage. Declared robustness estimator, reported
beside it and not substituted for it: max of the three seats, on the ground that "does English
hold such a rhyme" is an existential claim.
Stage F runs on 38 ghazals: Leaf's 28, plus the ten candidates of §8.
7.2 Stage S — realised coverage R, on the English
One call per item. The seat sees, for one ghazal, the N Persian rhyming hemistichs with the qāfiya marked, and separately the M English rhyme-bearing words of an anonymous verse translation, each with the line it ends. For each Persian qāfiya word the seat answers: does any word in the English list carry its sense — YES or NO — and names the English word when YES.
R(item) = mean over the three seats of (YES count / N).
Nothing identifies the translator, the century, or the fact that some items are controls.
7.3 Controls inside stage S
X— mismatched-ode control, 6 items. The Persian rhyme words of ode A shown with the English rhyme words of ode B, both drawn fromPRIMby the printed seed. Registered: pooledRonXmust be at least 0.20 below pooledRon the true items, or the instrument is not measuring sense correspondence andQ1is withheld.POS— a positive control that costs nothing extra.T-hafez-bekonad-R57-v1FREEis one of the two lead items; its own frozen log claims all eight senses were carried by one rhyme. Registered: it should scoreR ≥ 0.75. A failure here is reported as an instrument limit, and does not by itself withholdQ1.
8. The translation limb, and the wire
The wire, one sentence. The study limb's source-side measure chooses the poem the translation limb attempts, and the finished poem is then scored by the same blind instrument as Leaf's twenty-eight — so a measure computed from Persian alone is tested both across a published book and in the act of translating one poem it says should be hard.
Selection, frozen. From the 495-ghazal census, the radif-bearing ghazals of 7–9 bayts that are
neither Leaf's 28 nor already rendered by this project (sh124, sh118, sh4, sh131, sh187)
— 215 of them — ten drawn at even index intervals by build.py: sh2, sh47, sh78, sh133,
sh170, sh202, sh229, sh272, sh352, sh401. Stage F is run on all ten. The poem
translated is the one with the lowest A.
Blinding of the translator. The selection script prints the ten A values and nothing else.
Every stage-F body for the chosen ghazal — the glosses, the proposed rhymes, the named English
words — stays sealed until the translator's log is frozen. No published English of the chosen
ghazal is opened before the freeze; Payne 1901 rendered the whole Divan and
runs/RS-20260830-leaf-contract/payne_odes.json is in this repository, and it is not consulted
for the chosen slug at any point in this session.
Regime: R57, the FREE arm — clauses 3, 4, 5 (monorhyme at every rhyming position; the radif
mandatory; one English couplet per Persian bayt), with clause 2, the metre, released. Releasing
the metre is the point: S233's poem paid its price in content words squeezed by the metre and paid
nothing at the rhyme, so with the metre gone the rhyme is the only clause left to fail. The same arm
of the same regime rendered غزل ۱۸۷, which makes the two poems comparable in everything but A.
Registered prediction P4, written before the poem is chosen. The new poem's realised coverage
R, measured blind by stage S, will be lower than T-hafez-bekonad-R57-v1's. P4 is a
one-poem contrast by one hand and is not powered; it is reported as such and never as support
for Q1 on its own.
9. Predictions, gates and failure criteria — registered
Primary P1 — on PRIM (n = 19). Spearman rank correlation between A and R is positive.
Test: exact one-sided permutation over 100,000 random relabelings against printed seed 20260830,
α = 0.05. Registered direction: ρ > 0.
P2 — on ALL (n = 28), same statistic, reported as secondary whatever P1 does.
P3 — mechanical, no API cost. Across PRIM, A is positively correlated with the per-ode
rhyme-pass rate adjudicated blind at S233 (RS-20260830-leaf-contract, the L-RHYME audit) and
negatively correlated with Leaf's per-ode rhyme-word repetition rate. Reported as secondary.
Withholding criteria. Any one of these withholds P1 and the run says so.
W1— the interquartile range ofAacrossPRIMis below 0.10. A predictor with no variance cannot be tested and the design would be claiming a null it never had power for.W2— the mismatched-ode controlXfails its 0.20 margin.W3— more than 30% of stage-S or stage-F calls fail to parse after one re-parse.W4— the three seats'Rvalues disagree so far that the mean is not a measurement: mean pairwise absolute difference in per-odeRabove 0.30.
What would refute the conjecture. P1 returning ρ ≤ 0 on PRIM with W1–W4 all clear, and
P4 failing in the same run, is a refutation of the S233 log's prediction and will be written as
one.
10. Spend
Ceiling declared $2.40 against a UTC-day headroom of $3.579539500 (S233 spent $1.420460500 of
$5.00). Worst case from max_tokens alone, per note (abc): 224 calls × 2500 output tokens at the
dearest seat rate ($6.00/M) = $3.36, which exceeds the ceiling, so the run is staged and probed
per note (bsh): the first 6 calls (2 items × 3 seats) are dispatched alone, their billed cost
read from usage.include, and the remainder is dispatched only if the probe-extrapolated total
sits under the ceiling. If it does not, stage F's candidate pool is cut from ten to five and the
mismatched control from six items to four, in that order.
The translation limb, the census, the Ganjoor fetch, all extraction, all CMUdict screening and the contamination measurement are the lead's own and are not ledgered (charter §3, A4).
11. What this run cannot do
- Three model seats, no sense Tier-D calibrated (
config/models.md). Every judgment here is a descriptive lexical adjudication — does an English word carry a Persian word's sense — not a quality judgment, which is the shape S015 found the panel usable on in the failing direction; it is still not calibrated and the result page says so. - One hand, one poet, one language pair. A second hand — Payne 1901, who monorhymed all 573 ghazals
— is step 2 of
ARM-rhyme-familyand is not attempted here. - Leaf's book is OCR. Rhyme-bearing words absent from CMUdict are reported and are a named limit.
Ais measured by asking models to invent a rhyme family. That is a lower bound on what English holds, not a census of it, and the result page must not call it one.