Repository path: workshop/experiments/E-20260831b-radif-hands/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260831b-radif-hands |
| status | frozen |
| created | 2026-08-31 |
| updated | 2026-08-31 |
| senses | style-correspondence |
| provisional | true |
| internal-judgment-only | true |
| links | wiki/arms/ARM-radif-hands.md, wiki/findings/results/RS-20260829-radif-hands.md, framework/v0.2/README.md, config/models.md, config/budget.md, workshop/translations/hafez-sahar-bolbol/R55-v1/translation.md, runs/RS-20260831b-radif-hands/payne_all.py, runs/RS-20260831b-radif-hands/code_payne.py |
The predictor put to a whole book: Payne's 199 odes against the Persian radif census
ARM-radif-hands step 2 (T5). Frozen before any classification was bought and before any
matching beyond the ten-item probe. Nothing here is a judgment of quality; no reader is asked
anything; Tier D is NOT PASSED, so no sentence may say anyone hears, prefers or wants anything.
1. The question, and why step 1 could not answer it
framework/v0.2 §7.36.1 and §7.40 tell a translator, before the first English line, that certain
Persian radifs cannot be carried. RS-20260829-radif-hands refuted the printed impossibility on
nine matched items and two hands, downgraded both clauses to costly, at the price of an
inversion in an archaizing register (§7.41), and said in its own limits that nine items settle an
existence claim and cannot support a rate.
Step 2 is the powered version. Payne 1901 volume 1 is the whole of Brockhaus I–CC, 199 odes printed (he omits XIV as spurious, in his own footnote), and he did not choose his poems — he translated the Divan entire, which is the selection confound Leaf's twenty-eight cannot shed.
P1 (primary). Among the Ganjoor ghazals Payne translated in volume 1 that the frozen census gives a radif, does the radif's grammatical cell predict whether Payne carries a repeated English tail?
P2 (primary). Among the ghazals the census finds radif-free, how often does Payne's English carry a repeated tail anyway — i.e. how often does the English manufacture a repetition the Persian has not got?
P2 exists because step 1 saw Leaf do exactly this at two items, both of them ghazals whose repeated
element is a bound person-suffix that the census rule does not count as a radif, and registered
it as an observation on two items. 190 of the census's 495 ghazals are radif-free; the ones inside
Brockhaus I–CC give the first properly denominated estimate this project can make.
2. Materials, all already on the shelf or free
- The census —
runs/RS-20260829-radif-hands/census.json, frozen 2026-08-29: all 495 Ganjoor (Qazvini–Ghani) Hafez ghazals, radif and qāfiya by a mechanical rule. 305 carry a radif (0.616), 190 do not. Not rebuilt, not amended. - Payne 1901, vol. 1 — already stored whole at
workshop/experiments/E-20260830-leaf-contract/materials/payne1901_djvu.txt(public domain; Payne d. 1916). Segmented byruns/RS-20260831b-radif-hands/payne_all.pyinto 199 odes; the segmentation is checked against the 142 Roman headings the OCR leaves intact. - The ten known Brockhaus→Ganjoor pairs, read off the citations Leaf 1898 prints under his own
odes (
(No. 6. R. i p. 16.)and so on): 6→sh5, 8→sh3, 43→sh25, 44→sh26, 52→sh43, 79→sh63, 84→sh89, 123→sh143, 151→sh148, 198→sh117. These are the held-out accuracy set for the match and are never given to a seat.
3. Procedure
Stage A — carriage coding, mechanical, already run and frozen.
runs/RS-20260831b-radif-hands/code_payne.py, the rule and the bar declared in its docstring before
the numbers were read: an ode carries an English tail iff one trailing sequence of 1–3 words closes
≥ 0.80 of its extracted rhyme lines and ≥ 4 of them. Tokenisation and the Levenshtein-1 merge of
scanner variants are step 1's carriage.py unchanged. Result, frozen at commit time: 134 of 199
odes carry a tail (0.673), sensitivity 0.60→146, 0.70→142, 0.80→134, 0.90→107, 1.00→99.
Stage M — the match, bought from two disjoint seats. Each seat is shown the 495 Persian matla's,
numbered, and Payne's English opening couplet, and returns an index and a confidence. Payne's
coded English tail is never shown to either seat: the tail is half the measurement, and showing it
would let the match be made on the very correspondence under test. 199 items, 20 per call, 10 calls
per seat. Seats: P1 openai/gpt-5.6-terra and P2 google/gemini-3.6-flash — disjoint labs, per
RS-20260830b-rhyme-family's finding that a shared rater is this instrument's first fault.
A match is accepted only where the two seats return the same index. Disagreements and NONEs are dropped, not adjudicated by the lead: the lead wrote the rule under test.
Stage G — the grammatical cell, bought blind from three seats. Every distinct radif among the
accepted, radif-bearing matches is shown inside its own maṭlaʿ, with no English, no rule and no
prediction, to P1, P2 and P3 (x-ai/grok-4.5), against the closed checklist step 1 used
(RS-20260829-radif-hands §8): part of speech, is it a case particle, is it a finite verb, is it
transitive, is it a copula, does it stand in an ezāfe to the word before it, plus a gloss. Ten
radifs per call. The cell is the 2-of-3 majority on each bit.
Stage V — verification. verify.py recomputes every reported number from the raw JSON, and runs
with --mutate to confirm the checks can fail.
4. Registered predictions
framework/v0.2 as it stands after §7.41 is the thing being predicted from.
- PR1 — the case particle holds. Ghazals whose radif is or ends in
راare carried by Payne at rate 0. The census has 5 whose radif is exactlyرا; how many fall inside Brockhaus I–CC is not known in advance. - PR2 — the transitive-finite-verb cell is not zero. §7.36.1 as printed says 0; §7.41 says costly, not impossible. Predicted carriage in that cell > 0.50.
- PR3 — the grammatical cell explains little. Payne's overall carriage is already known to be 0.673 and the source's radif rate is 0.616; predicted difference in carriage rate between the cells §7.36.1 calls unreachable and the rest is < 0.20.
- PR4 — manufacture is rare. Predicted
P2rate — a tail in Payne where the census finds no radif — < 0.25.
Every one of these can fail, and PR3 and PR4 are the ones that would most change the framework. PR3 failing upward would restore a version of §7.36.1; PR4 failing upward would put a wholly new clause in, about English manufacturing a repetition Persian has not got.
5. Failure and withholding criteria, declared before the run
G1Match accuracy. The two seats' agreed matches must recover at least 9 of the 10 held-out Brockhaus anchors. Below that,P1is withheld entirely.G2Match coverage. At least 0.70 of the 199 odes must reach an accepted match. Below that,P1is reported as an underpowered subset with its denominator, not as a rate for the book.G3Classification agreement. Every checklist bit must have a 2-of-3 majority; bits without one are dropped and the items resting on them are dropped with them.G4Cell sizes. No carriage rate is reported for a cell with fewer than 8 items; smaller cells are printed as counts with their denominators and nothing is computed from them.G5No association statistic is computed unlessG2passes, and none is computed at all for cells belowG4's bar.G6The OCR bar. Every rate is reported at the frozen 0.80 bar and again at 1.00. If a reported comparison reverses between the two bars, the comparison is withheld.
6. What this design cannot do
- One hand. Payne is one translator with one register. Nothing here separates what English can host from what Payne's archaizing English hosts; §7.42 measured that price separately and this design does not re-measure it.
- The census's definition. A radif is a repeated word sequence, so bound-suffix repetitions
are radif-free by definition and land in
P2by construction. That is stated, not hidden:P2is a measurement of manufacture against this definition, and the definition is the standard one. - Brockhaus is not Qazvini–Ghani. Payne translated a different edition, which prints more couplets than Ganjoor at some poems (step 1 found this at two of nine). Carriage is coded on Payne's own printed lines and the census on Ganjoor's, so a poem can have a different number of positions on each side. Nothing in the coding compares the two counts.
- The seats classify grammar; they do not validate it. They are a reproducibility and independence diagnostic, exactly as in step 1.
- Not a claim about readers. Tier D is NOT PASSED.
7. Budget
Declared ceiling $2.20 for the UTC day 2026-08-31, which had $0.00 spent before this session.
Probe already spent: $0.026584 (P1, the ten anchors, 10 of 10 correct). Estimate: stage M
10 calls × 2 seats ≈ $0.60; stage G ~10 calls × 3 seats ≈ $0.35; stage C critic 2 calls ≈ $0.15.
Worst case is built from the max_tokens cap each call actually permits, per note (abc), and every
cap is probed on this task shape per note (bsf).