Repository path: workshop/experiments/E-20260828-purchased-figure/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260828-purchased-figure |
| status | frozen |
| created | 2026-08-28 |
| updated | 2026-08-28 |
| senses | accuracy, style-correspondence |
| provisional | true |
| track | T1 |
| links | wiki/arms/ARM-gulistan.md, workshop/translations/gulistan-bab2/R05-v1/translation.md, workshop/translations/gulistan-bab2/collation-chapter.md, wiki/base/anchors/A-gulistan-hands/README.md, wiki/findings/results/RS-20260822c-persian-hands.md, config/models.md, framework/v0.2/README.md |
E-20260828 — where a rhyme gets paid for
v2, frozen after the pre-run critic and before any judging call. ARM-gulistan step 1's
study limb. Translation limb: T-gulistan-bab2-R05-v1 span B, frozen and committed at b4daef0f
with its log before v1 of this design was written and before any English of the chapter was
opened (charter A4). The critic's findings and their disposition are §10; v1 is in the git history
at b4daef0f's successor commit and is not reproduced here.
1. The wire
The translating produced the question. Rendering 39 of Sa'di's bayts under a rule that forbids
buying a figure with a word the source has not got (V2), the translator refused six rhymes, and
every one of the six was refusable for the same reason: the English rhyme was one small added word
away, and the word would have stood at the end of the line (D24) — astray, at all, slain,
door, gown, hall. Nothing was refused for want of a rhyme; the rhymes were there.
That is a claim about where the cost of a formal constraint falls, and it is testable on a hand
that took the rhymes. Eastwick 1852 rhymes 38 of 42 bayts in this book (RS-20260822c, on
باب اول), and the rhyme coding below finds 17 of his 20 sampled line ends chiming with another
line end in the same stanza, against 0 of 20 for the unrhymed control.
PR-PRIMARY, in one sentence: in Eastwick's 1852 verse rendering of these ten tales, the word
at the end of a line is more often a word with no counterpart in the Persian than a word from the
middle of the same line is.
2. What this is not
- It is not a claim that Eastwick's rendering is bad. A verse translator who has undertaken to rhyme must put something in the rhyme position, and supplying it is the craft. The measurement is of where the constraint is paid for.
- It is not causal. The design does not randomise rhyme; it compares positions and hands. Both
critic seats pressed this and both are right: a positive
PR1licenses an association between line-final position and supplied words in this hand, and nothing about rhyme causing it, except so far asPR2's control arm carries that weight. - It is not a test of the lead against a published hand. The lead's arm is here because it is a rendering of the same 39 bayts under a policy that refuses the purchase — and its END−MID gap is partly entailed by that policy, which is why the critic struck it from the primaries (§10, A3).
3. Materials
| Persian | «گلستان» باب دوم حکایات ۱۱–۲۰, the 39 bayts of blocks b073–b155, from source-ganjoor-bab2-whole.txt (single witness, declared in collation-chapter.md) |
EAS |
Eastwick 1852 (2nd ed. 1880), archive.org gulistanorrosega00sadiuoft, chapter II stories XI–XX — 24 verse blocks, 76 lines, 70 eligible. Blocks found mechanically by Eastwick's own printed labels; OCR repair and prose truncation done by hand and declared in materials/verse_blocks.py; raw OCR kept in materials/eas_verse_raw.json |
UNR |
the control the critic required. The same 39 bayts rendered as two lines of unrhymed English verse each by x-ai/grok-4.5 — told not to rhyme and told nothing whatever about fidelity — make_unr.py, 39 calls, $0.150823200, 0 void. Deliberately not one of the three judging seats, so no model judges its own output. 78 lines, 73 eligible |
LEAD |
T-gulistan-bab2-R05-v1, span B — 79 lines, 68 eligible, frozen at b4daef0f. Descriptive only |
| excluded | Arnold 1899. His scan wraps nearly every verse line onto two OCR lines and his italics survive only as stray marks. The line end is this experiment's manipulation, so recovering Arnold's line ends would put lead reconstruction inside the manipulation. Gladwin 1806 and Ross 1823 print Sa'di's verse as prose and have no line ends at all |
Rhyme coding, mechanical, from tools/rhyme_pairs.py and its CMU dictionary, computed on the
sampled lines before any judging call: a line end is coded by its best relation to any other line
end in the same block.
| hand | STRICT |
NEAR |
NONE |
|---|---|---|---|
EAS |
17 | 0 | 3 |
UNR |
0 | 2 | 18 |
LEAD |
2 | 2 | 8 |
UNR is verifiably unrhymed and EAS is verifiably rhymed; the arm contrast is what it says it is.
LEAD's four are D22's and D23's — the rhymes that fell out of the plain sentence.
4. Arms — positions, not texts
Every item is one English word set against the Persian bayt(s) its block renders, and nothing else. The seat cannot see the arm, the hand, the line, or the hypothesis.
| arm | what the word is | n |
|---|---|---|
END |
the final word of an English verse line | 20 EAS · 20 UNR · 12 LEAD |
MID |
the content word nearest the midpoint of the same line — paired, not independently sampled | 20 · 20 · 12 |
DECOY |
a content word from a different block of the same hand, set against this block's Persian | 8 per hand |
CTRL-ANCH |
a known-answer item: an English word that renders a thing the bayt names outright (camel for بُختی, stone for سنگ, muezzin for مؤذّن …) | 12 |
Selection is mechanical and seeded (build_items.py, seed 20260828). A line is eligible only if
its final token and at least one interior token pass the same content filter, and an eligible
line contributes one END and one MID item — so the two arms are drawn from an identical line
set and cannot differ in lexical class by construction. 140 items, one shuffled order.
5. Procedure
Three seats, per config/models.md — P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash,
QR qwen/qwen3.7-max (P3 is the UNR translator and does not judge; P4 and P5 are out on notes
(bps) and (bne)). One call per item per seat, no batching, temperature 0. Each seat is asked, for
one Persian bayt and one English word:
Does the English word render something that is actually present in this Persian — a word, an image, an action, a thing named, a quality, a relation? … If you cannot read the Persian well enough to say, answer
UNSURE— that is a real answer and it is better than a guess.
Item verdict = majority of the three seats; a three-way split or any tie is UNSURE.
Cap probe first, note (brt): the six longest prompts, three seats, 18 calls.
Acceptance rule, registered (critic A7): all 18 bodies must parse and none may carry
finish_reason == "length". If either fails, the cap is doubled and the probe re-run once; a second
failure stops the run before the main block.
6. Registered primaries and predictions
Both two-sided, alpha 0.05, Holm across the two (critic A4). Significance by exact permutation over the arm labels within hand, 20,000 relabelings, seed 20260828, reported with the risk difference and its 95% interval. The test is associational, not causal (§2).
PR1—P(ADDED | END, EAS) − P(ADDED | MID, EAS) ≠ 0, paired within line. Predicted positive, onD24.PR2—[P(ADDED|END,EAS) − P(ADDED|MID,EAS)] − [P(ADDED|END,UNR) − P(ADDED|MID,UNR)] ≠ 0. Predicted positive.UNRis English verse with line ends, line-final stress, and clause closure — everythingEAS's line ends have except the rhyme — and it was told nothing about fidelity, so its gap is free to be anything. This is the primary that can fail informatively: ifUNRshows the same gap, the effect is about line ends in English verse and not about rhyme, andD24's claim is much weaker than it reads.
Reported, not registered as a primary: the LEAD gap, and the overall ADDED rate by hand.
The LEAD gap is declared partly entailed — V2 forbade the lead to supply an unanchored word
anywhere, so a small gap there is policy, not evidence (critic A3). It is reported because it is the
wire back to the translation limb, and it is interpreted only as a check that the policy was kept.
7. Gates — what withholds the primaries
G1(can the seats read the Persian at all).CTRL-ANCHitems codedANCHOREDat < 0.75 (fewer than 9 of 12) → both primaries withheld. This is the real calibration: known answers, blinded among the study items.G2(ceiling diagnostic). PooledDECOYitems codedADDEDat < 0.55 (fewer than 14 of 24… 13 of 24, the rule being strictly below 0.55) → both primaries withheld. The bar is low and the reason is written: decoys are drawn from the same book and the same ten tales, so a decoy word is "not in this Persian" only in the sense that no hand used it to render this block, and shared vocabulary will anchor some honestly.G2is a ceiling on the instrument, not proof of it;G1is the proof.G3(three-wayUNSURE).UNSUREabove 0.25 of non-control item verdicts → withheld.G4(saturation). Non-controlADDEDrate pooled at ≥ 0.95 or ≤ 0.05 → withheld.G5(floor). Any of the six hand × arm cells below 10 complete items, or fewer than 132 of 140 items complete → both primaries withheld.
A gate that fires is the result. No primary is reported past a fired gate, and no number from a withheld primary may be cited anywhere in this repository.
8. Budget
Ceiling raised from $1.60 to $2.40, and the reason is the critic: two BLOCKING findings
required a third translation arm (UNR, $0.150823200 spent) and forty more judging items.
Spent so far: critic $0.093254750 (C2 re-dispatched once at a larger cap after its first body
truncated), UNR $0.150823200. Remaining plan: 18 probe calls, then 140 × 3 = 420 judging
calls. Worst case built from the cap actually set and not from an assumed output length, note
(abc). Headroom on the day at design time: $3.769373050.
9. Limitations, written before the run
- Not causal, and one published hand.
PR1is a statement about Eastwick 1852's rendering of باب دوم حکایات ۱۱–۲۰ and nothing wider. Arnold's exclusion (§3) is what costs the generalisation; span C can pay it back from page images. - Alignment is at block level, not hemistich level. A word absent from its own hemistich but
present elsewhere in the same block codes
ANCHORED. This makesADDEDa stronger claim and depresses all arms equally. - The
UNRcontrol is a model, not a person. It is the right control for line-end convention without rhyme, and it is not evidence about what human unrhymed verse translators do. - The design is not disinterested. The lead produced two of the three arms and the hypothesis.
- Single-witness copy-text, inherited and declared.
G1is 12 items. A seat that reads Persian badly but is right about camel and stone will pass it. It is a floor, not a certificate.
10. The pre-run critic, and what was done with it
Two seats, both adversarial, both given the frozen v1 design, the item builder verbatim, eight items
as the seats would see them, and D24. C1 openai/gpt-5.6-terra returned NEEDS-REDESIGN, 7
findings, 2 BLOCKING. C2 google/gemini-3.6-flash returned NEEDS-REDESIGN, 5 findings, 2
BLOCKING. Raw bodies: raw/critic.json. Eleven of the twelve findings are accepted; one is
overruled.
| # | seat | severity | finding | disposition |
|---|---|---|---|---|
| A1 | C1 | BLOCKING | END is not a verified rhyme position, and line-final words differ from interior words in closure, stress and information structure — MID controls none of it |
accepted. Line ends are now rhyme-coded mechanically (§3) and a third arm, UNR, supplies unrhymed English verse line ends (§3, PR2) |
| A2 | C1 | BLOCKING | PR2 confounded by hand and material; END and MID independently sampled rather than paired |
accepted. Arms are now paired within line, from an identical eligible-line set |
| A3 | C2 | BLOCKING ×2 | The LEAD gap is algebraically near-entailed by V2, so the old PR2 restated PR1 |
accepted. LEAD is struck from the primaries and reported as declared-entailed; UNR takes its place |
| A4 | C1 | MAJOR | No alpha, no multiplicity, and permutation is associational not causal | accepted. Alpha 0.05, Holm, risk differences with intervals, and §2's causal disclaimer |
| A5 | C1 | MAJOR | G2 (item-level UNSURE) cannot catch seats that are confidently wrong |
accepted. CTRL-ANCH, twelve known-answer items, is now G1 |
| A6 | C2 | BLOCKING | The decoy exclusion used substring matching, so short words were rejected wherever their letters occurred inside longer ones | accepted. Token-set membership |
| A7 | C1 | MAJOR | The cap probe has no acceptance rule and cannot withhold anything | accepted. §5 |
| A8 | C1 | MAJOR | Decoys are not ADDED by construction, so the old 0.60 threshold has no validated meaning, and it fires only at ≤14/24 |
accepted. G2 reframed as a ceiling diagnostic, threshold and discreteness stated |
| A9 | C1/C2 | MAJOR/MINOR | Even if everything comes out as predicted, only a narrow associational claim is licensed | accepted. §1's PR-PRIMARY and §2 rewritten to the narrow claim |
| A10 | C2 | MAJOR | The content filter treated END and MID asymmetrically |
accepted, by the same change as A2 |
| A11 | C2 | MAJOR | Draw decoys from outside باب دوم, because these ten dervish tales share moral vocabulary | OVERRULED. C1's A8 says the opposite about the same set — that decoys already differ too much from END/MID candidates — and an out-of-corpus decoy would differ further in register and topic, which is the defect C1 names. The substance of C2's worry is met by G1, which does the calibration with known answers instead of with decoys, and by G2's lowered, explicitly diagnostic threshold |
| A12 | C1 | BLOCKING (the fix) | Obtain translations under randomised rhyme-required / rhyme-forbidden instructions from the same translators | noted as the successor design, not adopted. It cannot be done to a hand that died in 1883, and UNR is the nearest thing this run can build. Recorded for span C |