Repository path: workshop/experiments/E-20260826b-radif/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260826b-radif |
| status | frozen |
| created | 2026-08-26 |
| updated | 2026-08-26 |
| senses | style-correspondence |
| provisional | true |
| internal-judgment-only | true |
| links | wiki/arms/ARM-radif.md, workshop/translations/saadi-ghazals-radif/R51-v1/translation.md, workshop/translations/saadi-ghazals-radif/R51-v1/arms.md, workshop/experiments/E-20260826b-radif/critic-response.md, wiki/findings/results/RS-20260824c-run-placement.md, wiki/findings/results/RS-20260821b-matched-heard.md, wiki/goodness-senses.md, wiki/method-notes.md, config/models.md, tools/rhyme_pairs.py |
E-20260826b — a repetition and a chime in one slot, taken apart
Design v2, frozen before dispatch. v1 was frozen, sent to two adversarial critic seats
together with the exact materials, and rebuilt on their findings: fifteen findings, seven
BLOCKING, all accepted, none overruled (critic-response.md; v1 is in git history). The
translation limb and the four line-end arms were frozen and committed at 74cc7c83, before v1
existed. No seat has judged anything at the time this is frozen.
1. The question
wiki/goodness-senses.md §style-correspondence defines itself over marked formal features and
names repetition among them. Every measured statement on the entry is about chime: device
presence (+3.524, S142) and device extent (+0.321 / +0.214, RS-20260824c-run-placement,
S219). Nothing on the entry measures repetition, and the run that came nearest
(RS-20260821b-matched-heard, S209) established that at a matched shape the subtraction method is
not available, closing with the sentence this design answers: what remains available is varying
which resource carries the match with the position held fixed, which no run has yet built.
The Persian ردیف is where that is buildable: two resources in one slot on every line — a phrase repeated identically at every rhyming position, and a chime on the word immediately before it.
Which of the two is a reader responding to, and is the pair worth more than the two halves added up?
2. Materials, and what the manipulation actually is
workshop/translations/saadi-ghazals-radif/R51-v1/arms.md v2. Seven loci, four arms each,
word-identical within a locus except at the line-ends (one declared mid-line normalisation at L8,
applied to all four arms).
| arm | word before the tail | tail |
|---|---|---|
AB |
chimes | identical — the Persian shape |
RH |
chimes | varied, each variant truth-conditionally equivalent in its own line |
RD |
chimes with nothing, repeating exactly where AB repeats |
identical |
NN |
as RD |
as RH |
L1–L4 are depth 3, L6–L8 depth 2. L5 was withdrawn before any judging: at a comparative
radif every English bearer must be a comparative, every English comparative ends in the same
unstressed syllable, and the language will not permit a no-chime arm. That is a result, recorded on
arms.md.
The confirmatory set is L1, L2, L4, L6, L7 — the five loci where the hand and
tools/rhyme_pairs.py agree that the AB bearers chime (≥ 1 STRICT pair). L3 and L8 are a
labelled sensitivity set: their AB bearers are blame/claim and lane/slain, full rhymes to any
English poet, rejected by the tool's rime-riche guard because the phoneme before the rime is L in
both members. Every quantity is reported on the confirmatory five and again on all seven.
What the two factors are NOT. They are not isolated phonological changes. Removing a chime means
changing a word, and un-repeating a phrase means changing words; the substitutes differ in shading,
and a varied tail is longer than is there or arose however carefully chosen. E_rep and
E_chime estimate preference between constructed line-end packages; repetition and chime name
the intended manipulations, not isolated causes. This is the critic's own first remedy for its
BLOCKING 2, taken in full, and it governs every sentence written from this run.
3. Seats
Three, all non-Anthropic: P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash,
QR qwen/qwen3.7-max. P4, P5 are out on any task shape (notes (bps), (bne)); P3 is a
cost problem; GL is out on long prompts. max_tokens 3000 on every seat from the start, per
note (brr).
4. Procedure
Blind forced choice, one pair per call, temperature 0. Each call shows two versions of the same
passage, labelled A and B, and asks which reads better as English verse, with one sentence of
reason. No call mentions Persian, Sa'di, rhyme, repetition, line-endings, translation, or that
anything was manipulated.
Four pairs per locus — the edges of the 2×2: AB–RH · AB–RD · RH–NN · RD–NN.
Both orders at every seat: 7 × 4 × 2 × 3 = 168 judging calls.
5. The statistics, registered
For each locus and each pair let w be the fraction of the six judgments (3 seats × 2 orders) preferring the first-named arm.
| quantity | definition | range |
|---|---|---|
E_rep |
½·[ w(AB>RH) + w(RD>NN) ] − ½ |
[−0.5, +0.5] |
E_chime |
½·[ w(AB>RD) + w(RH>NN) ] − ½ |
[−0.5, +0.5] |
E_int |
w(AB>RH) − w(RD>NN) |
[−1, +1] |
The unit of analysis is the locus. Seats and orders are pooled inside a locus and are never
counted as independent observations (RS-20260809i BLOCKING 3).
Primary reporting is the per-locus effect and its dispersion — median, range, interquartile
range over loci — in the shape RS-20260824c used, because five loci do not support an interval
estimate. The exact one-sided sign test is secondary and is a descriptive consistency summary over
constructed loci, not evidence generalising across texts or readers: the seven loci come from three
ghazals and share radifs. Its floor on the confirmatory five is P = 0.031.
Tie rule, registered: a locus whose effect is exactly 0 counts as not supporting the one-sided prediction and stays in the denominator.
Registered predictions, one-sided:
P1—E_rep> 0. The entry's gap.P2—E_chime> 0. A replication check on the chime package, not a validity gate forP1(critic finding 8).P3—E_int> 0. The Persian form's own claim. Declared underpowered and reported with its dispersion only.
Depth-2 and depth-3 loci are reported separately for every quantity (critic finding 12).
6. Failure and withholding criteria, registered before dispatch
F1—P1's interpretability rests onC1's per-edge gate (§7), not onP2. IfC1fails its planted-error gate, all three quantities are reported descriptively only and nothing is written intowiki/goodness-senses.md.F3— an edge thatC1marksDIFFERENTat a locus is dropped from the quantity it feeds at that locus, and the locus leaves that quantity's denominator.nis reported per quantity.F4— order, per edge. For each of the four edges, the first-position win rate is computed over its 42 judgments. If any edge falls outside [0.35, 0.65], every quantity that edge feeds is recomputed within each order and both reported; if the two orders disagree in the sign of a quantity's median, that quantity is withheld.F5— parse. An empty or unparseable body is re-dispatched once; a second failure drops that judgment, reduces the denominator, and is reported as a count.F6— the arms must be what they claim.tools/rhyme_pairs.pyover the declared bearer lists must return, for every live locus, ≥ 1 relating pair inABand 0 relating pairs inRD. Run before this design was frozen: passes at 7 of 7 (grading-arms-v2.json). Registered here so a later reader can see it was a gate.- Saturation is a diagnostic, not a withholding rule (critic findings 4 and 13). If a pair returns the same preference in ≥ 95% of judgments the fact is reported beside the number.
7. Control C1 — a per-edge equivalence gate
The substitutions are defensible renderings of the same Persian but are not identical in content, and
at six of seven loci the RD bearer is the more literal. A seat preferring RD might be preferring
accuracy.
Procedure. Seat P2, one item per call, temperature 0, shown the two texts of one edge and
asked: Do these two passages assert the same thing? Answer SAME or DIFFERENT. If DIFFERENT,
say in one clause what the second asserts that the first does not, or what it contradicts. The
prompt states explicitly that a difference of wording that is not a difference of assertion is
SAME (critic finding 15).
Items: the four edges × seven loci = 28 real items, plus 6 planted items in which one side has been altered to change a number, change a referent, or drop a clause outright. 34 calls.
Gate. DIFFERENT on ≥ 5 of the 6 planted items. Below that the control is uninformative,
F1 fires, and that is reported rather than papered over.
If C1 fails its gate the judging stage is not dispatched at all.
8. Pre-flight cost
Worst case from max_tokens, not from expected output (note (abc)).
| stage | calls | worst case at list prices |
|---|---|---|
| pre-run critic, 2 seats, cap 9000 | 2 | $0.13 (actual: $0.10866075) |
control C1, P2, cap 3000 |
34 | $0.39 |
| judging, 7 × 4 × 2 × 3, cap 3000 | 168 | $2.39 |
re-dispatch allowance under F5 (10%) |
~20 | $0.29 |
| declared ceiling | $3.20 |
Headroom before this run: $4.783086950 of $5.00 (opening snapshot 144.764653501).
Staged so it can be stopped: critic → C1 → judging.
9. What this cannot establish, written before any number exists
- Nothing about hearing, and nothing spontaneous. This measures explicit comparative preference after a salient line-end substitution. Side-by-side presentation makes the changed words conspicuous however the prompt is worded. No sentence citing this may say a reader hears a repetition or responds to one unprompted.
- The factors are packages, not isolated causes (§2). A varied tail is longer and heavier than an identical one; a de-chimed bearer is a different word with different shading.
- Three poems. Seven loci drawn from three ghazals, two of them sharing a radif with another locus. This is not seven independent texts and the sign test is not a general inference.
n= 5 confirmatory. The one-sided exact test cannot go below P = 0.031, so this run can support a direction, never a strong claim.- Three model seats, no human reader. Tier D is NOT PASSED;
internal-judgment-only,provisional. E_intcannot establish the Persian form's claim at this n and is not reported as if it could.