Repository path: workshop/experiments/E-20260828-forced-half/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260828-forced-half |
| status | frozen |
| created | 2026-08-28 |
| updated | 2026-08-28 |
| senses | style-correspondence |
| provisional | true |
| internal-judgment-only | true |
| links | wiki/findings/results/RS-20260828-forced-half.md, wiki/arms/ARM-radif.md, workshop/experiments/E-20260828-forced-half/critic-response.md, workshop/translations/saadi-ghazals-forced/R54-v1/translation.md, workshop/translations/saadi-ghazals-forced/dependence-note.md, workshop/regimes/R54-forced-pair.md, wiki/findings/results/RS-20260826b-radif.md, wiki/findings/results/RS-20260827b-shown-or-told.md, wiki/findings/results/RS-20260824-eye-or-ear.md, wiki/method-notes.md, config/models.md, config/budget.md |
Which half of the radif does a reader take — and does the source change it?
ARM-radif step 2 (T2). This is v2. v1 was frozen, put to two independent critic seats, and
both returned NEEDS-REDESIGN — 16 findings, 14 accepted, 2 remedies substituted with the reason
written (critic-response.md). v1 is in git history at 9a31299b and no data call was
dispatched under it. The translation limb this runs on (T-saadi-ghazals-forced-R54-v1, six whole
renderings) was frozen and committed at 358b987b before v1 existed.
Tier D is NOT PASSED. These are three model seats, not readers. Every figure this design produces
is provisional and internal-judgment-only, and no sentence of the result page will say that a
human reader hears, prefers or wants anything.
1. What step 1 left owed, in its own words
RS-20260826b-radif §5 reported all three registered predictions unestablished and named the cause: a
forced choice between two passages differing only in three line-end words
"is, on these seats, reading position rather than text… which reads better as English verse gives a seat nothing to be right or wrong about, and position fills the vacuum. That, not a better statistic, is what step 2 has to fix."
ARM-radif §Step 2 records the choice that follows: re-ask the separation with an outcome the seats
can be right or wrong about, or close the arm retired. Two things have changed since.
RS-20260827b-shown-or-told(S227) built the instrument. Its within-cell contrast holds the two passages and their physical order identical on both sides of every comparison and varies only the material above them.- The material now exists. Step 1 could only vary three words inside one text.
R54has produced two complete renderings of record of each of three ghazals, one keeping the repeated tail and one keeping the chime — the choice a translator of any radif ghazal actually faces.
2. The question
At a Persian rhyming position
[…qāfiya][radif]that English cannot carry whole, the translator must keep the repeated tail or the chime and lose the other. Which does a reader take — and does putting the Persian in front of them, or telling them what it does, change which?
Why the second half is the sharp half. RS-20260827b found that telling a seat what the Arabic
does moved its choice and showing it the Arabic did not. So this design asks whether that
asymmetry is a fact about disclosure in general or a fact about the feature disclosed. The feature
there was saj', a rhyme. The feature here is an identically repeated word, the more robustly
recoverable of the two from a text one is merely shown: a rhyme must be reconstructed from the
writing; an identical string need not be.
What IS is, stated as the critic forced it to be stated (critic-response.md B5). v1 called
IS a visual shape condition, on the ground that a reader who cannot read Persian can still see a
repeated string. These seats read Persian. IS therefore supplies the source, semantics
included, exactly as S227's did, and the contrast between the two runs is the feature, not the
reader's access to it. The visual-shape framing is withdrawn and is used nowhere in the analysis.
This is a question about translating literature and about reading translations, not about the
project's instruments (wiki/tracks.md §The subject rule): what a translator of a radif ghazal must
decide is which half to keep, and what this asks is whether a reader with the source in front of them
decides it differently.
3. Materials
Nine windows, three per ghazal, covering all 21 bayts: W1 = bayts 1–3, W2 = bayts 4–5, W3 =
bayts 6–7. Each window is printed in two arms:
| arm | policy | score |
|---|---|---|
REP |
the repeated tail kept at every rhyming position, the chime given up | s = +1 |
CHI |
the chime kept to the source's own depth, the repeated tail given up | s = −1 |
N (no preference) scores 0. Both arms of every window come from R54-v1, unedited, in the same
bayt order.
The three pairs are NOT equally different, and the design obeys that rather than discovering it.
../../translations/saadi-ghazals-forced/dependence-note.md §2, run before v1 was written per note
(bry): the two arms of ۹۴ share 53 seven-word runs and a 19-token run, ۱۲۶ share 19 sevens, ۹۷
share 9. The shared material is the a-hemistichs, which carry no rhyming position. Registered
consequence: every primary is reported per poem as well as pooled, and no conclusion is claimed on a
pooled figure whose three poem-level figures disagree in sign.
4. Design — five conditions, one pair, one physical order per cell
| block | cells | calls |
|---|---|---|
M, main |
9 windows × 5 conditions (I0, I2, IS, IP, IL) × 2 orders × 3 seats |
270 |
A cell is a (window, seat, order) triple. Within a cell the same two passages are shown in the same physical order in all five conditions; only the material above them changes.
Every condition carries the same one-sentence provenance stem, so each disclosure adds exactly one thing and none is confounded with simply learning that the passages are translations of one poem:
stem, in all five: "Below are two English translations of the same poem by the thirteenth-century Persian poet Sa'di."
I0— the stem and nothing else.I2, told. Stem + "In the Persian, every one of these lines ends in the same word, repeated identically; and the word immediately before that repeated word rhymes, in every line, with the corresponding word in every other line. English cannot do both at once, and neither translation below does."IP, the control forI2. Stem + a true statement about the same Persian that is not about the line-ends: "The Persian original is a ghazal — a poem whose couplets are semantically independent of one another and have been read and anthologised singly for seven centuries — written throughout in a single fixed quantitative metre, the same in every line."IS, shown. Stem + the window's own Persian, in script and in a graphemic transliteration, introduced only as "the Persian original of the same lines." Nothing is said about what it does.IL, the control forIS— new inv2, on the critic's BLOCKING 3. Stem + a different Sa'di ghazal (۴۹, which is not in this experiment's material), in script and transliteration, at the same number of bayts as the window, introduced as "a different poem by the same poet, for comparison." It is length-, script- and form-matched toIS, and it is not the source of the passages. ۴۹ is itself a radif ghazal, soILdiscloses the form as fully asISdoes and differs from it in exactly one thing: whether the Persian shown is the original of these lines.
The question put in every condition is the same — which of these two do you prefer to read as English verse? — and that is deliberate. The design does not repair the question; it holds it fixed and varies what the reader knows, so the quantity estimated is the movement the knowledge causes.
Seats, per config/models.md: P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, QR
qwen/qwen3.7-max — the same three as step 1 and S227. P3 is out on cost, P4 on note (bps), P5
on note (bne), GL on note (brt). Temperature 0.
5. The registered analysis — written before any call
5.0 The estimand, narrowed and named (critic BLOCKING 2). The two arms are whole independent renderings, differing in diction, syntax, rhythm and idiom as well as in line-end policy. What is estimated here is therefore the movement in preference between two fixed whole renderings caused by naming or supplying the source's line-end shape — never an isolated device effect. The fixed quality difference between the arms is constant across every contrast below and cancels; what does not cancel is a disclosure × quality interaction, and that is declared here rather than hidden.
5.1 The coding rule, stated once; analysis.py imports it and nothing re-implements it
(note (bqb)).
- A response is parsed to
REP/CHI/N/VOID. s = +1 forREP, 0 forN, −1 forCHI. - For a contrast (C1 → C2), a cell's paired difference is d = s(C2) − s(C1) ∈ {−2,−1,0,+1,+2}. A
cell
VOIDin either condition is excluded, and the excluded count reported. - The inferential unit is the WINDOW MEAN (critic BLOCKING 1 and 4, and
P2's MAJOR 4): d̄_w = mean of d over all (seat, order) cells of window w. Averaging over the two balanced orders cancels a symmetric position shift exactly, including one that varies with prompt length; averaging over seats removes the pseudo-replication of treating 54 deterministic calls as 54 readers. - The registered primary test is an exhaustive sign-flip cluster randomisation test on the nine window means: all 2⁹ = 512 assignments of ± to the nine d̄_w are enumerated, the statistic is the mean of the nine window means, and the two-sided P is the proportion of assignments whose |statistic| is at least the observed. The exact binomial sign test on the nine window signs is reported beside it, two-sided.
- The cell-level figures are DESCRIPTIVE and carry no test (critic BLOCKING 4).
- The order split is reported for every primary (critic BLOCKING 1): d̄_w recomputed on
REP-first cells and onCHI-first cells separately. A conclusion is withheld where the two orders disagree in the sign of the pooled statistic. - Per-poem and per-seat figures are reported for every primary (§3).
5.2 The confirmatory set — four contrasts, all predefined, none selected after the fact.
| id | contrast | prediction |
|---|---|---|
P1a |
I0 → I2 |
two-sided. Being told what the Persian does at the line-ends moves the choice. |
P1b |
IL → IS |
two-sided. Being shown the source of these very lines, rather than an equally long, equally Persian, equally radif-bearing poem that is not their source, moves the choice. |
P2a |
IP → I2 |
two-sided. The line-end statement moves the choice more than a true non-line-end statement of similar length. |
P2b |
I0 → IP |
two-sided. Any true statement about the Persian moves the choice. This is the one whose null is wanted; a movement here says the effect is disclosure in general. |
Holm across all four at 0.05. These are reference tails under exchangeability of the condition
label, not Type-I guarantees: the arms are fixed texts, not randomly assigned treatments
(RS-20260825b §2).
All four are two-sided (critic MAJOR 1). v1 registered P1a and P1b as positive — toward
REP — on the argument that the repeated tail is checkable against the English page and the chime is
not. P2 supplied the opposite prior with the texts in front of it: told about the form, a seat may
penalise REP for its unrhymed, weak-ended repetitions. Both stories are stories. The
direction is the finding, and §7 writes out what each direction would mean.
5.3 The specificity claim. What moved the choice is the line-end information about these lines,
not the fact of being told or shown something Persian — claimed only if P1a survives Holm and
P2b does not, and P2a survives Holm in the same direction as P1a. Any other pattern is
reported as it falls and the claim is not made.
5.4 Registered secondaries, reported in full whatever they say.
Q0— descriptive only. Mean cell score s atI0, per seat and per poem. Confounded by construction, and the confound is named here rather than after the fact:R54-v1§7.1 records that theREPpolicy forces an unstressed line-ending at every position in ۹۷ and ۱۲۶ whileCHIends every line on a stressed rhyme. An untold reader is never comparing repetition against chime alone, which is whyQ0carries no prediction and why every primary is a movement, not a level.Q1— the direct pairedISvsI2contrast, same machinery, no prediction. This is S227's shown-versus-told comparison asked on a different feature.Q2, content-decided counts. A window is content-decided for an arm in a condition when that arm takes strictly more seat-votes than the other in both orders,Ncounting for neither. Reported per condition; at nine windows it cannot carry a primary.
6. Gates, with their consequences fixed in advance
F1— position. Pooled first-position rate over all 270 main calls, per condition. Bar [0.40, 0.60]. Consequence: failing withholds every pooled preference rate,Q0included. It does not withhold the primaries, because those are within-cell contrasts averaged over both orders.F1b— per-cell order consistency (note (brs)). Per condition, the fraction of (window, seat) pairs returning the same label under the order swap;N/Ncounts as consistent,Nagainst an arm as inconsistent. Chance is 0.50. Consequence: a condition at or below 0.50 has itsQ0andQ2figures labelled as resting on cells that behave at chance.C1— the sense-equivalence gate, seatP2, 15 items. The nine window pairs plus six planted content errors (one arm altered so it asserts something the other does not; the six are named inrun.pyPLANTSand are checked by assertion before dispatch). Bar: at least 5 of 6 planted caught. Consequence, revised on the critic's MAJOR 4: aDIFFERENTverdict on a real pair does NOT remove a window. Every primary is reported twice — with and without the flagged windows — and a conclusion is claimed only where the two agree. The gate's information is kept and the researcher's freedom to drop inconvenient windows is closed.F2— the operational floor, and nothing more. An intactREPwindow against the same lines re-ordered by a fixed seed: 2 windows × 2 orders × 3 seats = 12 calls. Bar: 10 of 12 choose the intact text. It is not evidence that the seats can perceive a chime or a repetition.F3— missingness. A cell with an absent body is re-dispatched once at double the cap; a present-but-malformed body is re-parsed, not re-dispatched (note (brx)). Bar: fewer than 10% of cells void.F4— repeatability. 18 stratified cells re-run at the end (note (brn)); the agreement rate is reported and carries no bar.- No keyword gate on whether the disclosure was "used." Note (brp): coding the subject's own
vocabulary measures displacement of register, not uptake.
P2a/P2bare this design's only manipulation check and they are behavioural.
7. Predictions, and what each outcome would mean
P1amoves towardREP— told that the source repeats a word and rhymes before it, a reader reaches for the rendering that keeps the repetition.P1amoves towardCHI— a reader told the source is formally elaborate reaches for the rendering that sounds elaborate in English, and penalises a repetition that reads as flatness. Both are substantive; the design is now two-sided precisely because the second was argued as well as the first.P1bmoves — the source of these lines does something an equally Persian, equally radif- bearing decoy does not.P1bdoes not move — then S227's shown-does-not-move result generalises beyond an audible feature, and what moves a reader is being told, not being shown.P2bdoes not move, andP2adoes — the specificity claim in §5.3 is available.P2bmoves — then any true statement about the Persian moves the choice, and nothing here is about the line-ends.Q0,Q1,Q2carry no predictions and are reported as they fall.
On what a non-significant result means (critic MINOR, accepted). At nine clusters this design can detect only a large and consistent movement. A contrast that does not survive Holm is reported as unestablished, never as no effect, and the arm-closing sentence — if it comes to that — says what this project cannot establish with these seats and this much material, not what is not there.
8. Stages, cost, and the stops
Note (abc): the worst case is built from the cap the request permits. Note (brt): a cap
verified on one task shape does not transfer — a cap probe runs first on this task shape with the
longest prompt (IS on the longest window) included, and the main run's caps are set from it. Note
(brw): cost is accumulated across every attempt. Note (brx): at temperature 0 a
present-but-malformed body is re-parsed, not re-dispatched.
| stage | calls | worst case | stop |
|---|---|---|---|
C pre-run critic, P1 + P2, cap 12000 — SPENT, $0.071679500, both NEEDS-REDESIGN |
2 | — | applied before any data call |
T cap probe, I0 + IS, 3 seats |
6 | $0.060 | caps written to raw/caps.json |
C1 sense-equivalence gate, P2, 15 items |
15 | $0.120 | recorded before the main run |
F2 floor |
12 | $0.090 | reported whatever it says |
M main, 5 conditions |
270 | $2.000 | — |
F4 repeat |
18 | $0.140 | — |
| total, critic included | 323 | $2.480 | declared ceiling $2.60 |
UTC day 2026-08-28 had no rows before this session: the whole $5.00 was available and $2.60 is
claimed. Actuals are recorded from "usage": {"include": true} and cross-checked against the
key-usage delta.
The stop that matters. If the probe shows a seat truncating on the IS prompt, that seat's cap
is raised and the worst case recomputed before the main run, not after.
9. Verification
verify.py recomputes every number the result page reports, from raw/*.json, importing nothing from
analysis.py. It re-derives the randomisation tails by exhaustive enumeration of all 512 sign
assignments rather than by sampling, recomputes the exact binomial tails by enumeration, recounts void
and tie cells, re-checks that every cell's five conditions were shown in the same physical order, and
re-checks that the six planted C1 items are the six the design named. Three mutation tests: a
flipped score sign, a swapped condition label, and a deleted void marker; each must be caught.