Repository path: workshop/experiments/E-20260827b-worth-paying/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260827b-worth-paying |
| status | frozen |
| created | 2026-08-27 |
| updated | 2026-08-27 |
| senses | style-correspondence, naturalness |
| provisional | true |
| internal-judgment-only | true |
| links | wiki/arms/ARM-worth-paying.md, wiki/findings/results/RS-20260825c-worth-paying.md, workshop/translations/maqamat-hulwan/R43-v1/translation.md, workshop/translations/maqamat-hulwan/R48-v1/translation.md, workshop/translations/maqamat-hulwan/R50-v1/translation.md, workshop/translations/maqamat-hulwan/loci-frozen.md, workshop/regimes/R43-prose-holds.md, workshop/regimes/R48-rhyme-first.md, config/models.md, config/budget.md, wiki/goodness-senses.md, wiki/method-notes.md |
E-20260827b-worth-paying — the both-orders statistic registered in advance, and a source shown rather than told
ARM-worth-paying step 2, T3. VERSION 2, and version 1 was never dispatched. Frozen
2026-08-27 before any data call. The texts it judges were frozen and committed at 86d1136b,
with their translators' logs, before this page existed.
Both pre-run critic seats returned NEEDS-REDESIGN on v1, and one of the two blocking findings
was arithmetic: v1's segment-level primary could not reach significance at any outcome. The
disposition of all sixteen findings is critic-response.md; the three that changed the design are
carried below. v1 dispatched nothing — the critic ran on the design alone, at $0.063360250.
1. What step 1 left owed, in its own words
RS-20260825c-worth-paying §8 limit 1: "The page's headline is a post-hoc statistic invented
after its gate failed… A successor that registers the content-decided statistic in advance is the
first thing owed." Limit 4: "I2 is still an intervention on the prompt… It remains a paragraph
the seat has just read." Limit 5: the run used one maqāma of fifty.
This design owes all three and pays them in one run: the statistic is registered below before any call, on a different maqāma in two renderings that did not exist when step 1's hypotheses were formed, and the disclosure is delivered in two ways — as a sentence about the source and as the source itself.
2. The question
Registered: does a blind reader's content-decided preference between a rhyme-first and a restrained English rendering of the same Arabic move toward the rhyme-first arm when the reader is given the source's form — and does showing the source do what telling about it does?
The second clause is the part nobody has asked. If only the sentence moves preferences, step 1's inversion is a fact about prompts. If the source itself moves them, it is a fact about what a reader does when they can see what the translator was up against.
3. Materials
Two arms, one hand, one text, both frozen before this page:
| arm | id | policy | loci chiming, of 58 | colon-end STRICT |
|---|---|---|---|---|
ORD |
T-maqamat-hulwan-R43-v1 |
restraint — the ordinary word for both bearers | 10 | 2 |
RHY |
T-maqamat-hulwan-R48-v1 |
the rhyme first, delivered at the colon end | 42 | 26 |
DCH |
T-maqamat-hulwan-R48D-v1 |
R49 — RHY with the sound removed at 49 colons and, by script, nothing else |
0 | 0 |
DCH is the placebo arm, added on the critic's blocking finding 4. RHY differs from ORD
both by chiming and by every wording the chime forced; DCH holds the second and drops the first,
so a disclosure effect that appears on RHY vs ORD and not on DCH vs ORD is about the sound.
Its R49 rule-4 audit finds DCH more faithful than RHY at 8 colons, so the placebo carries a
fidelity advantage and any shift toward RHY is measured against it.
BAL (T-maqamat-hulwan-R50-v1, Preston's balanced period) is NOT an arm of this run, and the
reason is measured, not editorial. ORD and BAL share 216 twelve-grams and a 48-token
identical run (R43-v1/translation.md §Contamination). A forced choice between them would be a
near-duplicate control, which is what step 1's F6 was. The two-arm design is what the measurement
leaves standing.
3.1 THE SEGMENT RULE — stated here once; every other mention cites this clause (note (bqb)).
segments.py implements it and nothing else does.
- A segment is a run of consecutive prose cola that does not cross a verse insertion point.
- Inside each maximal prose stretch, cut greedily into chunks of 12 cola; a final chunk under 8 cola is merged into its predecessor if that keeps it at or below 14, and is otherwise dropped.
- A chunk is admitted only if
ORDandRHYdiffer, at the colon-end grade, at 2 or more of the loci it contains. - Segments are labelled
G01… in colon order. All admitted segments are used; there is no selection among them.
Yield: 9 chunks, 9 admitted, 60–116 English words each, 2 to 5 differing loci each. The chunk size was chosen before any data existed by running the rule at 12/8, 10/7, 9/6 and 8/6 and taking the setting with the most admitted segments (9, 6, 7, 6) — a power choice made on segment counts alone, recorded here so it is not mistaken for a choice made on answers.
3.2 The four information conditions. The question put to the seat is identical in all four; only the material above it changes.
| id | what precedes the two passages |
|---|---|
I0 |
nothing |
I1 |
provenance only: "Below are two English translations of the same passage from a work of twelfth-century Arabic literary prose." |
IS |
I1, plus the passage's own Arabic, in the original script and in a romanisation, with no sentence describing it. |
I2 |
I1, plus a statement that the Arabic is rhymed prose (saj'), in which the clauses chime at their ends. No Arabic is shown. |
The "translators have disagreed about carrying it" clause step 1's I2 carried is STRUCK
(critic finding 3, both seats): it is meta-literary framing that IS has no counterpart for, and
keeping it would have confounded evidence-versus-assertion with framing-versus-none. All four
conditions now open with the identical I1 sentence and differ only in what follows it.
IS supplies evidence and no assertion; I2 supplies assertion and no evidence. Step 1's
I2 mixed the two — it asserted the form and printed four transliterated cola — so this I2 is a
purified version of it and figures are not compared with step 1's across designs.
THE ESTIMAND, restated as the critic required and narrower than v1's. IS is longer and more
visually salient than any other condition and invites a reader who can read Arabic to check the
renderings against the source. Those cannot be matched away without inventing filler. What IS
measures is therefore the total effect of putting the source in front of the reader — not a
decomposed "evidence" effect — and every sentence of the result page will say so. The one thing
that can be said about the fidelity-checking route is directional and is registered here: RHY
is the arm that bends words for the chime (its own frozen log lists four leaning words, and the
R49 audit finds eight), so a reader checking against the source should prefer ORD or DCH. A
shift toward RHY under IS runs against that route.
The romanisation is produced by translit.py, a fixed letter table applied by machine with no
per-word exceptions, so that no thumb of the analyst's rests on how alike two colon-ends are spelled.
It is a reading aid and not a scholarly romanisation; the copy-text is only partially vocalised and
the romanisation reproduces that unevenness inside the colon. The colon-ends, which is what the
condition exists to show, are pointed throughout.
3.3 The task. Both passages, labelled A and B, then: "Which of these two do you prefer to read
as a piece of English prose? Answer A, B, or N if you genuinely have no preference." Plus one
sentence of reason. JSON out. N is offered on every call.
4. Design
| block | cells | calls |
|---|---|---|
M, main |
9 segments × 4 conditions (I0,I1,IS,I2) × 2 orders × 3 seats, pair RHY vs ORD |
216 |
D, placebo |
9 segments × 2 conditions (I1,IS) × 2 orders × 3 seats, pair DCH vs ORD |
108 |
Every cell is run in both presentation orders (note (bro) remedy (i)).
Seats, per config/models.md: P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash,
QR qwen/qwen3.7-max — the same three as step 1. P3 is out on cost, P4 on note (bps), P5 on
note (bne), GL on note (brt). Temperature 0.
5. The registered analysis — written before any call, which is the whole point of this step
5.1 THE CODING RULE — stated here once; analysis.py imports it and nothing re-implements it
(note (bqb), and the critic's finding 5).
- A cell is a (block, segment, seat, order) tuple. Within a block and cell, the same two passages are shown in the same physical order in every condition; only the material above them changes. This is what makes the contrast position-proof: the order is identical on both sides of every paired comparison.
- A response is parsed to one of
RHY/ORD/DCH/N/VOID. Scores, in the main block: s = +1 forRHY, 0 forN, −1 forORD. In the placebo block: s = +1 forDCH, 0 forN, −1 forORD. - For a contrast (C1 → C2), the cell's paired difference is d = s(C2) − s(C1) ∈ {−2,−1,0,+1,+2}.
A cell whose response is
VOIDin either condition is excluded, and the excluded count is reported. - The test is an exact binomial sign test on sign(d) over the cells with d ≠ 0, tie cells (d = 0) excluded and counted. One-sided in the direction registered below, with the two-sided tail reported beside it.
- Cells are not independent — six share a segment. So a segment-clustered check is mandatory: the same contrast recomputed on the 9 segment means of d, exact sign test, reported beside the cell-level figure. A conclusion is claimed only where the two agree in sign.
- Per-seat figures are reported for every primary (critic finding 8).
5.2 The primaries.
P1a(told). Cell-level d forI1 → I2in the main block is positive — the assertion moves preference towardRHY.P1b(shown). Cell-level d forI1 → ISin the main block is positive — putting the source in front of the reader moves preference towardRHY.
Holm across P1a and P1b at 0.05. These are reference tails under exchangeability of the
condition label, not Type-I guarantees: the arms are fixed texts, not randomly assigned
treatments (RS-20260825b §2, RS-20260825c §3).
5.3 P3′, the specificity control — the run's identification, and it replaces v1's dose test.
The same I1 → IS contrast computed in the placebo block, DCH vs ORD.
The specificity claim — the shift is about the chime — is made only if all three hold:
(i) P1b is positive and survives Holm; (ii) the placebo's I1 → IS shift toward DCH does
not survive the same test; and (iii) the difference between the two blocks' mean d is in the
registered direction and clears an exact permutation test over the block label, 10,000 draws with a
fixed seed, or exhaustive enumeration where feasible. Any other pattern is reported as it falls
and the specificity claim is not made.
5.4 Registered secondaries, reported in full whatever they say.
Q1— the direct pairedISvsI2contrast in the main block, same machinery. No directional prediction is registered; the interesting outcomes are in both directions. (This replaces v1's untested "P1bnot smaller thanP1a" claim — critic finding 9.)Q0— descriptive only. The mean cell score s atI0, its sign-test tail, and the per-seat breakdown. No prediction, no threshold, no "near even" (critic finding 12).Q2, the both-orders statistic step 1 owed. A segment is content-decided for arm A in condition C when A takes strictly more seat-votes than the other arm in both orders,Ncounting as a vote for neither. The counts per condition are reported for all four conditions. Registered here in advance, as step 1's limit 1 required, and reported as a secondary — the critic showed it cannot carry a primary at nine segments.- The dose contrast (segments with ≥3 differing loci against those with 2) survives only as a labelled descriptive with no decision attached (critic findings 6, 7).
6. Gates, with their consequences fixed in advance
F1— position. Pooled first-position rate over all 324 main-and-placebo calls, reported per condition. Bar [0.40, 0.60]. Consequence:F1failing withholds every pooled preference rate. It does not withholdP1a,P1borP3′, because those are within-cell contrasts in which the presentation order is literally the same on both sides — not because any statistic is "immune to position", a claim v1 made and the critic correctly struck.F1b— per-cell order consistency (note (brs)). For each condition, the fraction of (block, segment, seat) triples returning the same label under the swap;N/Ncounts as consistent,Nagainst an arm as inconsistent. Chance is 0.50. Consequence: a condition at or below 0.50 has itsQ2count reported and labelled as resting on cells that behave at chance. It does not withhold the within-cell primaries, for the reason just given.F2— the operational floor, and it is nothing more (critic finding 11). IntactORDagainst the sameORDwords with the cola deranged, 2 segments × 2 orders × 3 seats = 12 calls. Bar: 10 of 12 choose the intact text. Passing shows the seats can tell English prose from its own scrambling. It is not offered as evidence that they can perceive a chime.F6— the near-duplicate scale, no bar.ORDagainstORDwith one word changed, 2 segments × 2 orders × 3 seats = 12 calls. Its only job is to put a number on the position rate: this is what first-position looks like when there is almost nothing to choose between (note (bro) remedy (ii)).F4— missingness. A cell returning an empty body or unparsable JSON is re-dispatched once at double the cap. Bar: fewer than 10% of cells void after the re-dispatch.F5— repeatability. 18 stratified cells re-run at the end (note (brn)).- No keyword gate on whether the disclosure was "used." Note (brp): a gate that codes the
subject's own vocabulary measures displacement of register, not uptake.
P3′is the only manipulation check this design has, and it is behavioural.
7. Predictions, and what would falsify them
P1aholds — the told condition moves preference towardRHY. This is step 1's finding asked prospectively on new material and a new inferential unit. Falsified by no shift, or a negative one.P1bholds — showing the source moves preference towardRHY. Falsified byISsitting atI1, which would say step 1's effect is a response to being told rather than to what the source is.P3′holds — the placebo does not move. Falsified by the de-chimed arm moving as much, which would say the disclosure is moving something other than the sound.Q0,Q1andQ2carry no predictions and are reported as they fall.
If P1a and P1b both fail, that is the finding and the arm closes on it: the inversion step 1
reported does not replicate prospectively, and framework/v0.2 §7.34 is amended to say so. An
honest null closes this arm as well as a positive does.
8. Stages, cost, and the stops
Note (abc): the worst case is built from the cap the request permits, not from an assumed output length. Note (brt): a cap verified on one task shape does not transfer. Note (brw): a runner that retries must accumulate the cost of every attempt.
| stage | calls | purpose |
|---|---|---|
C |
2 | pre-run critic — SPENT, $0.063360250, and it rebuilt this design |
T |
6 | cap probe — 3 seats × {I0, IS} on one segment, to measure actual completion tokens on this task shape before any fan-out |
F |
12 | F2 operational floor |
X |
12 | F6 near-duplicate scale |
M |
216 | main pair |
D |
108 | placebo pair |
R |
18 | F5 repeatability |
| 374 |
Declared ceiling for the whole run: $2.20, against a UTC-day headroom of $3.252152779 after
S226. Stages M and D are not dispatched until stage T's measured tokens produce an estimate
that fits inside what is left of the ceiling, and that estimate is written into
raw/cap-probe.json before dispatch.
The partial-dispatch rule (critic finding 13): no primary is analysed unless the M and D
blocks both complete after the F4 re-dispatch. A budget shortfall is a deferral to NEXT.md,
not a reduced analysis of whatever was bought.
Caps. T runs at 4000 for all three seats, the doubled figure note (brt) requires on a new
shape. The other stages run at whatever T shows to be sufficient with a 2× margin, floored at 1200.
Every attempt's cost is summed inside the cell record and an attempts count is kept — note
(brw), whose defect cost the previous session $0.206 of unexplained ledger. The dispatcher takes
an exclusive lock before running — note (brf).
9. Verification
verify.py recomputes every reported number from the raw records with no import from analysis.py,
recomputes the reference tails by exhaustive enumeration where the cell count permits it, and is run
against three mutations of the raw data that it must catch.
10. Addendum, written after the run — the one thing added post-hoc
Declared here so the frozen design carries it. The registered placebo (§4, block D) was bought
for the I1→IS contrast. I1→IS came back null, and the contrast that moved was I1→I2,
which therefore had no control. Its control cells — block D at I2, 9 × 2 × 3 = 54 calls,
$0.190531950 — were dispatched after P1a and P1b had been computed, under run.py mode
placebo2, whose docstring says so. They are labelled UNREGISTERED in analysis.json, in the
result page's §4 table and in its §5, and no registered claim rests on them.
Buying a control for the wrong contrast could not have been avoided by care — which contrast would move was the question. What could have been avoided is buying only one, and that is the lesson.