Repository path: workshop/experiments/E-20260825c-worth-paying/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260825c-worth-paying |
| status | frozen |
| created | 2026-08-25 |
| updated | 2026-08-25 |
| senses | style-correspondence, literary-quality, voice |
| provisional | true |
| links | wiki/arms/ARM-worth-paying.md, wiki/findings/results/RS-20260825b-flippancy.md, workshop/translations/maqamat-sanaa/R50-v1/translation.md, workshop/translations/maqamat-sanaa/R48-v1/translation.md, workshop/translations/maqamat-sanaa/R43-v1/translation.md, workshop/regimes/R50-balanced-periods.md, wiki/base/anchors/A-hariri-hands/README.md, framework/v0.2/README.md, config/models.md, config/budget.md |
E-20260825c — is the price worth paying? Three real policies on one Arabic text, and what a blind seat chooses
Frozen 2026-08-25 (S222) before any call was dispatched, and AMENDED after the pre-run critic and
before any data call. Both critic seats returned NEEDS-REDESIGN; fifteen findings were
accepted and the amendments are marked [An] below and argued in critic-response.md. Three
were accepted in part and three overruled in writing. ARM-worth-paying step 1, T3.
1. The question, and why it is not a question about this project
framework/v0.2 §7.33, written yesterday, ends with a sentence that names its own gap:
A translator who decides to carry a source's prose rhyme into English is buying a lighter, less graceful, more self-displaying English … Whether the price is worth paying §7 does not say, and does not have the evidence to say.
A price is not a verdict. RS-20260825b measured what a rhyme-forward rendering costs on three
rated attributes; nobody has asked whether a reader, offered the two texts, takes the cheaper
one. And there is a second gap, older and larger: Preston 1850 did not merely refuse the rhyme.
In the same sentence he printed what he would put in its place — clauses "though not rhyming
together … arranged as far as possible in evenly balanced periods, and never exceed a certain
length" — and that positive prescription has never been instantiated by anybody, so it has never
been possible to ask whether his remedy is better than doing nothing in particular.
The subject-rule sentence (wiki/tracks.md, continue-prompt.md §4.5): this unit teaches
whether a measured stylistic cost is one a reader will actually pay, and whether knowing what the
original does changes what a reader wants — which is the question a translator faces at every
ornamented sentence and the question §7.33 declines to answer. It is about translations and their
readers, not about this project's instruments.
2. Materials — three policies, one hand, one text, all frozen before this page
All three are complete renderings of al-Ḥarīrī's first Assembly from the same collated copy-text, by the lead, at $0, each with a translator's log frozen before any evaluation of it was designed.
| arm | translation | policy |
|---|---|---|
ORD |
T-maqamat-sanaa-R43-v1 |
restraint — take the ordinary rendering of both rhyme-bearers; keep a chime only if it comes free. 2 full colon-end rhymes. |
RHY |
T-maqamat-sanaa-R48-v1 |
the rhyme first — deliver a chime at every colon end the sense permits. 33 full colon-end rhymes. |
BAL |
T-maqamat-sanaa-R50-v1 |
Preston's own remedy — no rhyme at all, members balanced in syllable count and shape, no colon over fourteen words. 0 full rhymes, 5 near, 0 identical bearers; mean |Δ syllables| across the 69 saj' pairs 2.174 against ORD's 3.000 and RHY's 3.290. |
BAL was translated this session and committed at 8d819873 before this design existed.
Contamination measured clean against all three published English hands (longest run 9 / 5 / 11
tokens; 0 twelve-grams), and its overlap with the lead's own two other arms is declared on its page:
17 twelve-grams with ORD and 2 with RHY, against the 168 those two share with each other.
The estimand, stated in the critic's own words [A10]. What this run estimates is preference among these three particular translations — three policies as executed by one hand on one text. It is not an estimate of the effect of a rhyme, and no sentence of this design or of its result page claims otherwise. The arms differ in the policy and in everything the policy dragged with it: diction, syntax, compression, imagery. That bundle is what a translator actually chooses between, and it is also why nothing here isolates a cause.
Segments [A2, A14]. Cut at the spans E-20260825b already cut and used, so nothing about the
choice of passage is new here — and this is therefore a within-material follow-up on texts
whose earlier results informed these hypotheses, not independent confirmation of anything. The H
stage uses twelve: G1–G7 (p034–094, the sermon) and L1, L2, L3, L5, L6 (p096–135,
the swindle). L4 and L7 are held out so that a successor step has out-of-sample material.
B and V use eight: G1, G3, G5, G7, L1, L3, L5, L6. Verse blocks are identical in
all four arms of this text and are excluded.
Seats. P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, QR qwen/qwen3.7-max —
the three that carried E-20260825b at $0.616. They are seats, not readers [A15], and no
sentence of the result page will call them readers. What they are blind to is the question, not
the rhyme [accepted in part, P1 15]: RHY's chimes are audible to anybody, and the design never
supposed otherwise. No seat is told that rhyme is at issue, that a project is behind the question,
or that the passages share a source, except as the information condition states.
3. Procedure
The task. Each call shows two passages, A and B, and asks one question: which of these
two would you rather have read? The seat returns strict JSON: choice (A/B/N, no
preference [A7]), confidence (1–7), and reason (≤ 25 words). Nothing is scored on a scale;
nothing is rated. This is the one question RS-20260825b's instrument cannot ask, and it is why
the critic's proposal to replace it with continuous ratings was overruled: those ratings, on these
exact materials, already exist.
Three pairs. RHY–ORD (the headline: the rhyme against restraint) · RHY–BAL (the rhyme
against Preston's own remedy) · BAL–ORD (Preston's remedy against restraint, where neither
arm rhymes).
Three information conditions, identical in every other word of the prompt:
I0none. "Below are two passages of English prose." No mention of translation.I1provenance. "…two English translations of the same passage from a classic work of Arabic literature, al-Ḥarīrī's Assemblies (Basra, c. 1100)." Nothing about form.I2form disclosed [A3].I1, plus the fact stated flatly, four transliterated cola of the source showing the chime, and one sentence that is both true and symmetric between the arms: English translators have disagreed about this feature: some have tried to reproduce it, others have deliberately declined to. The first draft of this condition put "RHYMED PROSE" in capitals and ran the chime through "the whole work"; both critic seats called it an advertisement, and they were right. The disclosure now names the disagreement, not the desideratum.
I1 is the placebo that matters. Without it, any shift under I2 could be the effect of learning
that one is reading a translation of a monument, which is not the same fact as learning that the
monument rhymes.
Order. Every (segment, pair, condition, seat) cell is run twice, once with each arm first. Position bias is therefore balanced by construction and measurable directly.
| stage | what | calls |
|---|---|---|
C |
independent pre-run critic, 2 seats, no data — spent, $0.09754525 | 2 |
F |
floor control × 4 segments × 2 orders × 3 seats — GATE, run and read first | 24 |
F6 |
near-duplicate control × 4 segments × 2 orders × 3 seats [A8] | 24 |
T |
pilot: one segment, one pair, three conditions, three seats, one order | 9 |
H |
RHY–ORD × I0,I1,I2 × 12 segments × 2 orders × 3 seats [A2] |
216 |
B |
BAL–ORD × I0,I2 × 8 × 2 × 3 |
96 |
V |
RHY–BAL × I0,I2 × 8 × 2 × 3 |
96 |
R |
repeatability: 6 H cells re-dispatched once, 3 seats |
18 |
| total | 485 |
Pre-flight cost [A12]. The first draft asserted $0.0090 per call as a worst case and did not
derive it; the critic caught that the resulting total exceeded the ceiling it was declared under.
The worst case is built from max_tokens (note (abc)) at the panel's list prices, with ~900 input
tokens per call: P1 $0.0045 (600 out at $6.00/M), P2 $0.0029 (600 at $3.75/M), QR
$0.0066 (1200 at $4.425/M) — mean $0.00467. 483 data calls × $0.00467 = $2.26, plus the
critic's $0.0975 already spent = $2.36 absolute worst case; the expectation, from
E-20260825b's realised $0.002472 per scored call, is ≈ $1.20. Declared ceiling: $2.60.
Today's UTC headroom is $3.788902875, of which $0.0975 is now spent.
Hard stage stops, declared instead of an improvised cut [A12]. The stages run in the table's
order and billed cost is summed after each: > $0.40 after the controls and pilot → abort ·
> $1.60 after H → V is dropped and Q1 reported as not run · > $2.10 after B → V is
dropped · > $2.45 after V → R is dropped. Nothing else is cut mid-stage, and any drop is
named on the result page. The stops, not the max_tokens arithmetic, are what enforces the
ceiling — see the repair below, which pushes the theoretical worst case above it.
Two things fixed at the F gate, before any main-stage call — recorded rather than done quietly
The token caps were set too low and the gate caught it. At max_tokens 600/600/1200, 2 of the
24 floor calls came back finish_reason: length — P2 burns hidden reasoning inside the cap
(config/models.md records ~1.7k tokens of it), and E-20260825b had used 1200 for it. Raised to
P1 800 · P2 1400 · QR 1600, the two dead calls re-dispatched and both parsed. This changes
no prompt, no material, no analysis rule and no gate; it raises the theoretical worst case to
$0.00667 per call, $3.22 over 483, which is above the $2.60 ceiling and is why the stage stops
above exist and are checked. The realised rate over the first 48 calls is $0.00189.
F2 passed at ceiling: intact prose chosen over the same words shuffled in 24 of 24 calls, bar
20. The seats can tell composed English from broken English.
4. What is registered
The unit [A1]. The first draft made the unit a (segment, order) cell and treated sixteen of them
as independent. Both critic seats rejected that, and they are right: two orders of one segment share
the same text, and three fixed model seats are not sixteen readers. The unit is the SEGMENT,
decided by the majority of its six calls (3 seats × 2 orders); an N counts toward neither side
and a 3–3 split is a segment-level tie. So H has n = 12 segments and B and V have n = 8.
The test. Exact enumeration over the segments under exchangeability of the arm label —
every assignment of signs to the non-tied segments, one-sided in the registered direction. As in
RS-20260825b §2, and for the critic's reason there: the arms are fixed texts, not randomly
assigned treatments, so what this yields is a reference tail, not a Type-I error guarantee,
and it is labelled that way everywhere it appears. Every headline number is reported with its
per-seat and per-order breakdown beside it, and the call-level counts are published whole.
P1— the primary. On pairRHY–ORD, the count of segments preferringRHYis higher underI2than underI0. Tested on the segments that change between the two conditions (ORD→RHYcounts +1,RHY→ORDcounts −1; unchanged and tied segments are dropped and their number named). Direction registered: disclosure moves the seats toward the rhyme. A shift the other way is reportable and is not the registered hypothesis.P2— the decomposition [A5]. All three contrasts are registered and reported with their counts:I1 − I0,I2 − I1,I2 − I0. No causal attribution is drawn from a failure to rejectI2 − I1— the first draft's sentence assigning the effect to "provenance rather than form" on that basis is struck. What is reported is the three shifts and their sizes; whereI1 − I0is large andI2 − I1is not, the page says that and stops.P3— the specificity control, with an equivalence margin [A4]. The sameI2-vs-I0test on pairBAL–ORD, where neither arm rhymes.P1stands as rhyme-specific only ifP3's shift is less than half ofP1's (segment counts, same scale). At or above half,P1is withheld — a non-significantP3is not by itself evidence of specificity, which is what the first draft wrongly assumed.Q0— the baseline, secondary. UnderI0, onRHY–ORD: isRHYpreferred in fewer than half of the non-tied segments? This is Preston's implicit prediction — the price shows up as dispreference.Q1— the remedy against the rhyme, secondary.RHY–BALunderI0and underI2, reported separately.Q2— a Preston-like prescription, tested as a modern consequence [A11].BAL–ORDunderI0: does the balanced rendering beat ordinary restrained prose when the seat knows nothing? Preston described a practice; he did not predict that it would win a blind preference, so a result either way is a fact about the prescription's modern consequence and not a verdict on his claim, and the result page will say so in those words.- Holm across
Q0,Q1(I0),Q2.P1stands alone;P2andP3qualify it and are not corrected.
Power, stated rather than assumed [A2, and P1 11 / P2 5 accepted]. At n = 12 an exact
one-sided reference tail can reach 0.0002 but needs a near-unanimous pattern; a null on P1 at
this n is uninformative and will be reported as uninformative, not as evidence of no effect.
Reported, not tested: confidence means per arm and condition; N rates; the free-text reasons,
read and quoted.
5. Gates and failure criteria — every one fixed here
F1position bias [A9]. The share choosing the passage shown first, reported overall, and separately by pair and by condition — an aggregate inside the band can hide a condition-by-order interaction that manufactures the primary. Bar: overall within [0.40, 0.60] and the first-position rate underI2differing from that underI0on theHstage by less thanP1's own effect. Either failing,P1is withheld and every figure is reported with its order breakdown.F2the floor. Four segments (G2,G6,L2,L4),ORDintact againstORDwith its cola shuffled by a fixed seed into a full derangement — every word the same, the order of the clauses destroyed. Both orders, three seats,I0. Bar: intact chosen in ≥ 20 of 24 calls. Below it, the seats cannot tell composed English prose from broken English prose and no preference figure on this page is reportable. Run and read before anything else.F6the near-duplicate, which is the control that matters [A8]. The critic's point is that a forced choice manufactures discrimination; a floor made of scrambled prose cannot test that. Four segments (G4,G7,L1,L5),ORDagainstORDwith one word changed — a single near-synonym substitution, listed in the run script before dispatch. Both orders, three seats,I0. Reported, and it conditions everything: if a near-identical pair produces a decisive majority in either direction, or anNrate near zero, the instrument is inventing preferences and every figure on the page is read in that light. Bar for a clean reading: the majority direction is split across the four segments, orNis returned on ≥ 6 of the 24 calls.
RUN, BEFORE THE MAIN STAGES, AND IT DID NOT CLEAR — the registered consequence fires. N was
returned once in 24. Segment majorities went 3–1 to the unmodified word, not split. And the
number this control was really bought for: on near-identical texts the seats chose the passage
shown FIRST in 17 of the 23 non-N calls — 0.739. The pattern underneath is legible and is not
simple noise: where the substitution is a real register difference the seats hold their view
across both orders (station over position 6 of 6; companions over comrades 5 of 6), and
where it is near-equivalent they fall back on position (chief/greatest and rage/fury: A
chosen 5 of 6 and 5 of 6 respectively, in both orders). So the instrument is not inventing
arbitrary decisiveness — it tracks a one-word difference when it has a view and takes the first
passage when it does not. That is exactly why both orders are run, and 0.739 is the number
F1 must now be read against: a first-position rate near 0.5 on the main stages would mean the
materials are deciding, and one near 0.74 would mean they are not.
- F3a the disclosure is used — gating [A6]. Mechanical coding of the reason strings against
a keyword list that the I2 prompt does not supply: sound, echo, music, cadence, rhythm,
jingle, doggerel, sing-song, sonor, alliterat, ornate, ornament, lyric, poetic, verse, jangl —
case-insensitive substring match. Bar: the mention rate under I2 exceeds the rate under I0 by
≥ 0.15 on the H stage. If the bar fails, a null on P1 is uninformative and P1 is
withheld rather than reported as a null.
- F3b the echo-inclusive rate — reported, gating nothing. The same coding with rhym, chime,
arabic, translat, original, form, faithful added. These are words the I2 prompt puts in front
of the seat, so a difference here is compatible with pure lexical echo; the figure is published
beside F3a so the gap between them can be seen.
- F4 missingness. A call that does not parse after one re-dispatch is a dropped call. A
segment with more than two dropped calls in any (pair, condition) is excluded from that test and
the reference distribution recomputed at the reduced n, with the ids named.
- F5 repeatability, reported not gating. Six H cells re-dispatched once on all three seats;
report the share returning the identical choice. Note (brn) established that determinism is a
property of the task, so this figure is descriptive and no bar is set on it.
- T the pilot, with the only decision rule it is allowed [A13]. Nine calls, one segment, three
conditions. Its only permitted outcomes are proceed-as-designed and abort. No prompt,
material, analysis rule, gate or stopping rule may change after pilot data has been seen. It
exists to price the run and to confirm the JSON parses, and for nothing else.
What would make this run uninterpretable, stated in advance: F2 failing; F1 failing either
of its two legs; F3a failing while P1 is null; F6 returning a decisive one-directional
majority on near-identical texts. Any of these is written up as the result.
6. Predictions, written before the run
Q0will hold — under no information,RHYis chosen in fewer than 8 of 16 cells. The attribute costsRS-20260825bmeasured are the kind a reader acts on.P1will hold, and smaller thanQ0. Telling a reader that the original rhymes throughout is the defence a translator would actually offer, and it should move some readers; it should not move all of them, because grace is grace.P3will not fire. The disclosure is about a sound only one arm has.Q2is the one I cannot call. The balanced rendering may read as tidied prose with nothing gained.L6is the segment to watch. It is whereRS-20260825bfound its largest effect, andBALis the only one of the three arms that does not carry the dine ⁄ opine chime, whichORDfound by accident.
7. What this design cannot do
- Three model seats are not readers. Tier D is NOT PASSED; every figure is
provisional. The claim at issue is a claim about reception, and this instrument is not a reception instrument. - One hand. All three arms are the lead's. Holding the hand constant is what makes the policy comparison clean and is also why nothing here generalises past this hand's execution of the three policies.
- One work, one maqāma of fifty, one language pair.
- A forced choice is not a reading. Nobody is choosing what to spend an evening with.