Repository path: workshop/experiments/E-20260829b-inversion-price/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260829b-inversion-price |
| status | frozen |
| created | 2026-08-29 |
| updated | 2026-08-29 |
| senses | naturalness |
| internal-judgment-only | true |
| amended | 2026-08-29 v2 after pre-run critic |
| provisional | true |
| links | wiki/arms/ARM-inversion-price.md, workshop/regimes/R56-order-pair.md, workshop/translations/hafez-darad/R56-v1/translation.md, runs/RS-20260829b-inversion-price/items.json, framework/v0.2/README.md, config/models.md, config/budget.md, wiki/goodness-senses.md |
E-20260829b — the price of the line-end inversion, with the register held fixed
v2, amended before dispatch after two NEEDS-REDESIGN critic verdicts and 30 findings
(critic-response.md). v1 is in git at dcb25c13.
ARM-inversion-price step 1 (T3). The materials were written and frozen first
(T-hafez-darad-R56-v1, commit dcb25c13) and this design was written after them. The one change
the design has since made to the materials is the one the critic required: the archaic arms are now
generated by script from the plain arms rather than hand-written (critic-response.md, P1-2 /
P3-4 / P3-9). The rendering of record and the plain arms are untouched.
1. The question, and the sentence it is about
framework/v0.2 §7.41.3, published this morning:
"That order is available in an archaizing nineteenth-century verse English and not in a plain one, and Leaf, writing plainer, never uses it and drops every one of those radifs. … So the instruction a translator can act on is not 'you cannot', it is 'you can, at this price, in this register'."
The evidence for it is two published hands who differ in everything at once — era, verse form,
diction, and which poems they chose. RS-20260829-radif-hands §11.2 registers that the design which
separates capacity from register does not exist. This is that design: one hand, one source, one
sense, word order and lexis crossed.
2. Materials — frozen, and machine-checked before this page was written
T-hafez-darad-R56-v1: 18 items, one bayt each, from Hafez sh124 and sh118, both radif
دارد. Every item exists in four arms:
| arm | order | lexis |
|---|---|---|
DP |
direct (English clause order; the line does not end on the radif) | plain present-day |
IP |
inverted (the complement moves before the verb; the radif stands last) | plain present-day |
DA |
direct | archaic |
IA |
inverted | archaic |
The minimal-pair constraint is mechanical, not editorial. Within a lexis, D and I are the
same word multiset in a different order. Across a lexis the archaic arm is a positional
one-for-one substitution of the plain arm — same token count, every differing position licensed by
the frozen map in archaize.py — which is checked position by position, not as a multiset.
pair_check.py v2: 926 checks, 0 failures.
The frozen archaism map, applied by script with no editorial rescue: has→hath, your→thy,
you→thou/thee (role declared per site), pass→passest, shows→showeth, spills→spilleth,
means→meaneth, lacks→lacketh, holds→holdeth, have→hast (2sg auxiliary). Nothing else.
Frozen item codings (R56 §8): what has to move in the I arms — OBJ 11 · ADV 4 ·
MIXED 3; and whether the D arm is itself already marked in its order — CANONICAL 17 ·
MARKED 1 (G2-6).
3. Seats
Panel v1, three seats, config/models.md: P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash,
P3 x-ai/grok-4.5. P4 and P5 are excluded by standing notes (bps), (bne). Temperature 0.
Judgment is not parallelised (charter §6); calls are dispatched one at a time.
The lead wrote all four arms and judges none of them (charter §5, R56 §7). Every body reaches
a seat unlabelled, with no source, no author, no arm name and no mention of Persian, radif,
inversion or register.
4. Stage S — the within-idiom well-formedness rating
Amended after the critic (P1-1, P3-2). The DV is no longer the naturalness sense. Seats
rate well-formedness within the passage's own idiom and are told in the prompt not to reward or
penalise a passage for which kind of English it is written in. wiki/goodness-senses.md allows
naturalness three register anchors and none of them is passage-indexed (D-20260803-15, ratified
A), so this measure is declared outside the six senses, named within-idiom well-formedness,
and is internal-judgment-only. It is the only measure that can answer §7.41.3, whose whole content
is a claim about availability inside a register.
Each of the 72 arm bodies (18 × 4) plus 12 planted-fault controls (6 plain, 6 archaic) is
shown to each seat once. Each seat gets an independently shuffled sequence (seeds 20260829 /
20260830 / 20260831, drawn before dispatch, written to order.json), so order effects do not couple
the three columns (P1-6, P3-16).
Below is a two-line passage of English verse.
Rate it on ONE thing only: how well-formed the English is WITHIN ITS OWN IDIOM. Some passages here are written in present-day English and some in an older, more formal English. Do not reward or penalise a passage for which of those it is. Ask only this: taking the passage's own kind of English as given, is this a well-formed, idiomatic sentence of that English, or is it forced, awkward, or ill-formed?
This is not a question about whether the passage is a good poem, whether it is beautiful, whether it is accurate to anything, or whether you like it.
0 = ill-formed or badly forced, even for the kind of English it is written in. 10 = entirely well-formed and idiomatic for the kind of English it is written in.
{BODY}
Reply with exactly two lines and nothing else: SCORE:
REASON:
max_tokens 300, temperature 0. Note (bsf) binds: the cap is not treated as verified by a probe
on one item; every body's finish_reason is checked and a truncated body already carrying its
SCORE: line is re-parsed, not re-bought (note (brx)).
Repeat control. Six bodies are shown a second time to each seat, at a different position in that
seat's differently shuffled sequence: 18 extra calls. This estimates presentation-order
sensitivity plus sampling noise, which is what it is called; at temperature 0 it is not a
test–retest reliability estimate and is not reported as one (P1-10).
Stage S total: 90 bodies × 3 seats = 270 calls.
5. Stage F — the within-lexis forced choice
For each item and each lexis the two orders are shown together. A/B assignment is randomised
per (item, lexis, seat), balanced inside each front-type stratum, from the same frozen seeds
(P1-7, P3-14) — so every item is seen in both orientations across the three seats and a position
effect is estimable. The prompt no longer announces what differs between the versions, which primed
the dimension under test (P3-12), and it asks the same within-idiom question as stage S.
Below are two versions, A and B, of the same two lines of English verse. Both are written in the same kind of English, so which kind it is cannot be the answer.
A: {A}
B: {B}
Which of the two is the better-formed, more idiomatic sentence of the kind of English it is written in? Ignore beauty, meaning and accuracy.
Reply with exactly two lines and nothing else: CHOICE: REASON:
18 items × 2 lexis × 3 seats = 108 calls, max_tokens 300.
6. Predictions, registered
The analysis unit is frozen here (P1-12, P3-8): score(item, arm) is the mean over usable
seats; the item contrast is c_i = (DP_i − IP_i) − (DA_i − IA_i); the test is an exact
one-sample sign-flip permutation over the 18 item signs (2^18 = 262,144, enumerated), against
the registered bar, not against zero (P1-5) — the contrast is shifted by the bar before
permuting. This is not a randomization test and is not called one: its null is that the contrast
distribution is symmetric about the bar, and its assumption is item exchangeability (P1-4).
Manipulation check, non-withholding. M1′ — within plain lexis, WF(DP) − WF(IP) > 0.
M1 as a gate is removed: it would have withheld the primaries exactly when the hypothesis was
most strongly true (P3-1). Reported, gates nothing.
Primaries. Both are contrasts on this material pack; neither is a claim about English in
general, and the result page's §7.41 language is conditioned in advance (P3-13).
P1— the interaction, on the rating. The inversion costs less under archaic morphology:mean(c_i) ≥ +1.5on the 0–10 scale, and the sign-flip test against that bar atP ≤ 0.05. Bar escalation, registered now: if the repeat control's mean absolute difference exceeds 0.75, the bar becomes twice the observed value (P3-3).P2— the interaction, on the forced choice. Decisive answers only (P1-3,P3-7):share(D) = n(D)/(n(D)+n(I)), andshare(D | plain) − share(D | archaic) ≥ +0.15, sign-flip over item-wise swaps of the two lexis conditions atP ≤ 0.05. TheNEITHERrate by lexis is a separate registered quantity: if it differs between lexes by more than 0.20,P2is withheld and both rates are printed.
The decision matrix, written before any number exists (P3-12):
P1 |
P2 |
what is claimed |
|---|---|---|
| holds | holds | §7.41.3's register clause survives, narrowed to morphology, on this pack |
| fails | fails | §7.41.3's register clause is not supported on this pack and the framework says so |
| holds | fails | no claim either way; the two instruments disagree and that disagreement is the result |
| fails | holds | no claim either way; same |
Registered secondary, descriptive, not a test. S1 — within plain lexis, is the inversion
penalty smaller for ADV items (n = 4) than OBJ items (n = 11)? Reported as a rank comparison
with MIXED (n = 3) excluded explicitly, no bar, no P-value, no directional claim (P3-15). It is
here because D2 in the translator's log says which repertoire a position permits is decided by the
Persian, so a translator can read it off the source.
Primary item set. The 17 CANONICAL items (P3-10); the full 18 including MARKED G2-6 is
printed beside it and any difference between them is reported.
7. Failure criteria — the gates, and they are the only gates
- Positive control, plain.
mean(FAULT_plain) ≤ mean(DP) − 2.0. - Positive control, archaic.
mean(FAULT_archaic) ≤ mean(DA) − 2.0(P3-11,P1-8). Seats could otherwise be noise-flooring all archaic text and still pass gate 1, and gate 1 is whereP1claims to read a reduced penalty. If either control gate misses, everything is WITHHELD. - Compression rule for
P1(P3-5,P1-9), replacing v1's floor rule. All three must hold: - headroom:mean(DA) − mean(FAULT_archaic) ≥ 2.0; - dispersion:SD(DA) ≥ 0.6 × SD(DP); - no bottom pile-up: share of archaic-arm ratings at ≤ 2 is below 0.25. Any one failing withholdsP1as uninterpretable — a compressed archaic half of the scale produces the predicted sign for a reason that has nothing to do with English. Cell distributions are printed before the interaction is read.P2is scale-free and is not gated by this rule. - Complete pairs (
P1-11). An item missing any of its four arms is dropped from the paired primary; below 15 complete items of 18,P1is withheld. Missingness is reported by seat, arm and lexis; reduced denominators are never silently substituted into a paired test. - Noise. Repeat control mean absolute difference; tolerance 0.75, with the bar escalation in §6 as its consequence.
- Dead calls. No parsable
SCORE:/CHOICE:after two re-dispatches is recorded dead. Cost accumulates across attempts (note (brw)). - Seat drop-out. A seat below 90% usable bodies at stage
Sis reported separately and the pooled figures recomputed without it, both printed.
8. What this design cannot establish, written before it runs
- The archaic factor is morphology, not Payne's register. After the critic, the map is purely
substitutive —
hath,thy,thou/theeand the inflections they force — so "archaic register" and "the-thform of the final verb" are the same manipulation here and cannot be separated (P3-6). A null therefore refutes the morphology-carries-it version of §7.41.3 and not the whole sentence; a hold does not establish that Payne's diction, metre or verse convention add anything further. - Tier D is NOT PASSED. Nothing here is a claim about what a human reader perceives, prefers or
wants. Three model seats rating English well-formedness is
internal-judgment-only. - The
naturalnesssense is not measured. The DV is passage-indexed by necessity (§4), so no figure here is in this project'snaturalnessunits and none may be compared to one. - One hand, and the order factor is still hand-written. The mechanical map removes the
translator from the lexis factor entirely; it does not remove him from the order factor, where
the multiset check constrains the words but not the felicity of the arrangement. An inverted line
the lead could not write well is indistinguishable from an inverted line English does not permit
(
P3-4, partly unfixed and named). The design removes the between-hand confound that §11.2 named and replaces it with a within-hand one: an inverted line the lead could not write well is indistinguishable from an inverted line English does not permit. - No chime.
R56§5 refuses the rhyme, so this prices the alignment of the tail and not the full device (D7in the translator's log). - The
Darms are not the radif dropped. They keep the word and lose its position (D9). What is priced is alignment, not repetition, and the result page may not widen it. - Two poems, one radif, one poet, one language pair.
9. Pre-flight cost estimate
Worst case is built from the caps, not from expected output (note (abc)).
| stage | calls | cap | worst-case cost |
|---|---|---|---|
C pre-run critic, 2 seats — spent, $0.126422400, both NEEDS-REDESIGN |
2 | 12,000 | actual |
S rating, 84 bodies × 3 seats |
252 | 300 | $1.05 |
S repeat control, 6 × 3 |
18 | 300 | $0.08 |
F forced choice, 36 × 3 |
108 | 300 | $0.50 |
| re-dispatch headroom (note (brw)) | — | — | $0.30 |
| declared ceiling, whole run | 380 | $2.20 |
UTC day 2026-08-29 stands at $0.197706300 of $5.00 after S231, so $4.802293700 is
available and a $2.20 ceiling fits with room. The opening key snapshot for this session reads
156.738407873, which is $0.872971860 above S231's close — non-project spend on the key
between sessions, recorded in config/budget.md and not this project's. Key usage is snapshotted before and after and the
delta cross-checked against the per-request sum.
10. Verification
analyse.py --mutate recomputes every number reported on the result page from raw/, and plants
deliberate faults to confirm the checks can fail. Nothing is reported that the verifier does not
recompute.