Repository path: workshop/experiments/E-20260903b-line-end-order/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260903b-line-end-order |
| status | frozen |
| created | 2026-09-03 |
| updated | 2026-09-03 |
| senses | style-correspondence |
| provisional | true |
| internal-judgment-only | true |
| links | workshop/experiments/E-20260901-inversion-habit/design-v2.md, workshop/experiments/E-20260901-inversion-habit/critic-response.md, wiki/findings/results/RS-20260901-inversion-habit.md, wiki/arms/ARM-line-end-order.md, workshop/translations/hafez-shahed/R60-v1/translation.md, workshop/regimes/R60-free-line-end.md, framework/v0.2/README.md, config/models.md, config/budget.md |
Design — the withheld line-end primaries, re-opened under a gate that is not the one that stopped them
ARM-line-end-order step 1 (T3), 2026-09-03. This design does not replace
E-20260901-inversion-habit v2. v2's §4 cells, §5 coding instruction, §7 predictions PR1,
PR2, PR2b, PR4, §8 decision matrix and §10 limits are imported verbatim and are not
re-opened. What changes is §6 (the gates), which is what withheld the run, plus two additions:
a fresh validity set, and the lead's own line ends put through the same instrument.
1. What is being repaired, and the part of it that is post hoc
RS-20260901-inversion-habit bought 523 items × 3 seats and withheld every word-order figure
because one seat of a required two cleared both registered gates. The gate that did the
excluding was the 24-item synthetic floor, and three things are on the record about it:
- It excluded the seat with the highest real-item agreement in the run.
P2agreed with the lead's blind gold at 36 of 40 against the voting seat's 32 of 40, and failed the floor at 20 of 24 against a bar of 21. - One of the 24 floor items is defective and all three seats caught it.
CAL11m's hand-made "canonical" counterpart postposes an adjective and leaves an adverbial before the subject; it isINVERTED, the key saysCANONICAL, and every seat lost a bit to it. - Both pre-run critics said in advance that the artificial task would establish little
(
C1-A7,C1-A8,C2-A4,C2-A11). They were right about passing it. What nobody predicted is that failing it establishes little either.
The repair: the synthetic floor is demoted to a reported diagnostic, and agreement with real, hand-coded items becomes the gate.
The part of this that is post hoc, stated plainly and not buried. The decision to demote the
floor was taken knowing that P1 passed it and P2 did not, and knowing both seats' gold scores.
A gate rewritten after seeing which seats it excluded cannot license anything on its own. Two things
are done about that, and they are the reason this design exists rather than a re-analysis:
- A fresh validity set (§3). The gate is evaluated on 60 study items the lead codes blind now, disjoint from the 40 already used, frozen and committed before any seat is scored on them. The threshold is unchanged at ≥ 0.75 — it is v2's, not a new number chosen to fit.
- A seat whose codes do not yet exist (§2). One new non-Anthropic seat codes the whole item set under a gate registered before it was dispatched. If it qualifies, at least one voting seat is qualified by a criterion that could have excluded it.
What remains post hoc even so, and is declared as a headline limit on the result page: P1 and
P2's codes were bought before this gate was written. Their re-qualification on fresh items is an
unbiased estimate, but the choice to keep looking for a gate they could pass is not something a
fresh sample can undo.
2. Seats
| seat | slug | status here |
|---|---|---|
P1 |
openai/gpt-5.6-terra |
codes reused from S237 for the 499 study items and 24 calibration items; newly dispatched on the 36 lead line ends |
P2 |
google/gemini-3.6-flash |
same, at batch 12 (S237's probed size) |
P6 |
qwen/qwen3.7-max |
new, first reserve in config/models.md, never used on this construct; codes all items — 499 study + 24 calibration + 36 lead |
P3 |
x-ai/grok-4.5 at reasoning: {"effort": "low"} |
excluded from voting by registered rule, whatever any gate says — note (bsp): on the 40 gold items it called an INVERTED line CANONICAL 7 times of 19, against P1's 3 and P2's 0. Its codes are reported as a diagnostic |
Non-Anthropic throughout (charter §5). P4 and P5 are out under notes (bps) and (bne);
z-ai/glm-5.2 is out on long prompts, note (brt).
P6 is priced from the API before dispatch — config/models.md carries $1.475 / $4.425 read
2026-08-04, and the pricing gate that table's S182 entry describes applies to any reserve seat.
3. The fresh validity set — the gate
- 60 items, drawn by
random.Random(20260903)from the 499 study items, disjoint from the 40gold_uidsof S237. - Printed by
gold2_blind.pywith hand, ode, cell and every seat's code stripped, in shuffled order, exactly as S237's sheet was. - Coded by the lead on
ORDERalone, using v2 §5's three-step instruction verbatim, and frozen and committed togold2.jsonbefore any seat is scored against it and beforeP6is dispatched. - Gate: a seat votes only at ≥ 0.75 agreement with
gold2onORDER. Wilson 95% intervals are reported beside every agreement figure, so a threshold decision is never presented as a point. - The circularity is declared and is unchanged from v2: the lead wrote the coding definition, so this measures whether a seat applies this definition, not whether the definition is right.
The 24-item synthetic floor is computed and reported for every seat and gates nothing. The
defective item CAL11m is reported both included and excluded.
4. The lead's own line ends — the translation limb, wired
T-hafez-shahed-R60-v1 renders غزل ۱۵ whole — the no-radif counterpart of S237's sh77,
selected by the frozen mechanical rule in workshop/translations/hafez-shahed/source.md, in the
same declared plain modern register by the same hand. 36 line ends go into the coding
batches, indistinguishable from Payne's and Leaf's: 20 from sh15 and 16 from
T-hafez-daasht-R58-v1 (sh77), one couplet at a time, no hand, no poem, no rule.
The wire, in one sentence. The study limb asks whether a published hand who inverts where a radif obliges him also inverts where nothing obliges him; the translation limb puts the same question to the hand that wrote the log, by rendering a ghazal that obliges nothing and sending both poems' line ends through the study's own blind instrument.
Declared before any of it is coded (R60 §5, and the log's own §5): the hand knew the question
while writing sh15, and T-hafez-daasht-R58-v1's log did not. The unbiased side of the contrast
is the obligated one; any bias runs to widen the contrast, so the contrast is an upper bound.
5. Predictions
PR1, PR2, PR2b, PR4 are imported verbatim from E-20260901-inversion-habit v2 §7 and are
not restated or altered here. PR3 was reported at RS-20260901-inversion-habit §3 and is closed.
New, and registered before the lead's line ends are coded by anybody:
PR6(the hand's own spill, direction only). Among the unobligated line ends of the lead's two renderings — the first hemistich of each non-maṭlaʿ bayt — theINVERTEDcount is higher inT-hafez-daasht-R58-v1, where a radif is carried at every rhyming position, than inT-hafez-shahed-R60-v1, where nothing is carried. n = 7 and 9, so failure criterion 6 fires by construction: counts, not rates, and no interval is claimed. This is the translation limb's own version ofPR2and it is registered as an existence observation.PR7(the obligated slot). InT-hafez-daasht-R58-v1the rhyming line ends areINVERTEDat a higher count than inT-hafez-shahed-R60-v1. This is close to definitional — the radif is a finite transitive verb and English cannot end a clause on one without fronting its object — and is registered as the positive control on the lead's own material: if the seats do not see it, they are not reading the lead's lines the way they read Payne's.
6. Failure criteria
Imported from v2 §9 with one substitution and one addition:
CTRL+below 0.70INVERTEDon the voting seats ⇒ instrument broken, no primary reported, whatever the gates say. Unchanged, and it is the guard that does not depend on the gold set.- Any seat below 0.75 agreement with
gold2is dropped and reported. Fewer than two voting seats ⇒ noORDERfigure at all — the same haltanalyse.pyenforced in S237, and the same code path enforces it here. - Mean pairwise agreement on
ORDERbelow 0.70 among voting seats ⇒ the construct is reported as not reliably codeable and the primaries are withheld. - Drop rate above 0.20 in any cell ⇒ that cell is reported as a range over the dropped items.
- Final-token damage rates differing by more than 0.10 between the hands ⇒
PR1is reported as a range. (Measured at S237: Payne 0.0934, Leaf 0.1329, difference 0.0395 — does not fire.) - Denominator below 25 ⇒ counts, not rates. Fires by construction on
PR6andPR7. - Sensitivity analysis, always run: every primary recomputed excluding
UNCLEARitems and damaged-final-token items. - New — the qualification-provenance flag. If the voting seats are only seats whose codes
predate this gate (i.e.
P6fails), every reported primary carries that fact in the sentence that reports it, and the result page's headline says the figures rest on seats re-qualified after their codes existed.
7. Statistics
As v2 §7: Wilson 95% on every rate, Newcombe 95% on every difference, and an ode-clustered bootstrap (5,000 resamples of odes with replacement) on both primaries. Outcomes are three-way — supported / contradicted / indeterminate. Non-unanimous items among voting seats are resolved by majority; with two voting seats a disagreement is a drop, and the drop rate is printed per cell.
8. Verification
analyse.py --verify recomputes every reported number from the raw stage files and asserts body
integrity; analyse.py --mutate corrupts inputs and requires each mutation to change a figure
this page actually reports — the repair RS-20260901-inversion-habit §8 made after a mutation
passed harmlessly because the gate had withheld everything.
9. Pre-flight cost
UTC day 2026-09-03 has one prior session, S241, which spent $0.00. The whole $5.00 is available. Declared ceiling $2.60.
| stage | calls | worst case |
|---|---|---|
C pre-run critics, 2 seats, cap 12000 |
2 | $0.12 |
cap probe for P6 at the batch size to be dispatched — note (bsf), six firings |
2 | $0.14 |
B6 P6 blind coding, 559 items ÷ 24 ≈ 24 calls |
24 | $1.55 |
BL lead line ends, 36 items, P1 (2 calls) and P2 (3 calls at batch 12) |
5 | $0.28 |
| re-dispatch headroom, note (bsf) | ≤ 8 | $0.51 |
| total | ≤ 41 | $2.60 |
Built from max_tokens and the batch size actually dispatched, not from an assumed output length —
note (abc). The translation limb, the selection, the fresh gold coding, all sampling and every
interval are the lead's own and are not ledgered (charter §3, A4).