Repository path: workshop/experiments/E-20260903b-line-end-order/design-v2.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260903b-line-end-order-v2 |
| status | frozen |
| created | 2026-09-03 |
| updated | 2026-09-03 |
| senses | style-correspondence |
| provisional | true |
| internal-judgment-only | true |
| links | workshop/experiments/E-20260903b-line-end-order/design.md, workshop/experiments/E-20260903b-line-end-order/critic-response.md, workshop/experiments/E-20260901-inversion-habit/design-v2.md, wiki/arms/ARM-line-end-order.md, workshop/translations/hafez-shahed/R60-v1/translation.md, framework/v0.2/README.md, config/models.md, config/budget.md |
Design v2 — the operative design: a prospective re-run, both gates binding, no reused code in any primary
This supersedes design.md, which stays frozen as the record of what was proposed and refused.
Eight BLOCKING findings across two critic seats landed on one point and it is granted:
a gate rewritten after seeing which seats it excluded cannot license the figures it lets through,
and disclosing that does not repair it. critic-response.md carries the dispositions.
E-20260901-inversion-habit v2 §5's coding instruction and §7's PR1, PR2, PR4 are imported
verbatim. PR2b is dropped with its cells (§2 below); PR3 was reported at
RS-20260901-inversion-habit §3 and is closed.
1. The rule that governs everything else
No ORDER code bought at S237 enters any primary, any interval, or any sensitivity analysis.
S237's codes are read once, at the end, as a between-run diagnostic, and the verifier asserts that
no primary function ever opens stage_b_P*.json from that run.
Every voting seat is dispatched after gold2 is coded, frozen and committed, and after both
gates are registered in this file.
2. Item set — trimmed so that three seats fit the day
The cells are E-20260901-inversion-habit v2 §4's, unchanged in definition. L-RHYME, O-P and
N-P are dropped, which drops PR2b with them: PR2b was secondary and v2 §7 already said its
definitional component made it weak. What remains is exactly what PR1, PR2 and PR4 need, plus
the instrument control.
| cell | n | used by |
|---|---|---|
INT-P |
83 | PR1 |
INT-L |
81 | PR1 |
SPILL-Y |
110 | PR2 |
SPILL-N |
124 | PR2 |
CTRL+ |
72 | failure criterion 1 |
| unique study items | 387 | (cells are labels, not partitions) |
CALIB |
24 | the floor, §4 |
LEAD |
36 | §8, exploratory |
| total dispatched per seat | 447 |
Payne 306, Leaf 81. 37 of the 387 carry a damaged final token and are flagged, as at S237.
3. Seats — three, all dispatched after the gates are committed
| seat | slug | why |
|---|---|---|
Q1 |
openai/gpt-5.6-terra |
frontier seat, cheapest in practice on this shape at S237 |
Q2 |
google/gemini-3.6-flash |
batch 12, the size probed at S237; it still lost one batch of 44 there |
Q3 |
qwen/qwen3.7-max |
first reserve in config/models.md, never used on this construct |
New tags Q1–Q3 are used deliberately so that no file, figure or sentence can confuse a
prospective code with an S237 one. Non-Anthropic throughout (charter §5). P4, P5 and
z-ai/glm-5.2 are out under notes (bps), (bne), (brt). x-ai/grok-4.5 at
reasoning: {"effort": "low"} is out by registered rule, note (bsp), and at default effort it
prices at roughly thirteen times its S237 cost, which does not fit the day.
qwen/qwen3.7-max is priced from GET /api/v1/models before dispatch, as the gate in
config/models.md §S182 requires for any reserve seat, and the read is recorded in the cost table.
4. Two gates, both binding, both registered before any code exists
Gate 1 — the synthetic floor, RESTORED. 24 pairs, each one real printed line whose printed order is canonical plus a counterpart made by moving that line's finite verb to the end and changing no word. Bar: ≥ 21 of 24.
The one key correction, under a rule fixed before any new code exists. A calibration item's key
is corrected if and only if (a) the lead's independent reading finds the key wrong and
(b) all three S237 seats coded against it. Applied to all 24 items, exactly one qualifies:
CAL11m, whose hand-made "canonical" counterpart "To me the East wind yesternight hath brought
the tidings rare" postposes its adjective and leaves the adverbial before the subject. Its key
changes CANONICAL → INVERTED. The floor is scored on the corrected key, and both the
corrected and uncorrected scores are reported for every seat. The audit over all 24 items is
committed at runs/RS-20260903b-line-end-order/calib_audit.json.
A consequence of the correction, recorded because it weakens the floor slightly. Pair 11 now
has both members keyed INVERTED, so the construction property "exactly one member of each
pair is INVERTED" no longer holds for that pair. The floor is therefore 24 keyed items, not 12
certain-by-construction pairs, and CAL11m's key rests on the lead's reading plus three seats'
agreement rather than on construction. Pair members still never share a batch.
Gate 2 — real-item validity. 60 study items, drawn by random.Random(20260903) from the 387,
disjoint from S237's 40 gold items, printed with hand, ode, cell and every seat's code stripped
and the order shuffled, coded by the lead twice in two passes separated by a reshuffle, with
disagreements adjudicated and the adjudicated key frozen and committed before any seat is
dispatched. Self-agreement between the two passes is reported. Bar: ≥ 0.75 agreement on
ORDER.
A seat votes only if it clears BOTH gates. This is stricter than S237, which is the point: the charge the critics made is that the author loosened a rule after it fired, and the answer is to tighten it and buy fresh codes rather than to argue.
What decides is the observed point — floor ≥ 21 of 24 and gold2 agreement ≥ 0.75. Wilson 95%
intervals are printed beside both and decide nothing (C1-A16, C2-A12).
5. Concordance among qualified seats
- Three qualified seats ⇒ item code is the majority; no majority ⇒ dropped.
- Two qualified ⇒ a disagreement is a drop.
- Fewer than two ⇒ no
ORDERfigure at all (failure criterion 2). - Per-seat reporting trigger (
C2-A4): if any qualified seat'sINVERTEDrate in any of the four primary cells differs from the pooled rate by more than 0.15, every primary is reported per seat as well as pooled, and the result page says which seat is the outlier. A qualified seat is never silently outvoted.
6. Enjambment — a separate field, and mandatory stratified sensitivity
C1-A11, C1-A12 and C2-A8 are granted: the three-step instruction asks whether the marked
line's last word "would still be last" in prose, and for a line whose clause continues overleaf the
honest answer is not a word-order fact about that hand. Two changes, neither of which touches
ORDER's three values or the halt logic:
- A fourth returned field,
CLAUSE:COMPLETE— the clause containing the marked line's last word finishes inside the printed couplet;CONTINUES— it runs past it. - Both primaries are recomputed on
CLAUSE == COMPLETEitems only, and that recomputation is mandatory, not optional. If the two versions of a primary fall on opposite sides of the registered bar, the primary is reported as indeterminate and the enjambment difference is the finding. - The per-hand
CONTINUESrate is reported as a fact about the two books.
7. Failure criteria
CTRL+below 0.70INVERTEDon the qualified seats ⇒ instrument broken, no primary reported, whatever the gates say.- Fewer than two seats clearing both gates ⇒ no
ORDERfigure at all. Enforced byanalyse.pyreturning before any cell rate is computed, and asserted by the verifier. - Mean pairwise agreement on
ORDERbelow 0.70 among qualified seats ⇒ primaries withheld. - Drop rate above 0.20 in any primary cell ⇒ that cell is reported as the interval between
assigning every dropped item
CANONICALand assigning every dropped itemINVERTED, and the prediction outcome is read off the worse end (C1-A13). - Final-token damage rates differing by more than 0.10 between the hands ⇒
PR1reported as a range. (S237 measured 0.0934 and 0.1329, difference 0.0395; recomputed here.) - Denominator below 25 ⇒ counts, not rates.
- Sensitivity, always run: every primary recomputed (a) excluding
UNCLEAR, (b) excluding damaged final tokens, (c) onCLAUSE == COMPLETEonly.
8. The translation limb — exploratory, and labelled so everywhere
T-hafez-shahed-R60-v1 renders غزل ۱۵ whole, the no-radif counterpart of S237's sh77,
selected by the frozen rule in workshop/translations/hafez-shahed/source.md, in the same declared
plain modern register by the same hand. Its 20 line ends and the 16 of T-hafez-daasht-R58-v1 go
into the coding batches indistinguishable from Payne's and Leaf's.
PR6 is withdrawn as a registered prediction (C1-A9, C2-A6): at n = 7 against n = 9 a count
comparison has no decision rule worth the name. What is reported is the counts, per seat and
pooled, as a descriptive case observation.
PR7 is relabelled a manipulation check (C1-A10, C2-A7), not a control: that the rhyming
line ends of a poem carrying a finite transitive radif are coded INVERTED more often than those of
a poem carrying nothing is close to definitional, and it is reported to show the seats read the
lead's lines at all, never as evidence for anything.
The "upper bound" claim of v1 §4 is WITHDRAWN (C1-A8). The hand knew the question while
writing sh15; that awareness could have acted in either direction, and nothing here bounds it.
No confirmatory claim rests on the translation limb.
9. Verification
analyse.py --verify recomputes every reported number from the raw stage files, asserts body
integrity, and asserts that no S237 stage file is opened by any primary path.
analyse.py --mutate requires each mutation to change a figure this page actually reports.
10. Limits carried into the result page
- The gold key is one coder — the same coder who wrote the definition, double-coded but not independent. An independent human coder is unbuilt and has been on the record as unbuilt for twenty-three sessions.
- A null on
PR1is not fully separable from an instrument that cannot read natural verse (C1-A14,C2-A10). The restored floor shows a seat can detect a constructed contrast; nothing available to this project shows it reads subtle natural displacement. - Two hands, and register is confounded with hand.
PR2is the within-Payne primary and is where register is constant. - Eleven matched poems for
PR1: Payne's volume 1 stops at Brockhaus CC. - No page images. All couplets are printed in
runs/RS-20260901-inversion-habit/line-ends.md. - The translation limb is exploratory (§8), by one hand that knew the question.
- Tier D is NOT PASSED. No sentence may say a reader hears, prefers or wants anything.
11. Pre-flight cost
UTC day 2026-09-03 has one prior session, S241, which spent $0.00. The whole $5.00 is available. $0.089021400 is already spent on the critic pass. Declared ceiling $4.50.
| stage | calls | worst case |
|---|---|---|
C pre-run critics (spent) |
2 | $0.089021400 actual |
cap and price probe, Q3 at the batch size to be dispatched — note (bsf), six firings |
2 | $0.16 |
B Q1, 447 items, 19 batches of 24, cap 12000 |
19 | $1.10 |
B Q2, 447 items, 38 batches of 12, cap 12000 |
38 | $1.40 |
B Q3, 447 items, 19 batches of 24, cap 12000 |
19 | $1.35 |
| re-dispatch headroom, note (bsf) | ≤ 8 | $0.40 |
| total | ≤ 88 | $4.50 |
Built from max_tokens and the batch size actually dispatched, not from an assumed output length —
note (abc). If the Q3 probe prices the seat above $0.09 a batch, Q3 is dropped and the run
goes with two seats, which §5 permits and which the result page then says. The translation limb,
the selection, both gold passes, all sampling and every interval are the lead's own and are not
ledgered (charter §3, A4).