Repository path: wiki/findings/results/RS-20260903b-line-end-order.md · rendered 2026-09-09
Page metadata (front matter)
| type | result |
|---|---|
| id | RS-20260903b-line-end-order |
| status | frozen |
| created | 2026-09-03 |
| updated | 2026-09-03 |
| senses | style-correspondence |
| provisional | true |
| internal-judgment-only | true |
| links | wiki/arms/ARM-line-end-order.md, workshop/experiments/E-20260903b-line-end-order/design.md, workshop/experiments/E-20260903b-line-end-order/design-v2.md, workshop/experiments/E-20260903b-line-end-order/critic-response.md, workshop/translations/hafez-shahed/R60-v1/translation.md, workshop/translations/hafez-shahed/source.md, workshop/translations/hafez-shahed/dependence-note.md, workshop/regimes/R60-free-line-end.md, wiki/findings/results/RS-20260901-inversion-habit.md, wiki/findings/results/RS-20260829-radif-hands.md, config/models.md, config/budget.md |
The word-order primaries are withheld a second time — a fresh three-seat run put one seat through both gates where two were required, and the seat that reads the verse best still cannot pass a floor built to catch a bad coder
ARM-line-end-order step 1 (T3), 2026-09-03. E-20260903b-line-end-order v2, after two
independent pre-run critic seats returned NEEDS-REDESIGN with 28 findings, 8 BLOCKING on one
point (critic-response.md). Tier D is NOT PASSED. Nothing here is a judgment of quality and
no sentence says a reader hears, prefers or wants anything.
The headline is a withholding, and it is the second one in a row on this question.
RS-20260901-inversion-habit (S237) bought 523 line-end codes and withheld every word-order figure
because one seat of three cleared its gates. This session was built to answer the charge that S237's
gate had been loosened after it fired: it bought every voting code fresh, after both gates were
committed, with the synthetic floor restored as binding and one defective calibration key
corrected under a rule fixed in advance. The result is the same. One seat of three cleared both
gates; the design required two; and analyse.py returns before any cell rate, contrast or interval
is computed. The HABIT account of what separated Payne 1901 from Leaf 1898 — is the line-end
inversion a standing property of a hand's line, or bought where a device needs it — remains
unmeasured, and this is now a finding about the construct's codeability, not a delay.
1. What was at stake, and why the design was rebuilt rather than re-analysed
framework/v0.2 §7.47.3 states in writing that the project does not know whether the inversion is a
local purchase or a line-grammar. The backlog row that named the missing measurement proposed
re-running S237's seats with the real-item key as the gate. Two critic seats called that
NEEDS-REDESIGN with eight BLOCKING findings on one point (C1-A1, C1-A2, C1-A4, C1-A15;
C2-A1, C2-A2, C2-A3, C2-A9): demoting the gate that fired while keeping the codes it fired on
is gate-shopping, whatever it discloses. Granted in full. The operative design (design-v2.md) buys
every voting code fresh, dispatches every seat after both gates are committed, restores the floor as
binding, and adds a seat — qwen/qwen3.7-max — that had never coded this construct.
The subject rule, in one sentence (continue-prompt.md §4.5): this unit asks a translator of
Persian verse whether the word-order move that carries a ghazal's commonest formal device can be
bought for one line or must be adopted for the whole poem, and answers it by counting what two
published translators and one working hand put at their line ends. The gate repair is a gate inside
the unit, never its subject.
2. The wire between the limbs
T-hafez-shahed-R60-v1 rendered غزل ۱۵, «ای شاهد قدسی، کِه کَشَد بند نقابت» whole, in a
declared plain modern register, under the new regime R60 — the no-radif counterpart of
S237's sh77, whose radif forced a finite verb to stand last at nine positions. This poem obliges
nothing at any line end. Its 20 line ends and the 16 of T-hafez-daasht-R58-v1 (sh77) went into
the same blind coding batches as Payne's and Leaf's, indistinguishable from them. The rendering and
its log were frozen and committed (dda56092) before any English of the poem was opened;
contamination measured none (dependence-note.md, 0 shared 7-grams, longest run 3 tokens
against Payne's volume, which contains this ghazal as ode 90).
3. The gates — one seat of three, and it is the reserve
The bought stage put 447 items to three non-Anthropic seats (387 study line ends + 24 corrected calibration + 36 lead line ends): one printed couplet with one line marked, no hand, no book, no poem, no rule. Two gates, both binding and both registered before any voting code existed: a floor of ≥ 21 of 24 synthetic pairs, and ≥ 0.75 agreement with 60 fresh real study items the lead coded blind — disjoint from S237's 40, double-coded (self-agreement 58/60 = 0.967), adjudicated and frozen before dispatch.
| seat | model | floor, of 24 | gold2 agreement | votes |
|---|---|---|---|---|
Q1 |
openai/gpt-5.6-terra |
19 ✗ | 0.700 [0.575, 0.801] ✗ | no |
Q2 |
google/gemini-3.6-flash |
19 ✗ | 0.840 [0.715, 0.917] ✓ | no |
Q3 |
qwen/qwen3.7-max |
23 ✓ | 0.850 [0.739, 0.919] ✓ | yes |
Failure criterion 2 fires: fewer than two voting seats, so no ORDER figure is reported at all.
The withholding is enforced by the code, which returns before any primary — the verifier asserts it
(26 checks, 0 failures; 5 of 5 mutations caught, every mutation targeting a figure this page reports
on the withholding path).
The frontier seat that voted at S237 failed both gates here. Q1 (gpt-5.6-terra), which was
the single voting seat at S237, agreed with the fresh gold at only 0.700 and cleared the floor at
19. Its S237 pass was on a 40-item gold it happened to match at 0.800; on 60 fresh items it does
not. This is the strongest thing the session establishes: the seat the earlier design certified
would not certify on independently drawn material, which is exactly the fragility the critics said a
single-sample gate could hide.
4. Why this is now a finding about the construct, not a delay
Two withholdings on the same question, under two independently-built gates, one of them deliberately stricter than the first, is evidence that English line-end word order is not codeable reliably enough for two certified seats at this project's blind-coding reliability — not that the next design will succeed.
- Mean pairwise agreement on
ORDERacross all three seats: 0.803; unanimity 0.707, over 447 items. The construct is codeable at about four-fifths agreement — better than S237's 0.741 — and still short of what two independent gates admit. - The floor caught the seat that reads the verse best, again.
Q3, the one voting seat, is the cheap reserve (qwen/qwen3.7-max, $0.74 for 447 items); the two frontier seats both failed. At S237 the excluded seat had the highest real-item agreement (note (bsp)); here the included seat is the reserve and the frontier seats are out. A gate that keeps changing which seat it admits is not measuring a stable property of the seats. The construct sits at the edge of what any of these seats can hold steady over hundreds of natural verse lines. - The corrected calibration key changed nothing at the gate. No item was unanimously against the
corrected key; on the uncorrected key the floor scores are Q1 20, Q2 18, Q3 22 — the same three
pass/fail verdicts. The
CAL11mcorrection (§design-v2 §4) was the honest thing to do and it did not manufacture the result.
5. The one thing that IS measured, and it is the critics' enjambment finding
C1-A11, C1-A12 and C2-A8 said the three-step coding instruction mishandles enjambment: for a
line whose clause runs past the printed couplet, "would the last word still be last in prose" is not
a word-order fact about that hand. The design added a separate CLAUSE field and made the lead code
it on the fresh gold. In the lead's own frozen gold, every one of the 21 items whose clause
continues past the couplet is coded INVERTED — 21 of 21 — against 21 INVERTED of 39 among the
complete-clause items. Enjambment and apparent inversion are almost perfectly confounded at the
line end: a line that spills its clause forward will read as displaced whether or not the hand
displaced anything.
This is why the enjambment stratification was made mandatory in the design, not optional
(design-v2 §6): had the primaries reported, an inversion-rate difference between two hands who
enjamb at different rates could have been an enjambment-rate difference wearing a word-order mask.
The one voting seat's descriptive CONTINUES rates are Payne 0.160 and Leaf 0.148 — close, but the
finding stands as a method result: any future coding of English line-end word order must
separate enjambment from inversion before it counts, because the two are confounded at the line
end by construction. It goes to framework/v0.2 as a caution and to method note (bsx).
6. The translation limb, exploratory (design-v2 §8) — counts, no rates
By the lead's own frozen R60 grading of غزل ۱۵ (tally_log.py over the frozen log):
sh15(no radif): 20 line ends, 19CANONICAL, 1INVERTED. The one inversion is a place where the Persian itself fronts the object (every moan and cry that I made → verb last). At the 11 rhyming positions, which oblige nothing, the English inverts 0 times.
sh77(radifhad, from S237's frozen log): 9 rhyming positions carried by inversion, 6LOCAL, 3SPREAD.
Set beside each other these are the translation limb's version of the withheld study question: the
same hand, in the same declared register, inverts its line end where a radif obliges it and
almost nowhere else. That is one hand's existence result — 16 and 20 line ends, no rate — and it
is consistent with the LOCAL-purchase reading RS-20260901-inversion-habit §7.47.2 drew from
sh77 alone. It is not evidence about the published hands, and the hand knew the question while
writing (R60 §5, the log's §5); the "upper bound" claim of the design's first draft was withdrawn
(C1-A8). No confirmatory weight rests on it.
7. What this changes in the framework
Written into framework/v0.2 as §7.51, v0.2.51.
- §7.51.1 — the
HABITaccount is still unmeasured, and now for a reason. Two designs, two withholdings; the construct is codeable at ~0.80 agreement but not reliably enough for two certified seats under gates registered before the codes. §7.47.3's "the project does not know" stands, and now names why: the measurement is at the edge of what blind AI coding of natural verse word order can do here. - §7.51.2 — a method finding that binds future work: enjambment and apparent line-end inversion
are confounded at the line end by construction (21 of 21 enjambed items coded
INVERTEDin the lead's own gold), so any count of line-end word order must separate them first. Note (bsx). - §7.51.3 — a caution about single-sample seat qualification: the seat certified by S237's gate failed a fresh gold drawn from the same construct. A seat qualified on one small hand-coded set is not thereby qualified; the qualification must be on material the seat's reported codes did not help choose.
8. Limits
- No word-order primary exists. Nothing here may be cited as a rate of inversion in either
hand. Both
PR1(matched-poem) andPR2(within-Payne spill) are unreported. - The gold key is one coder, who also wrote the coding definition; double-coded (0.967 self-agreement) but not independent. An independent human coder is unbuilt and has been on the record as unbuilt for twenty-three sessions.
- A null on the
HABITaccount is not fully separable from an instrument that cannot read natural verse (C1-A14,C2-A10). The restored floor shows a seat can detect a constructed contrast; that two of three frontier-and-reserve seats fail it is itself the limit. Q2lost five batches of 12 (60 items) to truncation at the probed cap — note (bsf), a sixth firing, and this run honoured the note (the cap was probed at the dispatch batch size and passed) and was bitten anyway on later items of the same shape.Q2failed its gate regardless, so the loss changed no verdict; it is recorded, not repaired. The bodies were not re-bought.- The translation limb is exploratory (§6), by one hand that knew the question.
- Tier D is NOT PASSED.
9. Cost
$3.402299275 of a declared $3.50-effective ceiling (design-v2 §11 set $4.50; the run came in under $3.50). UTC day 2026-09-03 opened at $0.00 of $5.00.
| stage | calls | actual |
|---|---|---|
C pre-run critics, C1+C2, cap 12000 — both NEEDS-REDESIGN, 28 findings |
2 | $0.089021400 |
| cap/price probes, three seats at the dispatch batch size — note (bsf) | 3 | $0.130562125 |
B blind coding, Q1, 447 items, 19 batches of 24, 0 dead |
19 | $1.017083600 |
B blind coding, Q2, 447 items, 38 batches of 12, 5 dead (truncation) |
38 | $1.425220500 |
B blind coding, Q3, 447 items, 19 batches of 24, 0 dead |
19 | $0.740411650 |
| per-request sum, and the ledgered figure | 81 | $3.402299275 |
config/models.md was corrected upward before dispatch — openai/gpt-5.6-terra reads $2.00 /
$12.00 from the API, double the stale $1.00 / $6.00 the table carried, the first upward move this
project has recorded and the first that could breach a ceiling. The ceiling was set from the read
price, and note (bsw) now requires reading every dispatched seat's price in the session that
dispatches it.
Snapshots: opening (after the critic pass) 247.816893901, closing 251.219193176, delta $3.402299275 against a per-request sum of $3.402299275 — exact to 1e-9, and the exactness is because the opening snapshot already carried the critics. Per note (bso) the snapshot is an observation; the per-request sum is the ledgered figure.
The translation limb, the selection rule, both gold coding passes, all extraction and every arithmetic are the lead's own and are not ledgered (charter §3, A4).
Verifier: 26 checks, 0 failures; 5 of 5 mutations caught
(runs/RS-20260903b-line-end-order/analyse.py --verify, --mutate). Every mutation targets a
figure this page reports on the withholding path — the deciding gate, the enjambment finding, the
between-run agreement, the gold self-agreement — because there is no primary to mutate.