Repository path: workshop/experiments/E-20260901-inversion-habit/critic-response.md · rendered 2026-09-09
Page metadata (front matter)
| type | note |
|---|---|
| id | critic-response-E-20260901-inversion-habit |
| status | frozen |
| created | 2026-09-01 |
| updated | 2026-09-01 |
| links | workshop/experiments/E-20260901-inversion-habit/design.md, workshop/experiments/E-20260901-inversion-habit/design-v2.md, runs/RS-20260901-inversion-habit/critic.log |
Pre-run critic pass — what the two seats found and what changed
Both seats returned NEEDS-REDESIGN. 28 findings, 9 BLOCKING. C1
openai/gpt-5.6-terra (13 findings, 5 BLOCKING), C2 x-ai/grok-4.5 (15 findings, 5
BLOCKING), each shown the frozen v1 design and nothing else, at disjoint labs. Raw at
runs/RS-20260901-inversion-habit/critic.log, $0.078468 for the pair.
The seats converged, independently, on the same four things, and three of them changed the experiment rather than its wording.
1. PR2 was measuring the wrong contrast, and both seats said so (C1-A2, C2-A2) — ACCEPTED
v1's second primary compared Payne's rhyming hemistichs against his first hemistichs, so
position-in-bayt moved with obligation and the comparison could not isolate either. C2 also
noticed that v1 sampled N-P — rhyming hemistichs of odes with no tail, the correct control — and
then never used it.
Changed. PR2b is now O-P vs N-P, the same structural slot.
2. The obligated cell was partly definitionally inverted (C1-A3, C2-A3, C2-A12) — ACCEPTED, and it produced the new primary
This is the sharpest finding of the pass and it is correct. v1 defined obligation as Payne carries
a repeated English tail. Where the Persian radif is a finite verb and Payne carries it, the line
ends on that verb because the cell was selected that way, so an elevated INVERTED rate at
O-P is not evidence that he bought anything.
Two changes.
- Obligation is redefined from the Persian, not from Payne's English.
O-Pis now: rhyming hemistichs of odes whose Persian radif was coded finiteYES/ transitiveYESunanimously by the three blind seats ofE-20260831b-radif-handsstageG;N-Pis: rhyming hemistichs of odes whose Persian carries no radif at all (censusradif == ""). Both labels exist before any English final word is looked at, which is whatC1-A3asked for. - A new primary that cannot be definitional, and it is the better test anyway.
PR2is now the spill question: does a hand who commits to carrying a tail invert more in the rest of the poem, at line ends where nothing obliges him? Those line ends are not selected on any English feature. If carriage spills, the inversion is a grammar the poem gets written in; if it does not, it is bought where it is needed and nowhere else — which is exactly theLOCAL/SPREADdistinction the translation limb's frozen log grades at nine positions.PR2b, the rhyming-slot contrast, is kept as a secondary and is reported with its definitional component named.
3. PR1 cannot license HABIT across two different books (C1-A1, C2-A1, C2-A8) — ACCEPTED
Payne translated the Divan entire; Leaf chose 28 odes and writes a plainer English. A book-level rate difference confounds hand with register, with poem selection and with edition.
Changed, and the material turned out to allow the stronger version. Leaf prints his Brockhaus
number at the head of 26 of his 28 odes, and Payne's volume 1 is Brockhaus I–CC, so eleven poems
are rendered by both hands: Leaf III, IV, V, VI, VII, VIII, IX, X, XI, XII, XIII = Brockhaus 6, 8,
43, 44, 52, 79, 84, 121, 123, 151, 198. PR1 is now a matched-poem comparison over those eleven,
censused rather than sampled — every unobligated line end both hands wrote for the same poems.
Register remains confounded with hand and is named in the limits; poem selection no longer is.
4. Calibration by construction proves less than v1 claimed (C1-A7, C1-A8, C2-A4, C2-A11) — ACCEPTED
Both seats: hand-made verb-moves on canonical lines are saliency-matched and easier than natural
verse displacement; passing them shows a seat can detect an artificial contrast and nothing more.
C1 added that PR5 gated on CLASS, not on the primary outcome ORDER.
Changed. Two gates now, and the second is the one that decides.
- The 24 synthetic pairs are demoted to a floor: a seat that fails them certainly cannot code the real items. Gate ≥ 21 of 24.
- A real-item gold set of 40, drawn by the same seed from the actual study items, is coded by
the lead with hand, ode and cell labels stripped and frozen and committed before any seat is
dispatched. Registered gate: a seat votes only at ≥ 0.75 agreement with gold on
ORDER. The gold items are study items, so this costs no extra calls. - The circularity is printed, not hidden: the lead wrote the coding definition, so gold agreement measures whether the seats apply this definition, not whether the definition is right. That limit is in the result page.
5. The ORDER construct is elastic (C1-A6, C2-A7) — ACCEPTED IN PART
Changed: the seat is given a three-step decision procedure rather than a one-line criterion, and
must return WHAT_MOVED from a closed list — which forces the constituent analysis instead of
letting a global impression stand in for it. Not changed: no exhaustive edge-case manual, for
which there is no session budget; the real-item gold set is what catches a shared wrong rule, and
C2-A7's point that seat–seat agreement cannot catch one is why gold is now a gate.
6. PR4 rewarded the coder's own example (C1-A4, C2-A12) — ACCEPTED
Changed: PR4 is now stated over the independent WHAT_MOVED annotation
(OBJECT_OR_COMPLEMENT_BEFORE_VERB), not over CLASS, and it is reported separately for cells
where a carried verbal tail could force the answer.
7. Hand identity leaks through the diction (C1-A5, C2-A6) — ACCEPTED AS A PROBE
Payne's archaism is visible in any couplet of his, so nothing "blinds the hand" in the sense the word usually carries.
Changed: a hand-identification probe is dispatched before the coding run — each seat sees
four labelled couplets per book and then twenty held-out couplets, and guesses. The accuracy is
reported. It measures the leak; it cannot measure whether the leak biased ORDER, and the result
page says so. PR1's matched-poem restriction is what actually limits the damage, since both hands'
items in that comparison are renderings of the same eleven poems.
8. OCR is unaudited and it is the dependent variable (C1-A9, C2-A9) — ACCEPTED IN PART, AND NOT FULLY ANSWERABLE
The project holds the text layer only; there are no page images to audit against, exactly as at S236. What v2 adds:
- a mechanical final-token damage rate per hand and per cell — the share of line-final words
absent from CMUdict after the British-spelling and
-eth/-estrules, which is a proxy computed at precisely the dependent variable's location; - a registered withholding rule: if the two hands' damage rates differ by more than 0.10, the between-hand figure is reported as a range over the damaged items, not as a point;
- every sampled couplet printed in
runs/RS-20260901-inversion-habit/line-ends.md, so the coding is checkable against the book by anyone who has it; - a sensitivity analysis recomputing each primary with all
UNCLEARand all damaged-final-token items excluded.
This is weaker than either seat asked for and the result page says so. One thing does protect
the within-Payne primaries: PR2 and PR2b compare cells drawn from the same scan of the same
book, so scan damage is common to both arms.
9. The syllable instrument is known-weak for exactly the hand it measures (C2-A10) — ACCEPTED
CMUdict holds no -eth form, so the fallback carries more of Payne than of Leaf.
Changed: the fallback is validated against 30 hand-counted lines, 15 per hand, frozen before
PR3 is reported, and the counter's mean absolute error per hand is printed beside the result. The
stage-A miss rates are already measured — Payne 0.0921, Leaf 0.0780 — and are printed too.
10. Whole-book rates from a context-free lexicon (C2-A14, and C1 implicitly via PR5) — ACCEPTED IN FULL
Changed: the mechanical line-final word-class coder and every whole-book rate are removed from the
design, and PR5 is withdrawn. A context-free "can this word be a finite verb?" rule cannot
deliver a clause-level class, and a whole-book number that the seats never saw is not worth the
paragraph defending it. Only sampled and censused rates are reported, with intervals.
11. Statistics (C1-A11) — ACCEPTED
Changed: every rate carries a Wilson interval, every difference a Newcombe interval, and both primaries additionally carry an ode-clustered bootstrap interval (5,000 resamples of odes with replacement), because line ends within an ode are not independent. Outcomes are three-way: supported (interval excludes the bar on the predicted side), contradicted (interval excludes it on the other), indeterminate (interval spans it).
12. A null must be separable from a broken instrument (C1-A12, C2-A5) — ACCEPTED
Changed: a registered positive instrument control. Payne's rhyming hemistichs in odes where
he carries a tail that is itself a finite verb must come back overwhelmingly INVERTED — the
seats are looking at lines that visibly end on a verb with its object in front. If that cell is
below 0.70 INVERTED, the instrument is declared broken and no primary is reported, whatever the
gates say. This is the one cell whose definitional character, which §2 above treats as a defect,
makes it useful.
13. Thin Leaf cells (C2-A13, C1-A10) — ACCEPTED
Changed: Leaf's obligation split is dropped from the primaries entirely and reported as counts with a declared minimum denominator of 25 below which no rate is printed. Cell sizes are fixed before sampling.
14. Budget (C1-A13, C2-A15) — ACCEPTED
Changed: the ceiling is raised to $3.50 — the UTC day opened with the full $5.00 and no other session has spent — and three seats are fixed in advance, so no post-freeze fallback changes the voting rule. The worst case is now inside the ceiling.
What was not accepted
C1-A9/C2-A9's page-image audit — impossible with the materials the project holds, and recorded as a limit rather than as a discharged finding.C2-A6's "restrict to lexically neutralized lines" — neutralising Payne's lexicon would require rewriting his lines, which destroys the word order being measured.C2-A4's "double-coded by humans" — the project has one coder. The gold set is that coder, blinded to hand and cell, and it is labelled honestly as such rather than as a human standard.