Repository path: workshop/experiments/E-20260821-matched-shape/critic-response.md · rendered 2026-09-09
Page metadata (front matter)
| type | note |
|---|---|
| id | E-20260821-critic-response |
| status | frozen |
| created | 2026-08-21 |
| updated | 2026-08-21 |
| links | workshop/experiments/E-20260821-matched-shape/design.md, workshop/experiments/E-20260821-matched-shape/critic-findings.json |
Response to the pre-run critic — E-20260821-matched-shape
One P1 call, $0.109895, provider OpenAI, finish_reason: stop, dispatched against design
v1-pre-critic before any census body existed. Verdict NEEDS REDESIGN: 3 BLOCKING, 7 MAJOR,
1 MINOR. All eleven are accepted. Nothing was overruled. The design is v2-post-critic and the
materials were rebuilt; the run had not been dispatched when this was written.
Under the stopping rule (design §5, note bqp) no second critic round is bought. The rule permits one only if the first killed a numbered primary — it did — but the same rule forbids buying a round to widen a primary, and every remedy below narrows.
The three BLOCKING findings and what was done
1. A lead-built English window cannot separate the translator omitted this from the window
does not reach it, and a wide window can manufacture an answer as easily as reveal one. The
design's claim that width biases only toward ANSWERED was wrong, and the critic named F9 and F26
as the cases where a lead-ABSENT verdict was being handed to the seat pre-baked.
Accepted in full. The window is abolished. Every seat now receives the whole comparator chapter — Knatchbull ch. IX for Set A (2,278 words), ch. XII for Set B (2,552) — and finds the place itself.
materials/align_fanza.pyis superseded and kept only as a record of the rejected approach. The response schema also gains the critic's requested third state:unlocated, distinct fromomitted, so "I cannot find it" is not silently counted as "he cut it".
2. SHAPE was not operationalised at the level needed to audit a coding claim. Matched in
form leaves open whether shared inflection, shared derivation, matched word-class or a mere
parallel frame qualifies — and the lead's own three close calls (F59 answered on ‑tion, F38
refused on ‑ed/‑en, F34 refused on ‑ness/‑ment) prove the ambiguity is live. A seat number would
then measure the seat's threshold, not the lead's zero.
Accepted in full, and this is the change that most improves the run. The seat no longer sets a threshold at all. It reports (i) whether the material is rendered, (ii) the English members it finds standing at the matched positions, quoted, and (iii) which formal relation holds between them, chosen from a fixed per-class menu written before dispatch (
materials/build_items.pyRELATIONS). The strict and loose readings are then computed by the analyser from the seat's own reported relation.SHAPE's menu: identical-suffix and identical-inflection are strict; same-class-and-length is loose; frame-only and none are not matches. The lead's three close calls fall out of the menu mechanically rather than being argued.Not done, and declared: the critic also asked for the manual to be piloted on held-out loci before dispatch. That is a further round of spend on the apparatus and it was not bought; the consequence is that the menu's own reliability is unmeasured, and that is in §7.
3. Three prompted models voting cannot falsify or confirm a historical coding count. The critic
is right, and its own remedy names the alternative: limit this run to a descriptive
model-disagreement study; for a check capable of revising the published figure, recruit independent
bilingual human coders. Those coders are the thing this project has repeatedly named and cannot
reach (NEXT.md, named, not built).
Accepted in full. §4's falsification bar is struck. No published figure is revised by a vote. The run is a second, independent, non-lead coding, reported as such. The one route by which it can still amend
framework/v0.2§7.18 is not a vote and does not need one: if seats quote English at a locus and the quoted English exhibits a strict formal match, the quotation itself is the evidence, checkable by any reader against the printed 1819 text, and the handbook is amended on the quotation, not on the count. That is the only claim this run is permitted to make against the published figure, and it is written into §4.
The seven MAJOR findings
| # | finding | what was done |
|---|---|---|
| 4 | Set A's baseline rests on an unpreserved crosswalk, and dropping «اللبؤة»'s four SHAPE loci is not neutral |
Accepted. §2 now states the crosswalk and its assumption explicitly, and §4 no longer claims Set A validates 0 of 18: it covers 14 of 18 and is reported as such, with the four excluded loci named |
| 5 | Several glosses reproduced the Arabic's member structure and handed the seat an English template | Accepted. All 54 glosses rewritten as flat single sentences. Where the proposition is two things (G48, G52, F94) no gloss can avoid naming two, and that residue is declared — it leans toward finding a figure, i.e. against the finding |
| 6 | P2 compares raw rates across groups with different omission base-rates and different sampling |
Accepted. P2 is restated conditional on the material being rendered, is descriptive only, carries no test, and its 16-item arm's imprecision is stated with it |
| 7 | F1's SPLIT bar misses 2-to-1 disagreement; the 5-of-14 bar has ~13% power against a true 0.20 |
Accepted. F1 is rebuilt on majority availability and per-label agreement; the 5-of-14 bar is struck with finding 3; a 0-of-14 result is reported with its one-sided 95% upper bound of ≈0.19, never as confirmed |
| 8 | P3's 60% agreement bar is below the 85% a constant PLAIN answerer would score |
Accepted, and it was the sharpest finding in the set. P3 is replaced: the trivial baseline is computed and printed, and the prediction is stated on the six non-PLAIN cells of the lead's coding, with a full 3×3 confusion matrix reported |
| 9 | Set A was identifiable as a condition (no Arabic sentence, all one class, worse OCR) | Partly fixed. Presentation is now uniform: both sets carry the Arabic sentence, both use the same reflow of the same scan. Not fixed: the type label is shown on every item, and a model may recognise Kalīla wa-Dimna. Both are in §7 |
| 10 | Set A's comparator excerpts are OCR-corrupt, and differentially so | Partly fixed. Both sets now get the same treatment of the same 1819 scan, so the differential is gone; the corruption is not, and a checked diplomatic transcription of two chapters was not made. In §7, with G5 named |
The MINOR finding
11 — that a broad window raises false positives as well as sensitivity. Subsumed by the abolition of the window and by requiring the seat to quote the English it is judging, which makes every verdict auditable against the printed page.
What the critic bought
The design that goes out is not the design that came in. The primary is narrower — no falsification bar, no vote against a published figure — and the measurement is better, because the seat's threshold has been taken out of the number and replaced by a reported feature. The single most consequential line in the review is finding 8: the agreement bar as written could have been cleared by a procedure that answered PLAIN to everything.