Repository path: workshop/experiments/E-20260821b-matched-heard/critic-response.md · rendered 2026-09-09
Page metadata (front matter)
| type | note |
|---|---|
| id | E-20260821b-critic-response |
| status | frozen |
| created | 2026-08-21 |
| updated | 2026-08-21 |
| links | workshop/experiments/E-20260821b-matched-heard/design.md, workshop/experiments/E-20260821b-matched-heard/critic-findings.json, wiki/method-notes.md |
E-20260821b — the pre-run critic's findings, and what was done about each
Round 1. P1 openai/gpt-5.6-terra, one call, $0.096713, verdict NEEDS-REDESIGN,
12 findings: 8 BLOCKING, 3 MAJOR, 1 MINOR. Raw in critic.json, parsed in
critic-findings.json. v1 of the design was never dispatched.
Its summary, verbatim: "The design does not presently supply valid minimal pairs, independent blinded arms, or a defensible recovery measure … Re-audit and rebuild the stimulus set and allocation before dispatch; reporting limitations afterward cannot repair these defects."
One finding is overruled, on a fact about the runtime; eleven are accepted, seven of them changing the design.
| # | sev | finding | disposition |
|---|---|---|---|
| 1 | BLOCKING | each seat receives both versions of every passage, so the between-version comparison is within-subject and open to memory and demand characteristics | OVERRULED on its premise, accepted in its consequence. Each dispatch here is an independent completion at temperature 0 with no conversation history: a seat cannot remember its earlier answer because it is never shown it. What survives is the statistical dependence — three model seats are not 90 independent observations — and that is finding 8's clustering point, which is accepted. The overrule and its ground are recorded here rather than left to the result page. |
| 2 | BLOCKING | many plain wordings still carry a match under the same four-resource table |
ACCEPTED, and it is the finding that reshaped the design. §3.2 now excludes on an independent judgement rather than the lead's audit: a blind seat is shown each plain span's members and asked whether they echo in form, and any locus it says do is out of the primary. The lead no longer decides eligibility. |
| 3 | BLOCKING | the §3 audit contradicts itself (excludes F55, keeps F38) and misses framed plain spans |
ACCEPTED. The lead's audit is demoted from an eligibility rule to a record of what the lead thought, printed in the result beside the independent screen so the two can be compared. Two of the critic's specific charges are wrong on the facts — slaughter is disyllabic, and pursuit does not end in ‑tion — and are noted here without altering the disposition, because the general charge holds. |
| 4 | BLOCKING | M8/M9 have their two members in different passages, so no reader of one passage can recover them, yet they are counted as covered |
ACCEPTED. build.py no longer closes a window while a locus still has a span to come. 15 passages, no locus split, verified by the builder. |
| 5 | BLOCKING | the five adjust entries change content, grammar or naturalness and must not be called frozen refusals |
ACCEPTED IN FULL, by the critic's own first remedy. Every locus carrying an adjust field is excluded from the primary — F59, F65, F89, M9, M12 — and reported separately. 22 loci go into the eligibility screen, not 27. |
| 6 | BLOCKING | the content screen is one model on four planted errors and its failure does not stop the primary | ACCEPTED. G1 is now a hard gate: if fewer than 3 of 4 planted errors are caught, the primary is withheld, not caveated. A length check is added: the passage word counts of the two versions are reported per passage and the primary is withheld if any pair differs by more than 15%. |
| 7 | BLOCKING | the substring scoring rule is gameable — ["the", "him"] scores RECOVERED |
ACCEPTED. Scoring now matches annotated member positions, requires a shared content token (a frozen stopword list), and requires two distinct members. The critic's exact counter-example is used as a mutation test in verify.py. |
| 8 | BLOCKING | the sign test ignores magnitude, treats loci as independent though they share passages and three seats, and can pass while being uninformative | ACCEPTED. The primary is now a permutation test with the version label permuted within passage — the unit that was actually assigned — reported with the paired difference itself, and PR1 carries a minimum effect (≥ 0.30) and PR3′ a plain-arm ceiling (≤ 0.50). |
| 9 | MAJOR | the PR5 keyword classifier is not a valid identification measure and has no chance baseline |
ACCEPTED. PR5 is withdrawn as a prediction. The property strings are reported verbatim and classified only descriptively, with no threshold and no claim. |
| 10 | MAJOR | declaring the (C)/source confound does not make the resource question interpretable | ACCEPTED. PR4 is withdrawn and every resource-comparison prediction with it. §1's second question — is the answer the same for all four? — is struck from this experiment and left to the arm. The result is a material-set result. |
| 11 | MAJOR | retained excluded figures and multiple targets compete for the six-group cap | ACCEPTED. The cap is raised from six to twelve, truncation is counted and reported, and the rebuilt windows put at most four eligible loci in a passage. |
| 12 | MINOR | the prompt primes formal recurrence, so the estimand is prompted detection, not spontaneous noticing | ACCEPTED. The estimand is now named as prompted detection wherever it is stated. |
No second critic round was bought in this session. Note (bqp) permits one where the first
round kills a numbered primary, and this one killed two. The redesign is instead exposed in the
other direction: the eligibility screen is bought from a seat that is not the lead and not the
critic, and the primary is withheld by F2′ if that screen leaves too few loci — so the design can
fail on evidence that the lead does not control.