Repository path: workshop/experiments/E-20260727b-arm-identifiability/critic.md · rendered 2026-09-09
Page metadata (front matter)
| type | note |
|---|---|
| id | critic-arm-identifiability |
| status | frozen |
| created | 2026-07-27 |
| updated | 2026-07-27 |
| links | workshop/experiments/E-20260727b-arm-identifiability/design.md, config/models.md |
Independent pre-run critic pass — E-20260727b
Model: openai/gpt-5.6-terra (panel role P1, non-Anthropic), provider read off the response, temperature 0.2, max_tokens 12,000. Latency 48.2 s, 4,406 in / 3,474 out, $0.065876875 (per-response usage.cost). Raw response preserved at runs/critic-P1.json.
Verdict: FREEZE-AFTER-FIXES. Six mandatory fixes, four recommended. All six mandatory fixes are accepted and incorporated in design v2; three of four recommended are incorporated, and the fourth is recorded as a limitation with a reason.
The six mandatory fixes, verbatim in substance, with disposition
| # | the fix | disposition |
|---|---|---|
| 1 | §3's central claim is only partly correct. With an all-treatment preference vector Pearson's r is undefined, not zero; and same-sign deltas can still yield a nonzero correlation if preference varies with delta magnitude. Narrow the claim to: M5 cannot detect a purely directional / sign-only confound where preference does not vary across items. | Accepted in full. §3 rewritten. This is the fix that matters most: the design's headline methodological claim was stated one step too strongly, and stating it correctly makes it narrower and true rather than broad and sloppy. |
| 2 | The exact binomial p-value is not valid with half-credit ties treated as Bernoulli trials, nor with dependent rep-triples. Use a sign test conditional on non-ties and label it descriptive, or a cluster-respecting test. | Accepted. v2 reports the sign test conditional on non-ties, labels every p descriptive, and reports the per-cell breakdown as the primary form. |
| 3 | The doubled two-sided adjustment covers larger-vs-smaller for one cue, not the maximisation over five cues that produces A*. Pre-specify one primary cue, or correct for the max. |
Accepted. C1 words is pre-specified as the single primary cue; A* over five cues is retained as descriptive only and carries no p-value. |
| 4 | M1_s ≤ A* compares two aggregate rates and does not show the cue rule reproduces the jury's item-level choices. Measure item-level agreement; limit the conclusion. |
Accepted, and it upgrades the design. v2 adds Metric B′ — item-level agreement between each cue rule's predicted winner and each juror's actual choice, per sense and per juror. The aggregate comparison is demoted to a necessary-condition check. |
| 5 | P2 is not genuinely pre-registered — its rationale depends on the already-seen MAIN word pattern — and P4's threshold is chosen around an already-known result. A rule written after seeing the motivating result cannot make that result confirmatory. | Accepted. P2 relabelled [POST-DERIVED]; P4 relabelled and explicitly barred from being reported as a met prediction. |
| 6 | C4's regex is wrong as frozen. [.!?]+(?=\s\|$) requires whitespace followed by a literal pipe. It does not count sentences. |
Accepted — this was a real bug. The escaped pipe was a markdown-table artefact that became part of the frozen text. v2 states every regex outside any table, unescaped, and C4 is corrected before the run. Disposition 9 exists for exactly this, and the pass caught it. |
Recommended, and what was done
- A skim-only task showing the cue is perceptually available, not just computable. Not done, and recorded as a limitation. There is no human in this loop, and the lead cannot serve as the perceiver for a measurement about whether the project's jurors can perceive something. v2 §10 threat 1 now says the design measures computable identifiability and that perceptual availability is unmeasured.
- Per-cell descriptive scores as primary rather than the pooled 12-pair rate. Done (and it converges with fix 2).
- P3 moved to a separately labelled planned single-case report. Done — it is no longer in the prediction list's arithmetic and is reported under its own heading.
- State how the rule handles visually tied cues. Done — v2 records that a delta of 0 is not a cue at all, reports the tie count separately, and keeps half-credit as a statistical convention rather than a claim about a perceivable tie.
What the critic found sound
The distinction between graded association and directional arm separability — retained, once §3's Pearson claim is narrowed. Metric A's larger-vs-smaller formulation as a descriptive identifiability measure, including the half-credit convention. The §0 disclosure of seen information, the separate treatment of the lead pair, the frozen extraction definitions, and the independent recomputation requirement.