Repository path: workshop/experiments/E-20260805g-printed-switch/critic.md · rendered 2026-09-09
Page metadata (front matter)
| type | note |
|---|---|
| id | E-20260805g-critic |
| status | frozen |
| created | 2026-08-05 |
| updated | 2026-08-05 |
| internal-judgment-only | true |
| provisional | true |
| links | workshop/experiments/E-20260805g-printed-switch/design.md, wiki/arms/ARM-atelier-cycle.md, config/models.md |
E-20260805g — three pre-run critic passes, thirty-one findings, all accepted
Independent adversarial critic, x-ai/grok-4.5 (P3), which is not one of the five seats.
Every pass was dispatched before any seat call. Raw bodies:
runs/critic_grok-4.5_try1.raw, ..._re2.raw, ..._re3.raw.
| pass | verdict | findings | BLOCKING | cost | body |
|---|---|---|---|---|---|
| 1 | NEEDS-REDESIGN | 9 | 6 | $0.0169824 | critic_grok-4.5_try1 |
| 2 | NEEDS-REDESIGN | 13 | 9 | $0.0222444 | ..._try1_re2 |
| 3 | NEEDS-REDESIGN | 9 | 7 | $0.0210564 | ..._try1_re3 |
Total $0.0602832 — 14% of the run's declared worst case, and it changed the run's question twice.
No pass was dispatched after the seats. No fourth pass was dispatched, on the policy frozen in
design.md §9 before pass 3: a gate that can be re-run until it passes is not a gate.
Pass 1 — the primary measured orthography
Revision 1's primary asked seats, directly, whether anyone present could not follow what was said, and predicted YES on the arms with a Swedish string and NO on the arm without.
- F1 BLOCKING — true by construction. "Any model that can detect non-English orthography will answer YES on A1/A3 and NO on A2 without recovering social exclusion, reader-seat inversion, or Canth's device." ACCEPTED. This is the finding the whole design turns on. → the primary became an open retelling.
- F2 BLOCKING — the prompt planted the concept. "a cued hunt for code-switch exclusion." ACCEPTED. → the retelling was made uncued.
- F3 BLOCKING — ¶453 is a second route to the content. ACCEPTED. →
CONTENTanswers scored for source paragraph; later superseded when the cued block was deleted. - F4 BLOCKING — the
INTENTprediction was unfalsifiable at unknown n. ACCEPTED. - F5 BLOCKING — the arms are not one ordered factor. ACCEPTED. → A1 vs A2 and A1 vs A3 reported as two separate contrasts.
- F6 BLOCKING — the failure criteria could not fire on plausible garbage, and the
RANKcontrol is near-guaranteed to pass. ACCEPTED. →RANKeventually deleted. - F7, F8, F9 NON-BLOCKING — pre-register the full rejection region; the orientation's "parish pastor" / "town doctor" primed a professional-hierarchy reading; A3 was not a minimal edit. ALL THREE ACCEPTED; the orientation was stripped and A3 later rebuilt as a two-word insertion.
Pass 2 — the scene carries the inference without the line
- F1, F2 BLOCKING — the
SHUT-OUTcode was an artifact in both directions. A correct two-sentence recovery would score 0 under a same-sentence rule; "The pastor asked in Swedish whether Mari could be cured" would score 1 with no exclusion content at all. ACCEPTED. → coded over the whole retelling, and an explicit exclusion predicate required. - F3 BLOCKING — the cued block contaminates any use of its data. ACCEPTED. → quarantined, then deleted at pass 3.
- F4 BLOCKING — and this is the pass's best finding. "Doctor avoids Holpainen's questioning eyes; pastor gets no answer; full silence while the doctor writes; pastor's 'nothing more for us'; Heikura must re-ask in plain language. Readers can infer side conversation / layperson shut out of the medical exchange without ever coding language." ACCEPTED, and it rewrote the question: the Englished arm stopped being a control and became the baseline the run exists to measure.
- F5 BLOCKING — the primary can succeed while the question fails. ACCEPTED. → the
LANGcode demoted to a manipulation check with no inferential weight. - F6 BLOCKING — the content prediction is true by construction for any honest seat. ACCEPTED. → dropped as support.
- F7 BLOCKING — the event control cannot do the job assigned to it. ACCEPTED. → retained but stated as weak, with nothing resting on it.
- F8 BLOCKING — n = 5 thresholds are decorative. ACCEPTED, and acted on fully at pass 3.
- F13 BLOCKING — the plausible path where every criterion fires and nothing is learned. ACCEPTED.
- F9–F12 NON-BLOCKING — A3 a confounded double edit; the orientation still primes; the keyword list leaky and brittle; the cost estimate should be line-itemed per model. ALL FOUR ACCEPTED.
Pass 3 — the uncued primary was not uncued
- F4 BLOCKING, and decisive. "'Do not go back and alter Part 1' does not stop a single-pass model from reading the whole prompt first; Part 1 is therefore not uncued. EXCL list mirrors Part 2's wording, so keyword scoring rewards prompt-echo." ACCEPTED. This is plainly right and it is repaired at zero cost: Part 2 was deleted from the payload. The dispatched body is a retelling and nothing else.
- F2 BLOCKING — reader-opacity is not in-scene exclusion. "Seats can score EXCL on A1 because they
cannot read the quote, not because they inferred social exclusion." ACCEPTED. →
OPACITYandEXCLsplit into independent codes, andEXCLnow requires a person in the room to be named. - F1, F7 BLOCKING — the rejection region is nearly unreachable and the design is non-decisive. ACCEPTED IN FULL: the confirmatory test was withdrawn before dispatch. No p-value is computed anywhere in this run. The critic's own remedy — "treat A1 vs A2 as descriptive counts only with an explicit non-claim" — is what was done.
- F8, F9 NON-BLOCKING — confirm billable model ids; a retelling that introduces Swedish or Finnish
with no on-page warrant should be inspected as contamination even in A2. BOTH ACCEPTED; the
second became gate
G1.
The three findings that were NOT repaired, and travel as limits
Stated here so that the result page cannot quietly drop them.
- Pass 3
F3— the scene's four non-language routes cannot be stripped without rewriting Canth. True. The design's answer is to measure the routes rather than remove them, which is what the baseline is; it is not a repair. - Pass 3
F5— keyword coding of free prose is brittle in both directions. True. Blinded dual coding by independent seats is what would fix it, and it was not affordable in this session's remaining headroom.verify.pymeasures the brittleness with non-keyword acceptance tests instead of assuming it away, and every retelling is quoted verbatim so the coding can be checked by hand. - Pass 3
F1/F6— n = 5 supports no estimate of magnitude, and the baseline has no failure mode. True. The result reports counts out of five and claims nothing about size.