Repository path: workshop/experiments/E-20260804b-quixote-affect/critic.md · rendered 2026-09-09
Page metadata (front matter)
| type | note |
|---|---|
| id | E-20260804b-critic |
| status | frozen |
| created | 2026-08-04 |
| updated | 2026-08-04 |
| links | workshop/experiments/E-20260804b-quixote-affect/design.md |
Pre-run critic pass — E-20260804b
One pass, one call. moonshotai/kimi-k3 (P4), provider and cost in
runs/critic_kimi-k3_try1.meta.json — $0.07368, 201.0 s, 4,430 chars. Verdict
NEEDS-AMENDMENT: 5 BLOCKING, 5 ADVISORY. Raw body at
runs/critic_kimi-k3_try1.raw, verbatim text at runs/critic_kimi-k3_try1.txt.
All ten findings are accepted. None is refused. The amendments below were written into
design.md before any grading call was dispatched, and the design's SHA and materials hash were
re-frozen after the rebuild. Amendment ids A1–A10 correspond to findings 1–10.
| # | kind | the finding, compressed | amendment |
|---|---|---|---|
| 1 | BLOCKING | PR5 was near-unfalsifiable — with 42 picks over 6 arms, modal-FUNNIEST ≠ modal-MATCH is the default under noise, so the only falsifying outcome was exact modal identity |
A1: PR5 restated directionally — modal FUNNIEST ∈ {motteux,nudge} and modal MATCH ∈ {shelton,smollett,ormsby,lead} — and the uniform-null probability of a bare modal split is computed and reported alongside |
| 2 | BLOCKING | Even a clean PR5 could not support the inference, because Stages 1 and 2 differ in task, source-presence and key order at once | A2: a within-stage contrast added — Stage 1 now asks FUNNIEST-SP in the same call, on the same key order, immediately after MATCH. That pair is the crux test; Stage 2 becomes its source-absent replication. The cross-stage inference is explicitly softened in §1 |
| 3 | BLOCKING | Training-data recognition was undeclared — four canonical published arms, and Ormsby's introduction with its verdict, are almost certainly in every grader's training data, so letter keys do not prevent recognition and the seats may recite the record | A3: a recognition probe added at the end of Stage 2 (not Stage 1, so no authorship question is anywhere in the prompt while SIGNAL/MATCH is being answered): name the translator per key or write UNKNOWN. Rates reported. New declared limit L6 |
| 4 | BLOCKING | No prediction tested the motivating reception fact — nothing predicted that motteux beats anyone on funniness |
A4: PR6 added — motteux FUNNIEST picks > ormsby FUNNIEST picks |
| 5 | BLOCKING | F5's exclusions rested on one seat, one call, no replication — and S5 is exactly the site that tests whether that seat is right, since the Spanish there does carry the aside «y fuera mejor que se curara, porque fuera curarse en salud» | A5: Stage 0 run on two independent source-only seats — P5 and the config/models.md first reserve qwen/qwen3.7-max. A site is excluded from PR2/PR3 if either seat says the Spanish itself signals — excluded by default on disagreement, as the finding asks |
| 6 | ADVISORY | The PR4 manipulation confounded adds a narratorial signal with adds a new joke — "more fool he", "a pretty prayer it was" are themselves gags | A6: a seventh arm, nudge-flat — lead plus an intrusive narratorial comment of matched length that is evaluative but not comic. New predictions PR4a (the operator is detected when it is not funny) and PR4b (nudge > nudge-flat on funniness) |
| 7 | ADVISORY | F6's length check used a Pearson r over six points per site and never isolated the pair PR4 compares | A7: F6 replaced by a pair-level check — per site, nudge−lead pick difference against nudge−lead word difference; and nudge-flat is length-matched to nudge by construction, which handles it structurally |
| 8 | ADVISORY | SIGNAL carried PR1–PR3 with no seat-agreement criterion (F4 covered MATCH only) |
A8: F8 added — majority attainment computed over every (site, arm, ordering) cell; below 0.80, SIGNAL is descriptive only |
| 9 | ADVISORY | F3 was so aggressive that PR4/PR5 were likely to be withheld by chance | A9: F3 now fires only where the order effect reverses the sign of the relevant pick difference; otherwise both orderings are reported and pooled |
| 10 | ADVISORY | At S3 Motteux omits the passage and cannot signal there, depressing pooled SIGNAL(motteux) by an operation the design itself excludes from SIGNAL |
A10: PR2 pre-registered to be reported both with and without S3, rather than decided after the data exists |
What the pass cost and what it bought
Finding 3 is the one this session would not have caught. The design had a
RS-20260804-yardstick-shaped defect it had already guarded against in the obvious place — no
lead-authored gloss, an independent mechanism statement — and missed the other channel by which a
canonical text can carry its own verdict into a grader: the graders may already know both the
translations and the criticism of them. No amendment can remove that; A3 measures it and L6
declares it.
Finding 5 is the one that is materially right about a specific site: the Spanish at S5-no-heed is
No se curó el arriero destas razones (y fuera mejor que se curara, porque fuera curarse en salud),
and that parenthesis is a narratorial aside. If Stage 0 says so, S5 leaves PR2 and PR3 — which is
the design working, not failing.