Repository path: workshop/experiments/E-20260804g-yardstick-repair/critic.md · rendered 2026-09-09
Page metadata (front matter)
| type | note |
|---|---|
| id | E-20260804g-critic |
| status | frozen |
| created | 2026-08-04 |
| updated | 2026-08-04 |
| links | workshop/experiments/E-20260804g-yardstick-repair/design.md |
Pre-run critic — z-ai/glm-5.2, verdict NEEDS-REDESIGN
Dispatched against the frozen design at commit c1baaad, before any subject seat was addressed.
Ten findings, two BLOCKING, all ten accepted; the amendments are design §§13–14.
This is the second critic seat. P4 moonshotai/kimi-k3 was dispatched first, returned a
non-JSON keep-alive body (note (bgc), unbilled), and on retry returned finish_reason: length with
12,000 completion tokens, every one of them reasoning, and zero characters of content, billed at
$0.2128698. Under note (b) that is a seat failure and never a partial answer, and nothing from
it is adopted here. Note (bhf), fifth firing. The seat below cost $0.03770906.
Findings verbatim as returned:
FINDING 1 | BLOCKING WHAT: The REAIM arm is written by the lead after seeing A1's relation statement — the answer key — and is then placed into the same grading task as FROZEN, FORCED and DECOY. The design declares this but does not control it: graders now see five renderings where the original saw four, and one of the five (REAIM) is a near-guaranteed YES that can anchor the grader's sense of what "carrying the relation" looks like, depressing YES rates on the subtler arms. WHY: A grader who sees an obvious carrier (REAIM) alongside FROZEN may become more stringent on FROZEN, inflating the separation the run is measuring. DECOY is the stated detector, but DECOY >1 is a coarse gate that fires after the damage is done and only declares the run "descriptive" — it does not protect the primary. FIX: Drop REAIM. If FORCED scores low because it was aimed at a different relation than A1 names, that is itself the finding — it means the 2026-08-02 repair was self-serving. Note the divergence as interpretation, not as a new arm.
FINDING 2 | BLOCKING WHAT: A1's "independent" statement is seeded by two lead-written inputs the design holds fixed: the gloss and the construction-under-study pointer. The glosses already encode the relation at most sites — S2's gloss says "you[plural]" and "lovers," S4's says "little-dear-earth" and "little-wretched-yurts," S5's says "your-servant," S6's says "daughters rather than sons" and "feminine." The pointer for S2 ("the pronoun the speakers use for each other") names the device in everything but the banned word. WHY: A1 may write essentially the same relation the lead wrote, not because the relation is obvious from the source but because the gloss already states it. The "yardstick repair" then repairs nothing, and a reproduced separation would be uninterpretable: it could mean R1 works, or it could mean A1 was handed the answer. Conversely, if A1 writes something different, it may be because A1 is a weaker reader, not because independence matters. FIX: Give A1 the source and a minimal word-by-word parse that does not interpret (no "little-dear-earth," no "your-servant," no social context). Drop the construction-under-study pointer entirely, or replace it with a non-directive prompt ("what does this passage convey beyond the literal meaning of its words?"). This is more expensive but without it the variable is not actually varied.
FINDING 3 | ADVISORY WHAT: The leak screen is a banned word-list that catches "formal," "pronoun," "diminutive" etc., but a statement like "the speakers use the distant language of strangers with each other" passes the screen and hands the graders the answer just as effectively. The screen tests vocabulary, not content. WHY: A statement that passes the screen can still name the relation so plainly that any grader can match it to the rendering that carries it (FORCED with "Pray," FROZEN with plain "you"). The screen creates a false assurance that the yardstick is clean. FIX: Add a human or second-seat check: does the statement, read alongside the five renderings, make the correct answer obvious? Any statement where a reader can predict which rendering carries it from the statement alone should be re-requested. This is cheap (one pass, no API cost if lead-coded with a frozen rubric) and catches what the word-list cannot.
FINDING 4 | ADVISORY WHAT: The lead has registered an expectation that R2p and R3p fail. Every outcome is now a win: if they fail, the lead predicted it and gets credit for an honest null; if they hold, the lead was wrong and the warrant is stronger than expected. There is no outcome the lead has left himself unable to spin as informative. WHY: A pre-registered expectation of failure is not a prediction that can be falsified — it is a hedge that converts the null from a disappointment into a confirmation. The run's informativeness is supposed to come from the predictions, not from the lead's stance toward them. FIX: Separate the prediction from the expectation. State a specific outcome that would be uninformative or embarrassing and commit to reporting it as such — for example, "if FROZEN rises to exactly 2–3 and FORCED drops to exactly 3–4, the run is inconclusive and no amendment is made either way." Give yourself a lose condition.
FINDING 5 | ADVISORY WHAT: F1 uses POSITIVE (≥5/6) as half of its instrument-failure gate, but POSITIVE is written by the lead from A1's own words and can always be made explicit enough to pass. The POSITIVE half of F1 is nearly unfireable by construction, so F1 reduces to a DECOY-only check. WHY: If the grading instrument is broken in a way that affects POSITIVE and the primary arms equally (e.g., A1's statements are incoherent and graders default to YES on everything), POSITIVE will still pass because the lead can write an explicit gloss of anything. F1 then fails to catch the failure. FIX: Have POSITIVE written by a second independent seat from A1's statement, or add a floor on the absolute YES rate: if all five arms score YES at ≥5/6 sites, the task is too easy and the run is descriptive regardless of DECOY.
FINDING 6 | ADVISORY WHAT: R6p (REAIM ≥4/6) is not a prediction but a near-certainty: the lead writes REAIM knowing the exact statement the graders will score against, under a brief that requires the relation to appear. R6p can only fail if the relation is not expressible in English at all, which would be a finding about R1's text, not its warrant — and the design has no failure criterion for that case. WHY: A prediction that cannot meaningfully fail is decoration. Including it makes the prediction table look more rigorous than it is, and if R6p fails, the design has no pre-committed consequence. FIX: Either drop R6p (if REAIM is dropped per Finding 1, this is moot), or add a failure criterion: if R6p fails, R1's text is qualified, not just its warrant.
FINDING 7 | ADVISORY WHAT: The two orderings per seat are the first ordering and its exact reverse. A seat that says YES in ordering 1 is highly likely to say YES in ordering 2 because the content is identical and only sequence changes. The "both orderings" requirement is treated as a robustness check but is closer to a test-retest of the same model on the same inputs. WHY: The recovery threshold (≥2/3 seats, both orderings) is stricter than it appears in one direction (it filters out order-sensitive noise) but does not provide the independence that two genuinely different orderings or two different prompts would. A seat that is systematically biased toward YES will pass both orderings and count as a recovery. FIX: Use two structurally different prompt templates (not just reversed order) for the two orderings, or acknowledge in the design that the two-ordering check detects only sequence effects, not prompt sensitivity.
FINDING 8 | ADVISORY WHAT: Two measurements — the A1/A2 stability coding (SAME/OVERLAPPING/DIFFERENT) and the A1-vs-log divergence coding — are coded by the lead and committed before grading, but nothing checks the coding itself. A bad-faith or self-flattering lead could code everything as SAME (showing stability) and as matching the log (showing the original repair was well-aimed). WHY: These measurements are declared as supporting evidence in the run's report. If they are lead-coded with no second check, they carry the same conflict of interest the run is supposed to be repairing. FIX: Have A2 or a third seat code the SAME/OVERLAPPING/DIFFERENT judgment independently, or drop the measurements from the report and keep them only as internal diagnostics. If they are reported, they need the same independence standard as the relation statement.
FINDING 9 | ADVISORY WHAT: A null result (no separation) is ambiguous between two causes the design does not distinguish: (a) the yardstick's authorship matters — the lead's statements were self-serving; (b) A1 writes vaguer or less precise statements than the lead did, collapsing all arms toward the same score. The design treats a null as evidence for (a) but it could be (b). WHY: If A1's statements are broader ("the passage conveys social distance") than the lead's ("these two people address each other in the manner reserved for people who are not on intimate terms"), graders will say YES to more renderings regardless of authorship. The amendment to the framework release would then be based on a confound. FIX: Add a precision check: have a third seat rate each A1 statement and each lead statement (from 2026-08-02) for specificity on a fixed scale, committed before grading. If A1's statements are systematically less specific, report the null with that caveat rather than amending the release outright.
FINDING 10 | ADVISORY WHAT: The consequence for a null — amending the framework release to say R1's machine-read warrant "does not reproduce" — is stated as a single binary action, but the run's own §9 says six sites are a census, not a sample, and three LLMs are not a survey. A null from three graders on six sites with a possibly-seeded yardstick (Finding 2) and a possibly-vaguer yardstick (Finding 9) is thin ground for amending a published release. WHY: The consequence is too strong for what the run can support if the yardstick is compromised. It is too weak if the yardstick is clean and the null is real, because it leaves R1's text standing "with less under it" — but if the warrant was self-serving, the text itself is suspect. FIX: Make the consequence conditional on the yardstick's independence being verified: if A1's statements are shown to be independent (Finding 2 addressed) and comparably specific (Finding 9 addressed), amend; otherwise, report the null as suggestive and run a second replication before amending.
VERDICT: NEEDS-REDESIGN