Repository path: workshop/experiments/E-20260728f-nonlead-items/critic/dispositions.md · rendered 2026-09-09
Page metadata (front matter)
| type | note |
|---|---|
| id | E-20260728f-critic-dispositions |
| status | frozen |
| created | 2026-07-28 |
| updated | 2026-07-28 |
| internal-judgment-only | true |
| links | workshop/experiments/E-20260728f-nonlead-items/design.md |
Pre-run critic pass — findings and dispositions
openai/gpt-5.6-terra, provider OpenAI, finish_reason: stop, in 16,391 / out 7,521 of which 1,552 reasoning, $0.082018. Raw: critic.json, critic-reply.txt.
Critic choice, declared. The critic is also rater P1. That overlaps the preference stated at S045 and the reason is on critic.py's docstring: every model that is not a rater failed this task in this session or the last one — deepseek/deepseek-v4-pro returned finish_reason: length with no content at all here ($0.039651 for nothing), z-ai/glm-5.2 returned finish_reason: error truncated mid-object here, and moonshotai/kimi-k3 did the same twice at S044 for $0.291582 under the ledger's own "drop, do not retry" rule. The mitigation is checkable and is in the analysis: the three rater pairs are reported separately, so if the critique shaped the prompt to suit P1, P1's two pairs will behave unlike the P2–P3 pair.
Instrument caution, and it is the S043 caution recurring verbatim. The critic cited 27 item id / position pairs and 12 of them are wrong — W24 is position 30 and not 36, W33 is position 1 and not 33, W39 is position 34 and not 39, and so on. Its structural findings were then re-checked against the real item set and hold; several of its examples do not. This is the second consecutive experiment in which this model produced sound structure on fabricated particulars, and every claim below was re-derived before being accepted.
One of its factual claims was checked and is true: three source sentences carry two items each — 「今日車中貴介,寧復識戴笠人哉?」 (W33, W34), 「竭力辦裝…始得歸。」 (W38, W37) and 「異史氏曰:「置身青雲…」 (W32, W31). Its ids for the third pair were wrong; the fact was right. 37 distinct source sentences carry the 40 items.
Accepted and acted on before dispatch
| # | finding | what was done |
|---|---|---|
| 1.1 | There is no evaluable headline test. Prohibit "survives the leak fix", "replicates S043", "supports R1" | Accepted in full. Amendment A3 writes the prohibition into the design; the result page uses none of that vocabulary. analyse.py records Q2 as UNEVALUABLE rather than as a boolean |
| 1.2 | A1 silently moved Q3's population. Q3 is stated on stratum P; calling it "evaluable over the whole 40" is a different comparison | Accepted, and A1 was wrong. Q3 is now UNEVALUABLE alongside Q2 and Q4. The whole-set sham comparison is reported with exactly the secondary, non-pre-registered status the whole-set A→B comparison has. This is the finding that changed the most |
| 1.3 | The item set has no defined target population; agreement can rise because R1 offers an easy neither route on out-of-scope items |
Accepted in substance, declined in form. The blinded eligibility audit is declined (below); the reporting half is implemented — neither rate per condition, full A→B and A→C transition matrices, and a collapse diagnostic counting A-split → B-unanimous and A-split → B-unanimous-neither |
| 1.4 | composite does not test clause (c) without structured sub-parts |
Accepted as the relabel the critic itself offers. Q5 is now "use of the composite response category", not a test of clause (c), and the result page says clause (c) remains untested for a second run |
| 1.5 | The treatment is not isolated from instruction-force; report "R1 increased consensus" separately from "R1's boundary caused it" | Accepted. Transition matrices added; the result page separates the two claims |
| 1.6 | Fixed order; three sentences carry two items each; the item bootstrap treats dependent rows as independent | Accepted. A sentence-clustered bootstrap was added beside the item-level one (resampling whole source sentences), and a first-half/second-half position check. B2 keeps the identical order deliberately: it is measuring determinism at temperature 0, which is what the S043 figure it is compared against measured |
| 2.2 | Several English renderings leak a likely label — libation, road-offering, fat post, The Chronicler of the Strange, smacking and gulping | Accepted as a stated limitation, in the result page's limits section and in §9 of the design's own known limitations, which already carried the weaker form of it |
| 3 | Six checkable confounds and six uncheckable ones | Accepted. All six checkable ones are computed; the six uncheckable ones are listed on the result page as limits |
| 4 | Q4's inference is unsound anyway — lexical alternatives can be read as "alternatives at that site", and a raw vote share is not a measure of clause (a) firing | Accepted as a limitation of the frozen prediction. Q4 is unevaluable here in any case; the objection is recorded so a later design does not re-use the same inference |
| 5(b) | Reporting whole-set A→B as secondary is legitimate only if it never inherits Q2's rhetorical status | Accepted, and the critic's own suggested wording is used nearly verbatim on the result page |
Declined, with reasons
-
1.3's blinded eligibility audit (an independent model coding each item as containing a candidate formal marker / cultural item / both / neither, before the labels are looked at). Declined on three grounds, and the first is the strongest: it is another lead-briefed model coding lead-selected material, which is the very move whose credibility this run exists to question — a second instrument with the same defect does not repair the first. Second, the audit's purpose is to bound how much of any A→B gain comes from out-of-scope items, and the collapse diagnostic and the transition matrices bound that directly from data the run already produces. Third, this is the arm's last declared session and an extra call to shore up a comparison that is already labelled exploratory is not where the remaining budget belongs. The cost of declining is stated on the result page: there is no eligibility denominator, so no statement about what fraction of the item set is in scope for the seam is available.
-
1.4's structured-JSON
compositeschema. Declined. Changing the response schema across three models mid-run risks parse failures that would cost the whole run, and the five-way scheme is frozen in §4. The critic's own fallback — relabel Q5 — is taken instead, and it costs nothing but the claim. -
5(c)'s option 1, "stop and report failed material construction". Declined, and this is the judgment call of the session. The critic is right that the run cannot answer its headline and right that A1 left rhetorical room for a conclusion the data will not license. But the run can answer a different question that this item set is unusually well suited to, and that S043 left open:
RS-20260727e§3 records that under R1 chance-corrected agreement FELL while raw agreement rose, because 92.5% of votes went to one label. An item set that is largely out of scope for the seam is a strong test of whether R1 buys agreement by funnelling answers into a default. That is a mechanism question, not the headline, and the collapse diagnostic and transition matrices are built to answer it. The run dispatches for that, and for nothing else.