Repository path: workshop/experiments/E-20260728d-jeli-marks/critic/dispositions.md · rendered 2026-09-09
Page metadata (front matter)
| type | note |
|---|---|
| id | E-20260728d-critic-dispositions |
| status | frozen |
| created | 2026-07-28 |
| updated | 2026-07-28 |
| links | workshop/experiments/E-20260728d-jeli-marks/design.md, config/budget.md |
Pre-run critic dispositions — E-20260728d
Critic: openai/gpt-5.6-terra (P1), one call, finish_reason: stop, in 4,493 / out 3,878,
$0.036104375 (worst case from max_tokens 8000: $0.131 — landed at 28%). Raw body:
../runs/critic__openai_gpt-5.6-terra.attempt1.json. Key snapshots: ../runs/snap/.
Verdict: NEEDS-REDESIGN. Ten findings, four blockers. Its one-sentence summary — "G8 cannot settle
the two accounts because P4 is non-exhaustive, lacks account-specific probabilistic predictions, and
converts one three-model site verdict into a causal adjudication" — is correct, and the design has been
changed to say so rather than argued with.
The single most consequential finding is G2, and it is the same defect this project shipped last
session. §5 said "the U6 row is settled by whatever P4 returns, including SPLIT, which settles it as
undecidable on this instrument." A rule that fires on every possible outcome has a false-alarm rate of 1.
That is note (bcs) — S046's B2 statistic, which could not come out zero — recurring one session later
in a different dress. It is repaired below.
| # | sev | accepted? | what changed |
|---|---|---|---|
| G1 | blocker | accepted | §7's 0.000250 is re-derived by the critic and confirmed (0.000250293), but it is relabelled: it is a random-response reference value, not a false-alarm rate under the substantive null. The substantive null's response distribution is unmeasured and the design now says so. Independence of models, sites and calls is now stated as an assumption, not a fact. |
| G2 | blocker | accepted in full | P4 is made exhaustive with a no-information branch. A → D62's necessity claim fails; B → consistent with D62; D → a mechanism neither account named; C, E or SPLIT → no information, and U6 is recorded as not settled by this instrument. The sentence "settled by whatever P4 returns" is deleted. A null outcome is a null, not a settlement. |
| G3 | blocker | accepted; repaired in part, and the rest conceded | Prior attribution is now mechanically defined (the marked English string, ≥4 words, recurring earlier inside attributed speech) and verify.py computes it rather than reading it off the table. The criterion decides exactly the two sites it must decide — G7 yes, G8 no — and is inapplicable at the three bare-NP sites, where "no prior attribution" remains the translator's reading and is now labelled as such. Proverb remains a translator judgment against a written rule. No independent annotator was available in this environment; the finding is conceded, not answered, and the affected predictions are labelled accordingly. |
| G4 | blocker | accepted; the design's central claim is withdrawn | G8 is no longer a crucial experiment between two accounts. §2 shows the proverb account has no instance under V7, and the critic is right that this makes it unfit to be an equal rival rather than a fit one. G8 is now a one-sided test of D62 alone: D62 claims prior attribution is what makes the marks work, so an A at a site with none would show it is not necessary. The alternatives the critic names — retained source punctuation, idiom marking, quotation convention — are added to §8 as unexcluded explanations of an A. |
| G5 | major | accepted | P5's unit is defined (mean of 3 model Q3 ratings per site, then mean over sites in each set) and P5, P6, and Q3 throughout are labelled descriptive, with no test. P6's SPLIT handling is specified. No prediction claims computable power, and §7 now says so in one line instead of implying it. |
| G6 | major | accepted | The erratum decision rule is removed in the no-change direction. P6 is one-sided: a change in modal category is a positive signal that mark type costs something; no change establishes nothing, for want of an equivalence margin and power. The words "whether the divergence costs anything is measurable" are struck. |
| G7 | major | accepted | Q1 gains a fifth substantive option — D = fixed expression (a proverb or set phrase; conventional language rather than any particular person's words) — because the critic is right that a proverb is borrowed language without being anyone's actual words, and A conflated the two. Null arithmetic recomputed on five options. On the artefact risk (showing a phrase already in marks makes A salient): the design's answer is that this is exactly what the bare-NP arm of the gate controls — if A were an artefact of displaying marks, G3–G5 would return A too and the gate would fail. That answer is now written into §7 rather than left implicit. |
| G8 | major | accepted | Question order changed: Q2 (whose words — open, unprimed) is asked first, then Q1, then Q3. Q2 becomes the primary measure for D62; Q1 stays primary for D21. Order is fixed, not randomised — randomising would triple the calls — and the residual priming of Q1 by Q2 is declared. The context-window point (a prefix-fed model is not a sequential reader) is conceded and added to §8. |
| G9 | major | accepted | The claim that more context favours A, and therefore that a B at G8 is "the harder result", is deleted. Position is confounded with context length, plot, salience and attention, and no direction is asserted. |
| G10 | major | accepted; not repairable here | G8's English is written by the translator who knows the predictions, which were frozen first. The repair the critic names — a blinded independent translator producing pre-registered alternatives — does not exist in this environment. G8 is therefore labelled exploratory, never confirmatory, the rendering constraints that bind it (V3 literalism, D63's bound horns, V7's mark, no frame, no gloss) are declared in §3 before it is written, and the residual freedom is recorded as an unremoved confound. |
Nothing was declined. Ten findings, ten accepted, four of them by withdrawing something the design
claimed. What survives as pre-registered and confirmatory is P1, P2 and P3 — tests of D21's written
cost claim and of D62's own mechanism at its own site, on prose frozen in earlier sessions that the lead
cannot now change. G8, which is the reason the experiment was designed, is the weakest thing in it,
and that ordering is the critic's, not the lead's.