Repository path: workshop/experiments/E-20260810z-idiom-reach/critic.md · rendered 2026-09-09
Page metadata (front matter)
| type | note |
|---|---|
| id | E-20260810z-critic |
| status | frozen |
| created | 2026-08-10 |
| updated | 2026-08-10 |
| links | workshop/experiments/E-20260810z-idiom-reach/design.md, wiki/arms/ARM-idiom-reach.md, wiki/method-notes.md |
E-20260810z — pre-run critic, findings and adjudication
openai/gpt-5.6-terra, reasoning disabled, cap 16,000, finish_reason: stop, $0.02704075.
Dispatched over the frozen design.md, the frozen R24 regime page, the frozen code.py, and
every prompt string that would be sent, including the REVISE block whose absence is the
manipulation under test. Raw body: runs/stage0-critic.json.
Verdict:
NEEDS-REDESIGN. 16 findings — 7 BLOCKING, 7 MAJOR, 2 MINOR.
Adjudication: 12 accepted, 4 accepted-in-part, and 5 individual remedies overruled in writing. Every amendment below was applied before any generation call went out; the only body that existed when the critic ran was the critic's own.
The findings, and what was done
| # | sev | the finding, compressed | ruling |
|---|---|---|---|
| 1 | BLOCKING | Asrc vs the frozen ∅ conflates the +I clause with independent-pass variability: ∅ was generated in a different call, at a different time |
ACCEPTED — A1 |
| 2 | BLOCKING | source-first generation removes the revision constraint as well as adding the permission, so a positive Asrc cannot be attributed to the permission |
ACCEPTED — A1 |
| 3 | BLOCKING | G1 checks only respelling; an arm could gain loc through added material or mistranslation |
ACCEPTED IN PART — A2; the "exclude non-compliant cells / fail the run" remedy overruled |
| 4 | BLOCKING | G2 treats any byte difference as evidence the manipulation happened; a synonym satisfies it |
ACCEPTED — A3; the "blinded item-level coding" remedy overruled |
| 5 | BLOCKING | Lsrc is not a valid positive control: it is instructed to maximise the very thing loc measures |
ACCEPTED — A4 |
| 6 | BLOCKING | G3 (loc(B) ≥ 0.40) validates the instrument on respelling, which is not what P1 is about |
ACCEPTED — A5 |
| 7 | BLOCKING | all arms of a site in one call makes arm identity inferable and licenses comparative scoring | ACCEPTED IN PART — A10; the "one rendering per call" remedy overruled |
| 8 | MAJOR | the 60 ⟨site, hand⟩ denominator ignores clustering by site; McNemar has no multiplicity handling | ACCEPTED — A8 |
| 9 | MAJOR | F3/F5 allow the denominator to change after arm-specific failures are observed |
ACCEPTED — A7 |
| 10 | MAJOR | a 90% return gate permits systematic missingness that could move the primary | ACCEPTED — A7 |
| 11 | MAJOR | P1b cannot show minimal revision caused S150's null without the contemporaneous crossing |
ACCEPTED — A1 |
| 12 | MAJOR | P2's "affordance" reading is confounded by forced or unfaithful markers |
ACCEPTED IN PART — A6: redefined as prompt-following reach |
| 13 | MAJOR | a non-empty loc string is counted without checking the named marker is in the text |
ACCEPTED — A6 |
| 14 | MAJOR | neither judge is calibrated for standard-spelled country/class/region/period markers | ACCEPTED — A5 |
| 15 | MINOR | P2/G5 are one quantity in two roles, so P2 corroborates nothing |
ACCEPTED — A5: G5 deleted |
| 16 | MINOR | the cost ceiling does not show the base-plus-retry arithmetic | ACCEPTED — A11 |
The amendments, in force
A1 (findings 1, 2, 11) — the design becomes a 2 × 2 and the primary moves inside it.
A contemporaneously generated source-first −I arm, Nsrc, is added. Nsrc and Asrc are
dispatched in the same stage, to the same models, at the same temperature, over the same items,
in the same format, and their prompts differ in exactly one clause — machine-checked: the
unified diff of the two prompt strings is two lines, and both are the clause.
−I |
+I |
|
|---|---|---|
| minimal revision | ∅ (frozen) |
A (frozen) |
| source-first | Nsrc (new) |
Asrc (new) |
P1 becomes loc(Asrc) − loc(Nsrc) ≥ +0.20 on both judges — the permission effect within
the source-first procedure. P1b = loc(A) − loc(∅) is the same effect within the revision
procedure. The interaction — whether the permission buys more from the source than from a
revision — is the reading, and no conclusion about "S150's null was procedural" may be drawn from
anything less than the full crossing.
A2 (finding 3) — an add-excluded sensitivity, not a compliance gate. P1 is recomputed
over the ⟨site, hand⟩ cells at which neither judge returned a non-empty add for either arm of
the comparison, and both figures are reported; if they differ in sign the exclusion governs and the
result says so. The remedy asking for compliance to gate the run is overruled, for the reason
ARM-low-pole recorded twice and note (bkr) states: at register-marked sites a content-parity
checker flags the published hand at 0.68, so a parity gate cannot license a claim about these
sites, and gating on add would repeat a failure this project has already paid for.
A3 (finding 4) — G2 is demoted from gate to sanity check, and F2 is deleted. The critic
is right that byte divergence between two independently generated arms is guaranteed and shows
nothing. G2 is retained only to record the divergence rates, including the frozen A's 0.07
against ∅, which is the number this run exists to explain. There is no separate manipulation
gate for P1, and this is stated as a limit rather than papered over: the located-idiom
manipulation cannot be verified independently of the outcome without a second rating stage the
budget does not hold. The mechanical corroborator is P4, promoted from descriptive to a
registered jury-free measure. The remedy asking for a blinded item-level coding stage is
overruled on cost — it is a third panel stage, and P4 plus A6's grounding gives a
text-grounded reading of the same thing for nothing.
A4 (findings 5, 12) — a real positive control, at no generation cost. Two arms are added:
NLEAD— the lead'sT-botchan-R22-v1English at each site:−I, standard-spelled, written under a rule that forbade located means, frozen at S150 before this question existed.PLEAD— the lead'sT-botchan-R24-v1English at each site:+I, standard-spelled, with a frozen per-site log naming the located item taken at 45 unprimed sites.
This is the "prevalidated positive-control set containing known standard-spelled placeable markers
embedded in otherwise comparable translations, rated blind in the same batches" the critic asked
for, and it exists already. The alignment from paragraph to site is the lead's, made once, and
code.assert_lead_sites proves every span is a contiguous substring of the frozen translation
— the lead can have chosen the boundaries and cannot have changed a word. Three of the thirty
spans are byte-identical across the two arms (JA-11, JA-14, JA-27) and that is reported.
These two arms are instrument controls. They enter G3b and nothing else; no primary, no
comparison with any hand, no claim about the lead's translation.
A5 (findings 6, 14, 15) — the instrument gate is replaced. G5 is deleted as a gate.
The new gate is
G3b—loc(PLEAD) − loc(NLEAD) ≥ +0.20on both judges. Can this instrument see a standard-spelled located idiom that is documented to be there?
G3 (loc(B) ≥ 0.40) survives as a batch anchor and reproduction — it says the batch is the
one S145 and S150 measured in — and no longer pretends to validate the instrument for P1's
question. F1 now fires on G3 or G3b: G3 failing means the batch is wrong; G3b
failing withholds P1's null reading and not its positive one, which is F4 relocated.
A6 (findings 12, 13) — loc must be text-grounded, and P2 is narrowed. A non-empty loc
answer counts as a placement only if the marker the judge names is present in the English it was
shown: code.grounded(), whole-word, with a frozen metalanguage stop-list and a frozen
inflection map, self-tested. Every rate is reported twice — all non-empty calls, and grounded
calls only — and the grounded figure governs any disagreement in sign. P2 is restated as
prompt-following reach and is no longer offered as a measurement of what the material affords;
the affordance question is answered, as far as this run answers it, by G3b and by the lead's
W5+ NONE rate, juxtaposed and not pooled.
A7 (findings 9, 10) — the denominator is fixed before the data. The primary is computed on
the complete-case site set: sites at which every arm returned from every hand and both judges
rated every one. G4 becomes 100% of the primary-arm cells (Nsrc, Asrc, ∅, A × 2 hands
× 2 judges) after at most one re-dispatch; below that P1 is withheld. Any exclusion is reported
as a failure and, if it changes a number, as a sensitivity — never as the primary. F3's
recompute-on-the-surviving-hand rule is deleted.
A8 (finding 8) — the site is the analysis unit. The supporting inference becomes a two-sided
exact site-level sign-flip permutation over the 30 site-level differences (Asrc minus Nsrc
placements, summed over hands), per judge, α = 0.05. The "both judges must pass" rule is a
conjunctive requirement, which is conservative; it is not and is not offered as a multiplicity
correction, and the result says so.
A9 — four rating blocks. 420 items over 4 site-blocks × 2 judges = 8 calls, ≤ 112 items per
call, item order shuffled across the whole block with seed 8155.
A10 (finding 7) — accepted as a limit; the remedy overruled. One rendering per call would be
420 calls and would break comparability with S145 and S150, and RS-20260806g measured loc/REG
as batch-sensitive by a factor of three, which is why the arms of a site are kept together.
Mitigations already in force: arm labels are never shown, item order is shuffled across the whole
block so a site's arms are not adjacent, and the source is reprinted with every item. The
strongest thing that can be said for it is empirical: this exact batching produced loc(A) −
loc(∅) = −0.0167 and 0.0000 at S150 — if seeing competing versions inflated differences, it
inflated nothing there. It is recorded as the run's largest instrument limitation.
A11 (finding 16) — the ceiling, with arithmetic. Base: 1 critic (spent, $0.02704075) + 6
generation calls at cap 6,000 + 8 rating calls at cap 16,000. Worst case from max_tokens:
$0.03 + $0.30 + $0.56 = $0.89. Maximum permitted retry path: one full re-dispatch of
generation and rating = +$0.86. Ceiling $1.75, against headroom $2.736 after the critic.
It fits with $0.99 to spare.
What was overruled, gathered
- Finding 3's compliance gate — note (bkr); an
add-based gate cannot license a claim about register-marked sites. Replaced byA2's sensitivity. - Finding 4's blinded item-level coding stage — a third panel stage the budget does not hold;
replaced by
A3's registeredP4andA6's grounding. - Finding 5's bespoke positive-control set — replaced by
A4, which is better: real translations by a real translator with a frozen per-site record of what was put in them. - Finding 7's one-rendering-per-call rating — 420 calls, and it would break the batching rule
RS-20260806gestablished. Recorded as a limit instead (A10). - Finding 6's "acceptable sensitivity and specificity before interpreting
P1" —G3bis a sensitivity threshold; a specificity threshold onNLEADalone would need an absolute bar this project has no basis for setting, andG3b's difference form does the work without one.