Repository path: workshop/experiments/E-20260805e-naturalness-wording/critic.md · rendered 2026-09-09
Page metadata (front matter)
| type | note |
|---|---|
| id | E-20260805e-critic |
| status | frozen |
| created | 2026-08-05 |
| updated | 2026-08-05 |
| links | workshop/experiments/E-20260805e-naturalness-wording/design.md, workshop/experiments/README.md, config/budget.md |
Pre-run adversarial critic pass — dispositions
Seat: x-ai/grok-4.5 (P3) — a panel model that is not a juror in this run (the jurors are
P1, P2, P5), which is what the S053 role-collision fix requires. One pass, max_tokens 24,000,
$0.0456044, 8,101 prompt tokens. Full text: critic.txt; raw body: critic.raw.json.
Verdict: NEEDS-REDESIGN. Sixteen findings, eight BLOCKING. All sixteen accepted — fourteen in
full, two in part with the refused remedy reasoned. Nothing had been dispatched against the design
when the amendments were made, and the amended version is design.md v2; a run may proceed
only against v2.
Two dead bodies before this one, and they are on the ledger. The seat was z-ai/glm-5.2;
attempt 1 (max_tokens 12,000) returned finish_reason: length with 11,595 reasoning tokens and
empty content, $0.03519676; attempt 2 (32,000) returned the same with 32,436 reasoning tokens,
$0.07863676. Both preserved as critic.attempt1.json / critic.attempt2.json. Raising the cap a
third time is not a fix — the seat spent its entire allowance on hidden reasoning at both caps —
so the seat was changed. Note (bhf)/(bit), ninth firing, seventh distinct slug.
| # | finding | BLOCKING | disposition |
|---|---|---|---|
| 1 | The REVISED string makes three edits at once, so the run discharges the obligation only as a bundle test; §6.1 and P4p nonetheless attributed outcomes to "the struck clause" | yes | ACCEPTED IN FULL. Every licensed sentence in v2 attributes only to the §3.2 REVISED string as a whole (three inseparable edits). §0.5 states the bundle limit; §10 records the single-edit design a clause-specific discharge would need |
| 2 | Naming a register anchor is a second treatment that can cancel or mimic the clause strike | yes | ACCEPTED IN PART; both offered remedies refused with reason. (a) three arms does not fit the day's headroom; (b) putting the anchor into both arms destroys the byte-identical retest, which findings 4 and 13 make more load-bearing, not less. Adopted instead: the anchor is named as the third inseparable edit everywhere (finding 1's remedy applied here), it is assigned blind by a third party rather than by the lead (finding 3), and the wording effect is reported split by anchor group (finding 10) so an anchor-driven effect is visible rather than hidden |
| 3 | The provenance→anchor rule is a date rule, not a register rule, and is confounded with cell | yes | ACCEPTED IN FULL, by the critic's own preferred remedy. §3.3 is replaced by stage 0: a non-juror seat reads the six undamaged references, unlabelled, and assigns each one of the three anchors verbatim from wiki/goodness-senses.md, knowing nothing of the design. Whatever it returns is used. P3p is restated as conditional on the new cell's assigned anchors |
| 4 | FC3 used the wrong noise comparator — distance from a single published point estimate is a random throttle, not a stability gate | yes | ACCEPTED IN FULL, by the critic's "better" option. Split into FC3a, a reported retest statistic |OLD−1.12|, and FC3b, a gate defined entirely within-run: the permutation test fires and |effect| ≥ δ_min = 0.25 |
| 5 | The sign test can fire only on near-unanimous non-ties and its threshold package is incoherent | yes | ACCEPTED IN FULL. P2p is now an exact permutation test on the 12 paired δ (all 2¹² = 4,096 sign assignments, magnitudes kept, zeros handled by construction). The sign test is retained as a secondary statistic, stated correctly as one-sided with ties dropped |
| 6 | The result→claim map leaves outcomes unmapped and rows 1 and 4 over-reach | yes | ACCEPTED IN FULL. §6.1 is a full decision table gated on FC1, FC2, FC4 and this run's own OLD value; row 4 is softened to a noise-bounded non-claim; the P1p∩P4p cell has its own row; the process rule about future designs is moved out of the map to §10 |
| 7 | Overlap with a published English can still move numbers if an edit site sits inside or beside the shared run | no (BLOCKING if unchecked) | ACCEPTED, AND THE CHECK IS DONE. The 19-token run occupies characters 1332–1424 of N8-A's reference. No edit site intersects it; the nearest (site 7) ends 83 characters before it. Recorded in §3.4 with the offsets. N8-A and N8-B are reported separately throughout |
| 8 | FC1 makes a drop uninterpretable, but §6.1 assigned verdicts without gating on it | yes | ACCEPTED IN FULL. Every row of §6.1 is gated on FC1 firing in both carried cells; otherwise only uninterpretable under FC1 may be written |
| 9 | FC2's 0.75 is borrowed from the naturalness specificity bar and is not argued as a spillover bound | no | ACCEPTED IN FULL. FC2 is now proportional to this run's own OLD accuracy drop: |Δ drop(accuracy)| ≤ 0.25 × drop(accuracy)ᴼᴸᴰ. The absolute value is reported too |
| 10 | A 2–2 anchor balance still mixes anchor with item and translator inside the pooled δ | no | ACCEPTED IN FULL. A pre-registered exploratory split by assigned anchor group is added; the primary stays pooled and §6.1 may not claim homogeneity |
| 11 | The seat probe consumes stage-1a data under a stop rule that can destroy the retest, and does not bound the longest payload | no | ACCEPTED, ADAPTED. The probe is now the longest payload in the whole run (N8-B REVISED, order 0), dispatched first and counted as data; and a partial-stage rule is added: the obligation cell completes only if all 16 carried payloads score, else no §6.1 sentence may be written |
| 12 | P3p over-reads a 6-unit mean | no | ACCEPTED IN FULL. P3p is descriptive by default; the words reproduces and extends beyond Russian require a pre-registered rule (FC1 fires in both new cells and ≥ 5 of 6 non-tied δ positive) |
| 13 | Temperature 0.2 drift compounds FC3 | yes (subsumed under 4) | ACCEPTED, discharged by finding 4's remedy: resolvability is now within-run only |
| 14 | FC1's detection rule is overall preference, not naturalness-targeted or accuracy-targeted detection — a construct mismatch inherited from the prior run | no | ACCEPTED IN FULL. §5 states the mismatch explicitly as legacy continuity and adds a reported check: drop(accuracy) > 0 in the same cell, with a flag if preference fires and the accuracy drop does not |
| 15 | Net word-count inflation is a naturalness cue confounded with accuracy dose | no | ACCEPTED IN FULL. §9 and P4p's licensed language now say the inseparability claim is about O4 as implemented, including its length-changing sites, not about pure semantic error with no surface disruption |
| 16 | §8's "byte-identical except for nothing at all" should be three separate assertions | no | ACCEPTED IN FULL — and already implemented: analysis/verify.py §§1–3 assert (i) carried strings by SHA-256, (ii) OLD payloads against the S086 builder, (iii) REVISED differing in exactly one line, that line being - naturalness:, and the assigned anchor appearing in it |
What the pass cost and what it bought. $0.0456044 on the accepted body, $0.1138 on two dead ones before it — 26% of the session's spend, and it changed the run's inferential structure before a single juror was paid. Findings 4 and 5 removed a resolvability gate that could have licensed a wording claim at random; finding 6 removed two sentences the run would otherwise have been entitled to write; finding 3 removed the lead from a judgment the design had it making.