Repository path: workshop/experiments/E-20260817e-mimetic-reading/critic-response.md · rendered 2026-09-09
Page metadata (front matter)
| type | note |
|---|---|
| id | E-20260817e-critic-response |
| status | active |
| created | 2026-08-17 |
| updated | 2026-08-17 |
| internal-judgment-only | true |
| links | workshop/experiments/E-20260817e-mimetic-reading/design.md, wiki/arms/ARM-mimetic-carriage.md |
Answer to the pre-run adversarial critic, E-20260817e
Seat P1 openai/gpt-5.6-terra, one call, $0.049393. Verdict NEEDS-REDESIGN, 18 findings, 12
BLOCKING. v1 was never dispatched. Full body: critic-v1.json.
The body was truncated by its own cap (finish_reason: length at 6,000 completion tokens, cut
mid-sentence after finding 18). Findings 19 onward were never seen. That is a defect in the v1
pass and it is why v2 is sent back for a second pass at a larger cap rather than dispatched on the
strength of an incomplete critique.
The findings, and what was done
| # | B? | finding, compressed | disposition |
|---|---|---|---|
| 1 | B | "no more and no less" demands exact cross-linguistic identity; NEITHER becomes the mechanically safest answer and the run goes vacuous |
taken, and it is the redesign. The forced choice is gone. Each arm is judged on its own on two orthogonal binaries, ADDS and OMITS; SAYS is derived. No answer is safest and both arms may get identical answers |
| 2 | B | agreement between seats could be agreement that neither English is exact, or agreement driven by English naturalness — not recoverability of a source reading | taken. The primary is no longer per-site agreement; it is a rate contrast between real mimetic sites and null sites. If naturalness drove the labels the two rates would move together, which the design now says in §7.3 |
| 3 | B | the null items test "will readers prefer vivid English", not "will readers hallucinate a mimetic assertion"; a correct reader may pick the enacting arm for source-sensitive reasons, so G1 can fire against a correct design | taken. The nulls are demoted from a pass/fail gate to the comparison condition. Under ADDS/SAYS, plausible but not stated is ADDS — which is the answer the critic says a good reader would give, and the design now wants that answer rather than punishing it |
| 4 | B | C1's stating arm "made my voice as big as I could" is a calque; a seat choosing bellowed may be choosing better English |
taken. Stating arm rewritten to spoke as loudly as I could. Flagged pl_repaired in items.json and declared in §2 as constructed material, not a quotation of the frozen chapter |
| 5 | B | C2 — scuttled out is plausibly licensed by 急いで引き揚げた, and beat a quick retreat adds escape implications |
partly taken. Kept, because under v2 "plausibly licensed but not stated" is exactly the ADDS label the baseline is made of. The uneven idiom is recorded as a limit |
| 6 | B | C3 — shooed him off is an ordinary rendering of 追っ払う and asserts nothing false |
same disposition as 5, and the point is now a feature: if seats label it SAYS, the baseline rises and the contrast shrinks, which is the honest direction |
| 7 | B | C4 — splash round and round is a plausible construal of exuberant swimming |
same disposition as 5 |
| 8 | NB | C5 is the strongest but came down smartly is itself marked |
taken. Rewritten to came briskly down, the critic's own suggestion. Flagged pl_repaired |
| 9 | B | G1's threshold turns a handful of disputed sentences into an all-or-nothing verdict on everything | taken in full — see 3. There is no longer a null-site gate. The gate that remains, G1 in v2, is inter-seat agreement, which is a property of the instrument and not of five contested sentences |
| 10 | B | identity items test string-identity recognition and instruction-following, not source reading; an A/B answer may be a format error, not confabulation | taken. G2 is demoted to a symmetry-and-attention diagnostic, says so in the design, and withholds the primary only if two seats fail; one failing seat gets the primary reported twice, with and without it |
| 11 | NB | on identity items both BOTH and NEITHER are defensible; the design must not imply BOTH is expected |
taken and obsoleted. Under v2 there is no BOTH/NEITHER: the check is that a seat's A-flags equal its B-flags when the strings are equal, whatever those flags are |
| 12 | B | M12 and M13 share a Japanese sentence and each arm changes both spans |
OVERRULED, with the frames quoted. M12's frame is "He wore a {X} gauze haori and snapped his fan open and shut, saying…" and M13's is "He wore a flimsy, flappy gauze haori and {X}, saying…". Each frame holds the other span fixed; only the nominated slot varies. The critic misread the materials. The residual is real and is declared: the held-fixed span sits in its enacting form in both arms of each item, priming the sentence toward enactment (§7.5) |
| 13 | B | M05 — went in with a splash is idiomatic, plunged in noisily is not; a preference may be an English-quality judgment |
acknowledged, not repaired. These are frozen chapter-2 materials and this project does not silently edit a frozen artifact. Under v2 the damage is much smaller, because each arm is labelled against the Japanese rather than against the other. Carried into the result page's limits with the critic's wording |
| 14 | B | M08 — grinning and grinning vs grinning … the whole time tests aspect, not enactment |
same disposition as 13. It is also one of the three sites RS-20260816e §3 already reported as failing propositional parity, which is the premise of this design rather than news |
| 15 | B | M09 — the enacting arm supplies two sound characterisations and the sentence then repeats the noise |
same disposition as 13 |
| 16 | NB | M01 — gave a hoot on its whistle may be too specific and is awkward |
acknowledged. Under v2 "too specific" is precisely what ADDS records, so this item is now informative rather than broken |
| 17 | NB | M03 — walked off slowly and heavily is awkward and may not equal のそのそ |
same disposition as 16 |
| 18 | NB | M04 — neither arm is a neutral statement of the other's content |
same disposition as 16. M04 is one of the three contested-class sites and is already excluded from every by-class split |
Twelve BLOCKING findings, nine taken, one overruled in writing with the evidence quoted, and two (5–7 group, 13–15 group) taken as declared limits on frozen material that this project does not edit after the fact.
What changed in the materials
prompts.pyREADINGreplaced: forced four-way choice → four independent binary flags.NULL_SITESC1andC5stating arms rewritten;C6,C7,C8added, raising the baseline from 5 items to 8.run.pyparser rewritten for the four flags; item count 23 → 26, reading calls 69 → 78; ceiling $0.80 → $0.95, stop-loss $0.70 → $0.82, worst case $0.62 → $0.72.design.md§1, §4, §5, §7, §8, §9 rewritten; §7.4 and §7.5 are new and exist only because the critic wrote findings 12–18.
The v2 pass — NEEDS-REDESIGN again, 16 findings, 12 BLOCKING, $0.06511625
Seat P1, one call, cap 14,000, finish_reason: stop — this one was not truncated. Full body:
critic-v2.json. It was shown the v1 findings and this response to them, and told to say where a
disposition was wrong rather than repeat what had been taken.
Two of its findings are decisive and both are correct.
- BLOCKING 1 — the 0.20 contrast threshold sits inside one standard error. At n = 15 and n = 8, SE(difference) = √(.25/15 + .25/8) ≈ 0.219; ~98 items per condition would be needed for 80% power at that effect; the rate increments are 1/15 = .067 and 1/8 = .125. A trivially noisy difference could clear the bar, and the design turned clearing it into a handbook conclusion.
- BLOCKING 2 — the null set was not a null set. "Contains no mimetic word" is not "states no
sound or manner." Seven of eight null sentences state a manner outright:
なるべく大きな声をして(C1),急いで(C2),追っ払う(C3),泳ぎ巡って(C4),威勢よく(C5),遠慮なく(C7),引き下がって(C8); and C6 (scrubbed vs rubbed) is a difference in the method of erasure, not a phonaesthemic move at all.
| # | B? | finding, compressed | disposition |
|---|---|---|---|
| 1 | B | the 0.20 threshold is below the sampling noise and is turned into a categorical conclusion | taken. No threshold survives. The primary is an exact two-sided sign test on paired items, registered with its p-values (9–0 → 0.004, 8–1 → 0.039, 7–2 → 0.180) |
| 2 | B | the null set is not null; eight items enumerated | taken, and it produced the redesign. The between-item null is gone. The control is now within item: the same English arms against the same Japanese sentence with the mimetic deleted |
| 3 | B | ADDS presupposes a stable cross-linguistic partial ordering the prompt does not supply; the examples invite finding additions in the marked arm |
taken as far as it can be, and neutralised by design. The example list is removed and the English target is marked. The residue is not cured but is made constant across the two conditions, so it cannot produce the contrast. A pre-annotated proposition inventory was refused: it would be lead-authored and would move the free choice rather than remove it |
| 4 | B | contradiction hard-coded as ADDS and OMITS conflates substitution with addition | taken. The rule is gone; ADDS now reads "something the Japanese does not tell its reader there, or something the Japanese denies" |
| 5 | B | raw agreement has no 0.50 chance level under skewed marginals; pooling four flags hides instability in the one that matters | taken. The pooled agreement gate is gone. G1 is now a floor/ceiling check on the primary flag; per-seat marginals and flag-specific agreement are reported alongside every figure |
| 6 | B | missing-data and tie rules unregistered | taken. Registered in §6: dead cell → two seats; 1–1 → cell unresolved; unresolved → pair dropped under G3; every drop named; unresolved cells never enter a denominator |
| 7 | B | "style choice and not accuracy" from a failure to clear a threshold is absence-of-evidence | taken. §9 now says a null closes the arm with "not shown to be recoverable by this design at this n", and explicitly refuses the stronger claim |
| 8 | B | nine real items still confounded (M05 M08 M09 M12 M13 N01 N07 M01 M04), and separate-arm judging does not cure it |
answered by construction. Every one of those defects is identical in both conditions of its pair and cannot produce a difference. M09, M13, N01 are additionally absent from the primary because their deletions are not grammatical. The list is carried verbatim into the result page's limits |
| 9 | B | M12/M13 are one clustered source sentence and are primed by the fixed enacting span |
taken in effect. M13 has no clean deletion and is absent from the primary; only M12 is paired, so the cluster is n = 1. Declared as an accident, not a repair |
| 10 | B | the reading prompt is leading; the English target is unmarked | taken. Examples removed, English target marked <<like this>>, a reason required per arm — and see 3 for why the residue no longer matters |
| 11 | B | the buildability prompt is a tutorial in the project's theory; "can only weaken" is not secured | taken in full. BUILD rewritten blind: it names no lexical class, no theory, no example, and asks only for up to three alternatives at least as accurate. §8 registers that a NONE has no evidentiary force and that classifying a returned candidate as phonaesthemic is a lead judgment |
| 12 | B | the dispatcher enforces neither the stage order nor the stop-loss nor the one-re-dispatch rule nor the response format | taken, all six. DEPENDS blocks a stage whose dependencies are incomplete; a dead row is written after one re-dispatch and is never re-bought; the hard ceiling is checked after each call; the parser requires all six lines, rejects duplicated flags, and requires both reasons |
| 13 | NB | G2 is a symmetry check, not an attention measure; its cutoff is arbitrary | taken in wording. Called a symmetry check throughout. The cutoff is kept and is listed in §11 as a free choice |
| 14 | NB | three model seats are not a reader population | taken. §4 says so, and "the source licenses" is barred from the result without the panel qualifier |
| 15 | NB | many choices called mechanical or frozen are substantively free | taken. §11 is a new decision log separating investigator choices from inherited constraints |
| 16 | NB | prediction 4 is registered as untestable and should not be a prediction | taken. Struck from the list |
Sixteen findings, twelve BLOCKING; all twelve taken, none overruled. The single overrule in this
experiment's whole critic history is v1's finding 12, on M12/M13's frames, and the v2 critic's
finding 9 partly reinstated it — which is why M13 is out of the primary.
No third pass was bought. Two passes have cost $0.115 and killed two designs; the second one's central objection built the design that runs, which is the strongest evidence available that the mechanism worked. A third would be buying reassurance, and the honest thing is to record that v3 has not been adversarially reviewed and that its own defects are therefore the lead's to have missed. That is stated on the result page, not only here.