Repository path: workshop/experiments/E-20260822b-echo-threshold/critic-response.md · rendered 2026-09-09
Page metadata (front matter)
| type | note |
|---|---|
| id | E-20260822b-critic-response |
| status | frozen |
| created | 2026-08-22 |
| updated | 2026-08-22 |
| links | workshop/experiments/E-20260822b-echo-threshold/design.md, workshop/experiments/E-20260822b-echo-threshold/critic-v1.json, workshop/experiments/E-20260822b-echo-threshold/critic-v2.json |
E-20260822b — the pre-run critic, and what was done about it
Round 1: P1 openai/gpt-5.6-terra, $0.052856, verdict NEEDS-REDESIGN, 5 findings — 2 BLOCKING,
3 MAJOR. Bought before any data call. The critic saw the frozen design, the frozen prompt strings,
the whole 69-item manifest and the mechanical rule that assigns the levels.
Every finding is accepted. Two of them changed the design's shape and one of them changed the
primary statistic. The design as dispatched is the amended one; the round-1 text stands in
critic-v1.json.
BLOCKING 1 — "measures models, states a conclusion about readers"
Model seats can apply learned rhyme conventions, follow the explicit request to search for "sound echo," or reproduce training-pattern judgments without undergoing human auditory or literary reading. … a positive or null result cannot overturn a claim about readers or support a handbook instruction about reader reception.
Accepted. The design's §2 already said the seats are models; it then went on to write predictions in the word reader, which is exactly the slippage the critic names.
Amendment A1. The word reader is struck from every prediction, criterion and consequence and
replaced by reading seat. §2's estimand is restated. And the handbook consequence is bounded in
advance: if F1 fires, framework/v0.2 §7 gains the instruction with its warrant named on the
face of it — three model seats, no human reader, one work, one language pair — and the existing
sentence in RS-20260822 §4 is corrected in the same edit whatever the outcome, because that
sentence asserted a fact about human hearing on no evidence at all and cannot stand either way.
The critic's alternative remedy — "replace the seats with readers" — is not available, is on the project's named-not-built list, and is stated as the arm's binding limit rather than pretended away.
BLOCKING 2 — level is confounded with the lexical identity of the filler
Each level is a different word, chosen non-randomly and often differing simultaneously in meaning, collocation, register, word frequency, morphology, syntactic fit, semantic parallelism, and poetic expectancy. The naturalness ranking can detect only one broad direction of one confound.
Accepted, and it is the finding that reshaped the run. The original design compared
(fixed, X) against (fixed, Y), so every difference between the words X and Y rode along with
the phonetic relation. A naturalness ranking cannot repair that, and the design leaned on it as
though it could.
Amendment A2 — the design becomes a 2 × 3 crossover. Each locus now carries two words in the
fixed slot: a, the original, and b, a second ordinary word for the same slot that
tools/rhyme_pairs.py scores NONE against every filler at that locus. Crossed with the three
fillers this gives six cells (four where the locus has no Z):
| filler X | filler Y | filler Z | |
|---|---|---|---|
| fixed a | NEAR | NONE | STRICT |
| fixed b | NONE | NONE | NONE |
The primary is now the interaction, not a difference of means:
Δ_N = [hit(aX) − hit(aY)] − [hit(bX) − hit(bY)]
Filler X appears in one NEAR cell and one NONE cell; filler Y appears in two NONE cells; fixed words a and b each appear in both columns. Everything the critic lists — meaning, collocation, register, frequency, morphology, syntactic fit, semantic parallelism — is a main effect of the word, and a main effect of either factor cancels in the interaction. What does not cancel is the phonetic relation, which is present in exactly one of the four cells.
The manipulation check M1 is rebuilt the same way on Z. The item count goes from 69 to 138.
What A2 does not fix, and is disclosed: the interaction still assumes the X-versus-Y difference is the same under fixed word a as under fixed word b. That is an additivity assumption, not a randomisation. It is much weaker than what the original design assumed, and it is named in the limits rather than buried.
MAJOR 3 — the permutation p-value is a relabelling distribution, not a randomisation test
Permuting F, N, and O labels asserts that the observed fillers could have occupied one another's levels under the null, although their level is mechanically and purposively determined.
Accepted. The levels are not assigned at random and no permutation of them yields a
randomisation p-value; calling the number p invited exactly the reading the critic objects to.
Amendment A3. The primary uncertainty statement becomes a bootstrap over loci (the unit that is plausibly exchangeable, since the loci are a convenience sample of one work's parallel cola): 10,000 resamples, seed 20260822, reported as a percentile interval on Δ. The permutation number is still computed and is reported under its true name — a within-locus label-reshuffle reference distribution — with one sentence saying what it is not. No claim rests on it.
The critic's stronger remedy ("report descriptive differences without permutation p-values") is followed in substance: the criteria in §6 are now stated on effect size and interval, not on a p-value.
MAJOR 4 — no unprompted condition, so "upper bound" is unsupported
A null under that task does not establish that ordinary readers would fail to hear an echo unless prompted, and a positive result cannot be quantified as an upper bound without observing the unprompted response rate.
Accepted, and bought rather than argued. The word upper bound was doing real work in §2 and was an assumption dressed as a measurement.
Amendment A4 — stage U is added: the three fixed-a cells of the 19 three-level loci
(aZ, aX, aY), 57 items, dispatched to P1 under an open prompt that does not mention sound
at all — "Is there anything you notice about how this passage is written? Answer in one sentence."
— and coded mechanically for whether the planted pair is named. This gives a real unprompted rate at
STRICT, at NEAR and at the floor. If a full rhyme is not spontaneously remarked on, the
direction argument is measured rather than assumed; if it is, the prompted figures are the ones
that need the caveat.
MAJOR 5 — the control passages are not free of other phrase-end relations
A filler can alter competing echoes, and existing echoes can occupy the response slots … making N-minus-O reflect competition among unscored relations rather than detectability of the planted relation.
Accepted, and screened mechanically at no cost.
Amendment A5. Every one of the 138 passages is segmented at line breaks and at ; : , . ? !, the
last word of each segment taken, and every pair of those words scored by tools/rhyme_pairs.py.
The result is in the manifest as unplanted_echoes and is reported:
14 of 138 items, spanning 5 of 25 loci (
PR6,PN3,VE1,VE8,VE11), carry an unplanted phrase-end relation. Every one of them isNEAR; not one isSTRICT.
A registered sensitivity analysis recomputes Δ_N, Δ_F and every criterion with those five loci
dropped, and both figures are reported side by side. The critic's stricter remedy — rebuild the
carriers until no unplanted relation exists anywhere — is refused in writing: it would require
discarding PN4-class items whose whole value is that they are the pairs RS-20260822 §4 named, and
a screen plus a pre-registered exclusion answers the mechanism without letting the item set be
reshaped by what the authors want to keep. The scoring rule's other count measures the same risk
empirically from the seats' own answers.
Round 2
P1, $0.0784515, verdict NEEDS-REDESIGN again, 5 findings — 2 BLOCKING, 3 MAJOR. Bought under
note (bqp)'s second clause: A2 changed the primary statistic and A4 added a stage round 1 never saw,
so round 1's verdict could not stand as a review of what would be dispatched. Buying it was
right — round 2 found the confound that would have made the whole run about spelling, and it caught
a factual error in a refusal.
BLOCKING 1 (round 2) — orthography, not phonology
Seats receive text, not sound … a successful Δ_F or Δ_N does not establish that the models recovered a sound relation rather than spelling-, suffix-, or word-form similarity.
Accepted, and it is the most important finding either round produced. Measuring the item set as
built shows the confound is real and large: mean shared final-letter run between the two words is
2.63 at STRICT, 1.60 at NEAR, 0.13 at NONE. The independent variable and its orthographic
shadow are almost collinear.
Amendment A6 — stage P, the orthography/phonology probe. Twelve dedicated items on neutral carriers, the same detection prompt, 3 seats, 36 calls:
- 6
PHONitems — full rhymes spelled unlike: blade / weighed, near / austere, conceit / complete, just / nonplussed, abide / denied, dies / sighs (shared final-letter run 0, except dies / sighs at 1). - 6
ORTHitems — rule-NONEpairs spelled alike: rough / though, cough / dough, worm / storm, boughs / coughs, sword / word, beard / heard (shared run 3–5).
If the seats are matching spelling, ORTH gets the hits and PHON does not; if they are
recovering pronunciation, the reverse. Registered before dispatch: if hit(ORTH) > hit(PHON),
every figure in the run is disclosed as measuring orthographic resemblance and the phonetic reading
of M1, P1 and F1 is withdrawn.
Amendment A7, free: the shared final-letter run between the two words is recorded on all 138
items and reported as a covariate, and hit within the 94 NONE cells is reported as a function of
it. Building the probe also produced a fact worth having on its own: of 30 classic English
orthographic traps, the mechanical rule scores 25 as NEAR — the NEAR class and the
spelled-alike class very largely coincide, which is why the probe was buildable at all only from a
narrow residue.
BLOCKING 2 (round 2) — the interaction still carries the lexical pairing
the cell
aXuniquely contains the semantic, collocational, syntactic, register, expectancy, and contrastive interaction of those two particular words.
Accepted as a limit; the remedy is refused in writing. The critic asks for "comparable lexical-pair interactions … in both NEAR and NONE cells". No such stimulus exists: a word pair cannot be simultaneously a rhyme and not a rhyme, so in any design whatever the sound relation is perfectly confounded with being that particular pairing. This is a property of the question, not a defect this design chose.
Amendment A8 — what is done instead. The per-locus Δ is reported for all 25 loci, with the count of loci whose Δ is positive and a sign test across them. One lexical pairing can carry a spurious effect; twenty lexical pairings sharing nothing but the sound relation cannot, without twenty coincidences. The estimand in §2 is restated to say exactly this, and the limits carry it.
MAJOR 3 (round 2) — the flagged loci must leave the primary, and one reason given was wrong
retaining PN4 is not a reason to retain them because PN4 is not among the flagged loci.
Accepted, and the correction is the critic's. The round-1 refusal argued that excluding flagged
loci would cost the PN4-class items. PN4 is not flagged; the flagged five are PR6, PN3,
VE1, VE8, VE11, and all three of the printed-pair loci survive exclusion. The stated ground was
simply false.
Amendment A9. The primary is now computed on the 20 unflagged loci, and the 25-locus figure is reported beside it. The critic's second point — that the screen defines phrase-ends by punctuation while the prompt leaves phrase undefined — is accepted as a declared limit: the screen can only bound the risk it can see.
MAJOR 4 (round 2) — the bootstrap does not license a class-level claim
Accepted. The 25 loci are constructed, not drawn from a population of English phrase-ends.
Amendment A10: Δ and its interval are reported as descriptive of these constructed loci, and
P1/F1's conclusions are stated about them, with any wider reading marked in the text as an
inference and not a measurement. The handbook sentence, if F1 fires, carries "on 20 constructed
loci from one work, three model seats, no human reader" on its face.
MAJOR 5 (round 2) — P3 cannot say "manufacture"
Accepted, and it costs nothing. Amendment A11: every pair a seat names is rescored by
tools/rhyme_pairs.py, and yes_any on NONE cells is split into (i) the named pair is STRICT
or NEAR by the rule — the seat found an unplanted relation the rule agrees with — and (ii) the
named pair is NONE by the rule — only this supports the word manufacture, and only it is
reported under that word.
Why there is no round 3
Round 2's amendments add a stage and change which loci the primary runs on; they do not change the primary statistic, and A6 is a diagnostic that can only disclose, never inflate. Two rounds is already one more than this project normally buys, and a third would begin to price the design rather than test it. The dispositions above are the record.