Repository path: workshop/experiments/E-20260830b-rhyme-family/critic-response.md · rendered 2026-09-09
Page metadata (front matter)
| type | note |
|---|---|
| id | E-20260830b-critic-response |
| status | frozen |
| created | 2026-08-30 |
| updated | 2026-08-30 |
| links | workshop/experiments/E-20260830b-rhyme-family/design.md, workshop/experiments/E-20260830b-rhyme-family/design-v2.md, runs/RS-20260830b-rhyme-family/critic_res.json |
Pre-run critic pass — E-20260830b, and what design v2 changed
Two independent seats, both on the frozen v1 design, neither shown the other's answer:
C1 = P1 openai/gpt-5.6-terra, C2 = P3 x-ai/grok-4.5. Note (brr): P2 was kept off the
long critic prompt by construction. Both returned VERDICT: NEEDS-REDESIGN, C1 with 15
findings (12 BLOCKING), C2 with 12 (6 BLOCKING) — 27 findings. Raw bodies:
runs/RS-20260830b-rhyme-family/critic_res.json. Cost $0.112091900.
The two seats converge on the same five structural objections, arrived at independently, and four of them change the run. Design v2 is the amended design and is what was dispatched.
Accepted, and what changed
A1 — the predictor and the outcome came from the same raters (C1-2, C2-1, the finding both
seats put first). Stage F and stage S used the same three seats. A seat that glosses Persian
generously, or finds the Persian of one ode more legible than another's, raises A and R
together, and a positive ρ follows with no lexical mechanism anywhere in it.
Change: the two stages now run on disjoint seat sets from disjoint labs. Stage F is P1
(OpenAI) + P3 (xAI); stage S is P2 (Google) + PR qwen/qwen3.7-max (Alibaba), the first
reserve in config/models.md, promoted for this run for exactly this reason and recorded as
provenance. No seat contributes to both sides of the correlation.
A2 — stage F let one call invent the senses and then the rhyme that covers them (C1-1,
C2-7). A seat asked in one breath to gloss a word and to find a rhyme for its gloss will bend the
gloss toward what rhymes.
Change: stage F is split into two sealed calls. F1 asks for the contextual sense of each
qāfiya word and does not mention rhyme, English verse, or translation at all. Its output is
frozen. F2 is a second call to the same seat carrying that frozen sense list and the instruction
that the senses may not be changed; only F2 mentions rhyme.
A3 — R was set cover, not monorhyme carriage (C1-3, C2-9). Stage S asked only whether some
English line-ending carried each Persian sense. A rendering that abandons the monorhyme entirely
could score high.
Change: every stage-S YES is now screened mechanically. The English word named must (i) be one
of the translation's actual rhyme-bearing words and (ii) rhyme, under the same CMUdict checker and
the same non-rhotic licence used at S233, with the modal rime of that ode's own bearer set. A YES
failing either test is recoded NO. The screened figure R_rhyme is the primary; the unscreened
R_set is reported beside it.
A4 — the primary stratum was cut on a property of the translation (C1-5, C2-5). PRIM was
|N − M| ≤ 1, and M is Leaf's own position count, which reflects his omissions.
Change: the primary is all 28 odes. The alignment subset survives only as a declared,
explicitly exploratory sensitivity analysis.
A5 — the translation limb could not falsify anything (C1-10, C1-11, C2-3). The lead wrote
the conjecture, selects the poem as the lowest-A candidate, knows the registered prediction, and
then translates. P4 could only measure the lead's compliance, and it sat inside the refutation
clause, which made the refutation steerable too.
Change: P4 is withdrawn as a registered test. The translation limb is an illustration — a
frozen enumeration of what the rhyme actually cost, entered as craft evidence and
internal-judgment-only — and it is removed from the refutation rule, which now depends on the
primary alone. Its blind stage-S score is reported descriptively and is never support for Q1.
The wire between the limbs is now generates, not tests (continue-prompt.md §4).
A6 — Leaf chose his twenty-eight (C2-2). Conditional on his having published them under a
monorhyme contract, the low-A tail may simply be absent.
Change, and it is the best thing either critic gave this design: the selection is now measured
rather than conceded. Stage F already had to run on unselected ghazals to pick the translation
limb's poem; the pool is raised from ten to twenty, and P5 asks whether Leaf's 28 sit
higher in A than twenty ghazals nobody chose. The conclusion of Q1 is narrowed in writing to
Leaf's menu whatever P5 does.
A7 — no gate could fail on the mechanism (C1-9, C2-10). Change: a placebo-predictor
gate. Four source-side or match-side quantities that are not rhyme-family availability — Persian
position count, mean qāfiya character length, the S231 match ratio, and Leaf's position ratio
M/N — are correlated with R by the same statistic. If any of them predicts R at least as
strongly as A does, the mechanism claim is withheld and the run says so.
A8 — the mismatched-ode control is weak because the ghazal's rhyme lexicon is stock (C2-4).
Pairing the Persian of one Hafez ode with the English of another still offers heart, soul, wine,
dust.
Change: two graded control arms. X-MIS (4 items) is the in-field mismatch and is now
descriptive only; X-SHUF (4 items) pairs an ode's Persian with a bearer word drawn from four
different odes, destroying both poem coherence and any single rhyme family, and carries the
gate: R_true − R_SHUF ≥ 0.20.
A9 — the qāfiya unit may not be a word (C1-7, C2-11). Change: F1 and stage S both carry
an UNGLOSSABLE option, and a position so marked by a majority of seats is dropped from the
numerator and denominator of both A and R identically.
A10 — CMUdict is modern and American; Leaf is 1898 and British (C1-6, C2-8). Change:
the checker is analyse.rime_nor, which already carries Leaf's own non-rhotic licence; a word
absent from CMUdict is counted UNSCREENABLE and every screened figure is reported twice, once
counting them NO and once dropping them.
A11 — smaller, accepted verbatim. C1-13: the permutation test is called Monte Carlo, with the
generator, seed, the +1 convention and midrank tie handling printed. C1-14 / C2-12: the control
pairs and the candidate draw order are printed in the design, and the budget fallback truncates the
candidate list by a prefix of the frozen draw order, never by A. C1-15: P3 is labelled
convergent evidence for general rhymeability and not for sense coverage. C1-4 / C2-2: the
"difficulty is in the language pair and not in the translator" sentence is struck from §3.
Overruled, with the reason
O1 — C2-2's first option: pull Payne 1901 into this run as a co-primary second hand. Overruled
on materials, not on principle. runs/RS-20260830-leaf-contract/payne_odes.json is Internet Archive
OCR of visible poor quality (skinker, tiD, Djre, wajrfiurer), and no Payne↔Ganjoor match
exists; building one badly and correlating on it would manufacture exactly the noise C2-6 warns
about. Payne is step 2 of ARM-rhyme-family and the arm says so. C2's own fallback is taken
instead: the conclusion is narrowed to Leaf's menu in §3 and §11.
O2 — C1-1 and C1-7's remedies: a pre-registered searchable English rhyme lexicon, and a
Persian prosody specialist freezing every qāfiya span. Both are correct and neither is reachable
by this project (no WordNet or comparable lexicon offline; no human specialist). The construct is
therefore renamed rather than pretended: A is what a competent reader of both languages can
find, which is the translator's actual situation, and design v2 §3 forbids the result page from
calling it a property of the English dictionary or a census of English.
O3 — C2-6's remedy: raise n to 50+ or do not run. Overruled for this step and conceded in
writing. n = 28 detects ρ ≈ 0.45 at 80% power one-sided, and design v2 registers that number, marks
the minimum effect of interest, and forbids reading a non-significant P1 as evidence against a
moderate effect. That is the honest version of an underpowered first step, and raising n is what
step 2 is for.