Repository path: workshop/experiments/E-20260826b-radif/critic-response.md · rendered 2026-09-09
Page metadata (front matter)
| type | note |
|---|---|
| id | E-20260826b-critic-response |
| status | frozen |
| created | 2026-08-26 |
| updated | 2026-08-26 |
| links | workshop/experiments/E-20260826b-radif/design.md, workshop/experiments/E-20260826b-radif/critic-v1.json, workshop/translations/saadi-ghazals-radif/R51-v1/arms.md |
Pre-run critic, one round, two seats — and what v2 did about it
Run before any judging call existed, on design v1 and the exact materials, which is why
several of the findings are ones no reading of the design alone could have produced. Seats: P1
openai/gpt-5.6-terra (complete, 11,166 chars) and P2 google/gemini-3.6-flash (truncated at
max_tokens 9000, 3,093 chars — note (brr) in its fourth form; three findings are all this
session has from it). Cost $0.10866075. Raw: critic-v1.json.
Verdict: NEEDS REDESIGN. Fifteen findings, seven BLOCKING. Every one accepted; none overruled.
The BLOCKING findings and what changed
1. P1 — the varied tails were not meaning-preserving. "safe-conduct's whole claim departed"
for arose reverses the event; has ended up for has fallen changes the action. So AB–RH
and RD–NN measured repetition removal bundled with a changed proposition, and E_rep had no
repetition-only reading.
→ Accepted. arms.md rule 2 written: a varied tail must be truth-conditionally equivalent to
the AB tail in its own line. departed and has ended up are gone. The irony is worth
recording: those two tails entered in the hand's own repair of a v1 defect (two tails sharing the
word up), so a repair made to protect the manipulation broke it. The rule now says content-word
overlap between tails is forbidden and function-word overlap is permitted and declared — chasing
that residue is what caused the damage.
2. P1 — at L3 and L4, AB repeats a bearer word and RD did not, so the supposedly
chime-only contrast also removed a lexical repetition. E_chime and E_int were both contaminated.
→ Accepted, and this is the finding that most changes the run. arms.md rule 3: RD
reproduces AB's own pattern of repeated bearers. L3 RD is now reproach · warrant ·
reproach; L4 RD is penance · repute · penance — and penance is available at both of L4's
repeated positions because ندامت is remorse and غرامت is a penalty and penance is both.
3. P1 — C1 could not catch what it was for. An additions/omissions question does not detect
antonymy, changed implication, or altered image, and the retention rule tolerated one DIFFERENT per
locus, so a locus could survive the very check that invalidated one of its edges.
→ Accepted. C1 is rebuilt in v2 §7 as a per-edge equivalence gate asking directly about
sameness of assertion and contradiction, applied to each of the four edges rather than to each
arm against AB. An edge that returns DIFFERENT is dropped from the effect it feeds, and the
locus is dropped from that effect; n is reported per effect.
4. P1 — F2 was post-data discretion. It permitted withholding an edge after seeing the
preferences without specifying what then happened to the loci, denominators or tests.
→ Accepted. F2 is removed as a withholding rule. Saturation is reported as a diagnostic and
the registered data stand.
5. P1 — L3 and L8 were in the primary pool although the mechanical grader rejects their AB
chime, leaving a later reader free to pick whichever reading suits.
→ Accepted. v2 pre-specifies the confirmatory set as the five loci the tool and the hand
agree about (L1, L2, L4, L6, L7) and L3/L8 as a labelled sensitivity set. The cost
is real and is stated: n = 5 confirmatory, whose one-sided exact sign test cannot go below
P = 0.031.
6. P2 — L8's RD/NN broke an anaphora. Line 1's line-end lane is the antecedent of line
2's that lane; substituting street left the demonstrative dangling, and a seat would punish the
grammatical fault rather than the manipulation.
→ Accepted. All four L8 arms now read "and there many a one…" — the window differs from the
rendering of record by one mid-line word, in every arm equally, and it says so.
7. P2 — the varied tails were clumsy and register-clashing, so "which reads better as English
verse" would track prose quality and inflate E_rep.
→ Accepted. Every tail re-authored for register; arms.md's substitution key is the result.
The residue is not eliminated and v2 says so: a varied tail is longer than is there or
arose however carefully chosen, and that weight difference is a standing confound on E_rep,
listed in §9.
The SERIOUS and MINOR findings
| # | seat | finding | action |
|---|---|---|---|
| 8 | P1 |
P2-the-prediction is not a validity gate for P1-the-prediction; a positive chime result does not validate the repetition manipulation |
Accepted. F1 rewritten: E_chime is a replication check on the chime package, and P1's interpretability now rests on C1's per-edge gate, not on P2 |
| 9 | P1 |
the order diagnostic was global and could hide offsetting per-edge effects; and only E_rep had a consequence |
Accepted. F4 is now per edge, with the same consequence registered for all three quantities |
| 10 | P1 |
the sign test left ties undefined, and with denominator six a locus can land exactly on 0 | Accepted. v2 registers the tie rule: a zero counts as not supporting the one-sided prediction |
| 11 | P1 |
seven loci clustered in three ghazals with shared radifs are not seven independent texts | Accepted. v2 §5 calls the sign test a descriptive consistency summary over constructed loci and §9 says the material basis is three poems |
| 12 | P1 |
depth is confounded with locus length; pooling depth 2 and 3 hides it | Accepted. Depth-2 and depth-3 are reported separately for every quantity |
| 13 | P1 |
the saturation criterion leaned on whether volunteered reasons name a line-end property, which tests nothing | Accepted. Reasons are qualitative annotation only and gate nothing |
| 14 | P1 |
side-by-side presentation makes the manipulation conspicuous whatever the wording | Accepted. §9 (1) rewritten: this measures explicit comparative preference after salient line-end substitution, never spontaneous response |
| 15 | P2 |
C1 would return DIFFERENT routinely on near-synonyms and over-invalidate loci |
Accepted in part. The gate now asks about same assertion / contradiction, not about wording, and a near-synonym difference that is not a difference of assertion is SAME. The risk of over-invalidation remains and the reported n per effect makes it visible |
The one finding not acted on as its author proposed
P1's BLOCKING 2 says neither factor is isolated from lexical, syntactic and prosodic change, and
offers two remedies: narrow the estimand, or rebuild the arms so each factor changes nothing but
phonology — adding that "if that cannot be done in English, the device-decomposition experiment is
not buildable with these passages."
The second remedy cannot be taken, and the reason is the session's own finding. You cannot remove
a chime from English without changing a word, and you cannot un-repeat a repeated phrase without
changing words. L5 was withdrawn for exactly this — at a comparative radif English supplies a chime
whether or not one is wanted — and RS-20260821b-matched-heard reached the same wall from the other
side at a matched shape.
So the first remedy is taken, in full. v2 states that E_rep and E_chime estimate
preference between constructed line-end packages, that repetition and chime name the
intended manipulations rather than isolated causes, and that any sentence written into
wiki/goodness-senses.md carries that qualification at the same size as the number. It is the
honest reading of what these materials can support, and it is smaller than what v1 was reaching
for.