Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260730-grain-clause/runs/critic.output.md · rendered 2026-09-09

critic.output.md

VERDICT: NEEDS-AMENDMENT

FINDINGS: 1. F4 cannot fire through rule logic; NORULE is structurally unreachable. Test 4 ("load-bearing or furniture") returns a handling for every conceivable item, so a rule-following reader can never reach NORULE. F4 and the NORULE half of Q8 therefore measure only reader non-compliance, which F5 already handles. Severity: MANDATORY. Fix: re-scope F4 to count UNSURE flags (>2 in a condition → condition flagged) and re-scope Q8 to UNSURE only; no new materials needed. 2. Q1 is one sample and cannot support the gate F1 places on it. A single byte-identical repeat per reader yields one agreement number; passing within 0.09 once does not establish pass stability, and failing once does not establish instability — S049's +0.175 shows the variance is the object of interest. Severity: MANDATORY. Fix: run the C15 condition twice per reader (6 calls, ~$0.50, no new materials) and gate F1 on both repeats. 3. F1 gates raw agreement but Q2 is a κ claim. The pass can repeat within 2 sites raw while κ moves far more (9-category chance correction at n=23 is high-variance), so F1 can open the gate for a κ comparison the retest never validated. Severity: MANDATORY. Fix: add to F1 a κ-retest clause: C15 repeat κ within a registered band of 0.452, else no between-rule κ is reportable. 4. Q4 and F3's Jaccard ≥ 0.80 are dominated by the shared first limb. All three rules share test 3's borrowing clause verbatim, so every borrowing site is in all three firing sets, inflating all pairwise Jaccards and making 0.80 nearly automatic regardless of the second limb's behaviour. The warrant-preservation measurement is therefore largely unfailable as specified. Severity: BLOCKING for F3. Fix: pre-commit to computing Jaccard on second-limb decisions only — operationally, sites where test 3 fires and the item is not an established borrowing (codable from the frozen glosses without new materials), or equivalently on test-3-vs-test-4 flips relative to C15. 5. Arm A's "firing set" is computed per reader but the design never says how the two readers' sets are combined. P1 and P3 will disagree; Jaccard of whose set against whose? Underspecified primary statistic. Severity: MANDATORY. Fix: register "per-reader Jaccard, reported separately, prediction assessed on P1~P3-averaged Jaccard" before dispatch. 6. Arm J cannot deliver the pair conclusion it is announced for. Pair, passage, site count (34 vs 23), label distribution, and instruction block all differ, and Q6 compares raw agreement across those lists. Confound 2 admits this, yet §2's framing ("arm J is where it shows") and Q6/Q7 plan to draw the pair conclusion anyway. Severity: MANDATORY. Fix: demote Q6/Q7 to descriptive in the design text and pre-register that no pair-attributive claim will be made from arm J; report κ alongside raw agreement. 7. The sham's criterion has its own reproducibility hazard: romanisation choice. C15″ requires counting letters of "a romanisation of the source item"; transliteration schemes differ (хата → khata/hata), so readers can disagree under C15″ for reasons having nothing to do with mechanisation-vs-aptness. A C15″ κ below C15′ (the F3 pattern) is then confounded. Severity: MANDATORY. Fix: register this confound alongside confound 1 and pre-commit that F3 requires the arm-A retrodiction to hold and a stated caveat on transliteration ambiguity; no new materials needed. 8. Reserve substitution silently breaks the primary pair. F5 lets a failed P3 call fall through to P2, after which "P1~P3" is P1~P2 and is no longer the S056 instrument, invalidating Q1/Q2 as retests. Severity: MANDATORY. Fix: register that any reserve substitution in arm R voids the retest comparisons for that reader's conditions, which are then reported as new data like P5's. 9. Q2's ≥ 0.15 threshold sits inside demonstrated instrument drift. S049 moved raw agreement +0.175 on a byte-identical repeat; the design's own expected effect is smaller than observed noise. Findings 2–3 partly cover this, but the threshold itself should be justified against the retest variance measured this session, not asserted. Severity: ADVISORY. Fix: state in the design that Q2 is assessed only conditional on F1's strengthened gate. 10. Confound 4 (author wrote the control after reading the nine disagreement sites) is declared but not mitigated, and it points directly at the expected outcome. The session wants Q3 to hold and had full freedom to make C15″ crisp exactly where C15 was loose. Declaration is not control. Severity: ADVISORY (cannot be removed without new materials). Fix: add a registered sentence that Q3 holding is weak evidence for the mechanisation reading precisely because the control was tunable, and that F2's closure of the arm should note this. 11. Confound 1 does not make the manipulation unsound, but it makes Q2's effect size uninterpretable, not just its null. The residual "exact" in the EQUIVALENT label is identical across conditions, so between-rule comparisons are unbiased; but the design says the residual reading "must be offered wherever Q2's outcome is reported" only for a null. An attenuated positive is equally consistent with the residual. Severity: ADVISORY. Fix: extend the mandated caveat to all Q2 outcomes.

PREDICTIONS THAT CANNOT FAIL: Q8 (NORULE half — test 4 always returns a handling); Q4-as-specified and the Jaccard ≥ 0.80 clause inside F3 (shared first limb inflates all overlaps — near-unfailable until re-scoped per Finding 4).

FAILURE CRITERIA THAT CANNOT FIRE: F4 (NORULE structurally unreachable; only fires on parse-level non-compliance already covered by F5).

WHAT YOU WOULD MEASURE INSTEAD: Two C15 repeats per reader, gating on both raw agreement and κ retest. On arm A, per-reader sets of sites where each rule changes C15's test-3-vs-test-4 decision, compared between C15′ and C15″ — divergence from C15 is the measurable warrant signal, not raw firing-set overlap. On arm J, κ and test-3 citation rate reported descriptively with no pair attribution. UNSURE counts as the completeness metric. Reserve substitutions registered as voiding retest status.