Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260804b-quixote-affect/critic.md · rendered 2026-09-09

Page metadata (front matter)
typenote
idE-20260804b-critic
statusfrozen
created2026-08-04
updated2026-08-04
linksworkshop/experiments/E-20260804b-quixote-affect/design.md

Pre-run critic pass — E-20260804b

One pass, one call. moonshotai/kimi-k3 (P4), provider and cost in runs/critic_kimi-k3_try1.meta.json — $0.07368, 201.0 s, 4,430 chars. Verdict NEEDS-AMENDMENT: 5 BLOCKING, 5 ADVISORY. Raw body at runs/critic_kimi-k3_try1.raw, verbatim text at runs/critic_kimi-k3_try1.txt.

All ten findings are accepted. None is refused. The amendments below were written into design.md before any grading call was dispatched, and the design's SHA and materials hash were re-frozen after the rebuild. Amendment ids A1–A10 correspond to findings 1–10.

# kind the finding, compressed amendment
1 BLOCKING PR5 was near-unfalsifiable — with 42 picks over 6 arms, modal-FUNNIEST ≠ modal-MATCH is the default under noise, so the only falsifying outcome was exact modal identity A1: PR5 restated directionally — modal FUNNIEST ∈ {motteux,nudge} and modal MATCH ∈ {shelton,smollett,ormsby,lead} — and the uniform-null probability of a bare modal split is computed and reported alongside
2 BLOCKING Even a clean PR5 could not support the inference, because Stages 1 and 2 differ in task, source-presence and key order at once A2: a within-stage contrast added — Stage 1 now asks FUNNIEST-SP in the same call, on the same key order, immediately after MATCH. That pair is the crux test; Stage 2 becomes its source-absent replication. The cross-stage inference is explicitly softened in §1
3 BLOCKING Training-data recognition was undeclared — four canonical published arms, and Ormsby's introduction with its verdict, are almost certainly in every grader's training data, so letter keys do not prevent recognition and the seats may recite the record A3: a recognition probe added at the end of Stage 2 (not Stage 1, so no authorship question is anywhere in the prompt while SIGNAL/MATCH is being answered): name the translator per key or write UNKNOWN. Rates reported. New declared limit L6
4 BLOCKING No prediction tested the motivating reception fact — nothing predicted that motteux beats anyone on funniness A4: PR6 added — motteux FUNNIEST picks > ormsby FUNNIEST picks
5 BLOCKING F5's exclusions rested on one seat, one call, no replication — and S5 is exactly the site that tests whether that seat is right, since the Spanish there does carry the aside «y fuera mejor que se curara, porque fuera curarse en salud» A5: Stage 0 run on two independent source-only seats — P5 and the config/models.md first reserve qwen/qwen3.7-max. A site is excluded from PR2/PR3 if either seat says the Spanish itself signals — excluded by default on disagreement, as the finding asks
6 ADVISORY The PR4 manipulation confounded adds a narratorial signal with adds a new joke — "more fool he", "a pretty prayer it was" are themselves gags A6: a seventh arm, nudge-flat — lead plus an intrusive narratorial comment of matched length that is evaluative but not comic. New predictions PR4a (the operator is detected when it is not funny) and PR4b (nudge > nudge-flat on funniness)
7 ADVISORY F6's length check used a Pearson r over six points per site and never isolated the pair PR4 compares A7: F6 replaced by a pair-level check — per site, nudge−lead pick difference against nudge−lead word difference; and nudge-flat is length-matched to nudge by construction, which handles it structurally
8 ADVISORY SIGNAL carried PR1–PR3 with no seat-agreement criterion (F4 covered MATCH only) A8: F8 added — majority attainment computed over every (site, arm, ordering) cell; below 0.80, SIGNAL is descriptive only
9 ADVISORY F3 was so aggressive that PR4/PR5 were likely to be withheld by chance A9: F3 now fires only where the order effect reverses the sign of the relevant pick difference; otherwise both orderings are reported and pooled
10 ADVISORY At S3 Motteux omits the passage and cannot signal there, depressing pooled SIGNAL(motteux) by an operation the design itself excludes from SIGNAL A10: PR2 pre-registered to be reported both with and without S3, rather than decided after the data exists

What the pass cost and what it bought

Finding 3 is the one this session would not have caught. The design had a RS-20260804-yardstick-shaped defect it had already guarded against in the obvious place — no lead-authored gloss, an independent mechanism statement — and missed the other channel by which a canonical text can carry its own verdict into a grader: the graders may already know both the translations and the criticism of them. No amendment can remove that; A3 measures it and L6 declares it.

Finding 5 is the one that is materially right about a specific site: the Spanish at S5-no-heed is No se curó el arriero destas razones (y fuera mejor que se curara, porque fuera curarse en salud), and that parenthesis is a narratorial aside. If Stage 0 says so, S5 leaves PR2 and PR3 — which is the design working, not failing.