Repository path: workshop/experiments/E-20260804d-gounod-blunders/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260804d-gounod-blunders |
| status | frozen |
| created | 2026-08-04 |
| updated | 2026-08-04 |
| senses | accuracy |
| provisional | true |
| links | wiki/base/sources/S-nation-1896-gounod.md, wiki/arms/ARM-reception-claims.md, workshop/translations/memoires-dun-artiste/R04-v1/translation.md, config/models.md, wiki/goodness-senses.md, wiki/base/searches/SR-20260804-rival-review-sweep.md |
E-20260804d — are an 1896 reviewer's five numbered "blunders" there?
Frozen 2026-08-04 (S104) before any API call. The lead's own translation of the flagged
passages and its log were frozen and committed first (767a82a), so nothing in this design could
have been written to suit them.
1. Question
The Nation of 2 July 1896, reviewing the two English versions of Gounod's Mémoires d'un artiste that appeared that year, wrote that Hutchinson's is "more terse and idiomatic, and free from the occasional blunders which occur in the first (pp. 111, 118, 125, 127, 139)".
That is a rare thing in the reception record: a comparative judgement about two translations that is indexed to specific pages and therefore falsifiable 130 years later. The question is not whether the reviewer was a good critic. It is:
Do the pages a contemporary reviewer named as carrying mistranslations in fact carry them — and do unnamed pages of the same book not?
What this teaches about evaluating translations, in one sentence (the subject rule): it measures how much of a period critic's comparative verdict survives contact with the source, on the one kind of claim — a located factual error — where "survives" has a determinate meaning.
2. Materials
| Source | Gounod, Mémoires d'un artiste, Calmann Lévy, 3rd ed. 1896. PD. archive.org memoires-dun-artiste-par-charles-gounod-calmann-levy-dvg |
| Arm A | Annette E. Crocker, Memoirs of an Artist, Rand McNally, 1895/96. PD. archive.org memoirsanartist00goungoog |
| Arm B | W. Hely Hutchinson, Charles Gounod: Autobiographical Reminiscences, Heinemann / Lippincott, 1896. PD. archive.org charlesgounodaut00goun |
| The claim | The Nation (New York) vol. 63 no. 1618, 2 July 1896, p. 18 (scan page 20, col. 2). Read off the page pixels via IIIF, not from OCR — S-nation-1896-gounod |
Loci. Five flagged = the reviewer's five pages. Five control = the pre-registered rule
in materials/locate.py (--controls): the five printed pages in [105, 145] chosen greedily to
maximise minimum distance from any already-chosen page, ties to the lower number → 105, 108, 114,
133, 145. The rule is in code, was written before the controls were seen, and does not use any
property of the text.
Alignment. Arm A needs none: a locus is a Crocker printed page (printed page N = scan N+12,
asserted on every recovered header numeral). The French span behind each page is found by
proper-noun fingerprint under a monotone token-position fit, with the route recorded per locus
(fingerprint / fingerprint-weak / fit). Arm B has no page handle and is aligned to the
French, never to Crocker, by fingerprint plus a longest-increasing-subsequence monotonicity
repair — because Hutchinson's volume interleaves family letters with the memoir and a single linear
fit put two loci in the correspondence.
3. Declared materials hazards
- OCR is a single witness. Only one scan of each of the three books is on archive.org, so note (aa)'s two-scan precondition cannot be met and is not claimed. Mitigation, and it is the binding one: every error a seat reports is verified against the page image via IIIF before it is counted. An error that is an OCR artifact is void (F3).
- The French OCR is known to be wrong in at least one place. At L125 the scan reads 1857 and
the page reads 1851 (verified on the page image, word box 780,1211,884,1171 of French scan
0135; translator's log D4). Every four-digit numeral inside the ten loci is image-checked before
the run and the corrections are stored in
materials/ocr-fixes.json. - L125 and L127 overlap by about thirty words, because the flagged pages are two apart and the emitted spans carry a margin. Not smoothed: the loci are the scored unit.
- Arm B spans are wider than arm A spans and may carry neighbouring sentences. Stage 2 is secondary for this reason and its instruction says so.
- The canary item I05 and the control item I03 are the same locus, one seeded and one not. A seat may notice the near-duplication. This is declared; it is also what makes the canary a matched control rather than a free-standing one.
4. Procedure
Stage 1 (primary). 11 items — 5 flagged, 5 control, 1 canary — each presented as the French
span and one English rendering (arm A). No arm is named, no page number is shown, item ids
carry no set information. Seats P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash,
P5 deepseek/deepseek-v4-pro (config/models.md), each in two orderings (forward and
reversed) → 6 calls. Per item a seat returns ITEM=Inn ERRORS=k and, for each error, a line
carrying a verbatim English quote and the French it misrenders.
Stage 2 (secondary). The same 10 real loci with arm B instead of arm A, 3 seats × 1 ordering → 3 calls.
Stage 3 (control on the instrument, one call per seat). After stages 1–2, each seat is asked whether it can name either translator. S103 found recognition total on a Victorian Russian pair; whether it holds for two obscure 1896 Gounods is a condition on how any preference here is read.
Seats are used as CODERS of factual error, not as jurors. No quality judgement is asked for and
none is reported. Tier D is NOT PASSED; every evaluative sentence downstream carries
provisional: true.
P3 is mechanical and costs nothing: English tokens per French token, arm A vs arm B, per locus.
5. Registered predictions
- P1 — the reviewer's claim is specific. A locus counts as carrying an error when ≥2 of 3 seats flag the same French point (majority coding, pooled over both orderings). Predicted: ≥3 of 5 flagged loci carry a majority-coded error in arm A, and ≤1 of 5 control loci do.
- P2 — the reviewer's comparison holds. At the five flagged loci, arm B carries fewer majority-coded errors than arm A.
- P3 — "more terse" is measurable and true. Arm B's English-tokens-per-French-token is lower than arm A's at ≥7 of 10 loci.
- P4 — the lead's frozen log agrees with the reviewer. Log D24, written before the reviewer's list was consulted for content, ranks L125 and L127 as the hardest and names L111 and L118 as presenting "nothing I would call a trap". Predicted: the two loci the log calls trap-free carry no majority-coded error.
6. Failure criteria — what withholds a result
- F1 (power). If the canary is not flagged by ≥2 of 3 seats, the instrument is not shown to detect this class of error, and P1 is reported as underpowered rather than as a negative.
- F2 (quote grounding). Any reported error whose English quote is not present verbatim in the stored arm text is void. Reported as a rate.
- F3 (OCR). Any surviving error that the page image shows to be a scanning artifact is void, and named.
- F4 (specificity of the canary). If the canary's seeded error is flagged but the unseeded copy of the same locus (I03) is also flagged at the same point, the seats are flagging the locus, not the error, and the canary licenses nothing.
- F5 (discrimination). If ≥3 of 5 control loci carry majority-coded errors, the reviewer's list does not distinguish flagged pages from unflagged ones, and that is the result.
- F6 (seat agreement). If pairwise seat agreement on locus-level "any error" is below 0.50, the primary is reported as unreliable whatever it shows.
- F7 (order). If the flagged−control difference reverses between the two orderings, the primary is reported conditional and both orderings printed.
7. Budget
Pre-flight worst case, built from max_tokens and the worst plausible provider (note (abc), and
the routing caution in config/models.md): critic 1 × 12,000 → $0.22; stage 1 six calls ×
6,000 → $0.42; stage 2 three calls × 6,000 → $0.21; stage 3 three calls × 1,500 →
$0.05. Worst case $0.90. Today's headroom at session start: $3.998 of $5.00.
8. What this cannot show
It cannot show that the reviewer was right about the book: five pages of 223 are a sample he chose, not a census, and a page he did not flag may still contain an error he did not notice or did not think worth printing. It cannot rank the two translations. It says nothing about "idiomatic", which is not a located claim, and it does not test the reviewer's actual overall verdict — "Both are well done" — which is not falsifiable in this form.
9. Amendments A1–A10, from the independent pre-run critic pass
runs/critic_kimi-k3_try1.txt, P4 moonshotai/kimi-k3, NEEDS-AMENDMENT, 5 BLOCKING + 5
ADVISORY, all ten accepted, applied before any scoring call was dispatched.
- A1 (finding 1, BLOCKING) — the recognition probe moves BEFORE stage 1 and now asks about the
review. It becomes stage 0, and adds
REVIEW=andFLAGGED_PAGES=. A seat that can name the 1896 review, or its five page numbers, is replaying the verdict rather than coding errors. Pre-registered discount: if any seat returns the flagged page list, the primary is reported CONFOUNDED and stage 1 is re-read as recall, not detection. - A2 (finding 2, BLOCKING) — P2 is WITHDRAWN as a prediction. Arm B is abridged and omission is out of scope, so "fewer errors" is structurally favoured whatever its accuracy. Stage 2 is reported descriptively, with a coverage figure — the fraction of each locus's French content words with any counterpart in arm B — printed beside every count, and the abridgement confound is named in §8.
- A3 (finding 3, BLOCKING) — P3 is DEMOTED to a manipulation check with no evidential weight, and the design states that it cannot fail: the materials were described as abridged before it was registered.
- A4 (finding 4, BLOCKING) — the aggregation rule, stated mechanically and implemented in
analysis/score.py, not by hand. (i) A reported error belongs to the locus of its item id. (ii) Two reports are the same point iff their normalised FRENCH anchors share ≥1 content token or their English QUOTEs share a ≥3-token run. (iii) A seat flags a point if it flags it in either ordering (union); the intersection (both orderings) is computed and printed as the strict variant. (iv) No step is adjudicated by the lead, and the code never readsarm_set. - A5 (finding 5, BLOCKING) — the seats are told the FRENCH is OCR too, in both stage headers, and F3 is extended: an error is void if either side of it is a scanning artifact, checked against the page image.
- A6 (finding 6, ADVISORY) — the L125/L127 overlap. An error whose French anchor falls in the shared thirty words counts at the earlier locus only; the primary is printed both ways.
- A7 (finding 7, ADVISORY) — the two-control case is registered now. P1 is met iff flagged ≥3 and control ≤1. Exactly 2 controls positive: P1 NOT met, reported as "the list discriminates weakly" only if flagged ≥4, otherwise simply not met. F5 still fires at ≥3.
- A8 (finding 8, ADVISORY) — a second canary,
Beethoven → Schubertat L145, alongsideRossini → Verdiat L105. F1 now needs both. Presentation order is a seeded shuffle (20260804) constrained to keep each canary ≥3 slots from the unseeded copy of its own locus. Language downgraded where n is still small. - A9 (finding 9, ADVISORY) — error opportunity is measured, not assumed.
analysis/score.pycounts error-eligible tokens (proper nouns, dates, numerals) per locus and prints the flagged vs control comparison beside the primary. If the flagged loci are richer in them, the primary is reported with that ratio attached. - A10 (finding 10, ADVISORY) — the span emission is re-derived, not trusted.
locate.pyis re-run and its output asserted byte-identical to the storedflagged.json/control.jsonbefore the run, and the assertion is part ofanalysis/verify.py.