Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260804d-gounod-blunders/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260804d-gounod-blunders
statusfrozen
created2026-08-04
updated2026-08-04
sensesaccuracy
provisionaltrue
linkswiki/base/sources/S-nation-1896-gounod.md, wiki/arms/ARM-reception-claims.md, workshop/translations/memoires-dun-artiste/R04-v1/translation.md, config/models.md, wiki/goodness-senses.md, wiki/base/searches/SR-20260804-rival-review-sweep.md

E-20260804d — are an 1896 reviewer's five numbered "blunders" there?

Frozen 2026-08-04 (S104) before any API call. The lead's own translation of the flagged passages and its log were frozen and committed first (767a82a), so nothing in this design could have been written to suit them.

1. Question

The Nation of 2 July 1896, reviewing the two English versions of Gounod's Mémoires d'un artiste that appeared that year, wrote that Hutchinson's is "more terse and idiomatic, and free from the occasional blunders which occur in the first (pp. 111, 118, 125, 127, 139)".

That is a rare thing in the reception record: a comparative judgement about two translations that is indexed to specific pages and therefore falsifiable 130 years later. The question is not whether the reviewer was a good critic. It is:

Do the pages a contemporary reviewer named as carrying mistranslations in fact carry them — and do unnamed pages of the same book not?

What this teaches about evaluating translations, in one sentence (the subject rule): it measures how much of a period critic's comparative verdict survives contact with the source, on the one kind of claim — a located factual error — where "survives" has a determinate meaning.

2. Materials

Source Gounod, Mémoires d'un artiste, Calmann Lévy, 3rd ed. 1896. PD. archive.org memoires-dun-artiste-par-charles-gounod-calmann-levy-dvg
Arm A Annette E. Crocker, Memoirs of an Artist, Rand McNally, 1895/96. PD. archive.org memoirsanartist00goungoog
Arm B W. Hely Hutchinson, Charles Gounod: Autobiographical Reminiscences, Heinemann / Lippincott, 1896. PD. archive.org charlesgounodaut00goun
The claim The Nation (New York) vol. 63 no. 1618, 2 July 1896, p. 18 (scan page 20, col. 2). Read off the page pixels via IIIF, not from OCR — S-nation-1896-gounod

Loci. Five flagged = the reviewer's five pages. Five control = the pre-registered rule in materials/locate.py (--controls): the five printed pages in [105, 145] chosen greedily to maximise minimum distance from any already-chosen page, ties to the lower number → 105, 108, 114, 133, 145. The rule is in code, was written before the controls were seen, and does not use any property of the text.

Alignment. Arm A needs none: a locus is a Crocker printed page (printed page N = scan N+12, asserted on every recovered header numeral). The French span behind each page is found by proper-noun fingerprint under a monotone token-position fit, with the route recorded per locus (fingerprint / fingerprint-weak / fit). Arm B has no page handle and is aligned to the French, never to Crocker, by fingerprint plus a longest-increasing-subsequence monotonicity repair — because Hutchinson's volume interleaves family letters with the memoir and a single linear fit put two loci in the correspondence.

3. Declared materials hazards

  1. OCR is a single witness. Only one scan of each of the three books is on archive.org, so note (aa)'s two-scan precondition cannot be met and is not claimed. Mitigation, and it is the binding one: every error a seat reports is verified against the page image via IIIF before it is counted. An error that is an OCR artifact is void (F3).
  2. The French OCR is known to be wrong in at least one place. At L125 the scan reads 1857 and the page reads 1851 (verified on the page image, word box 780,1211,884,1171 of French scan 0135; translator's log D4). Every four-digit numeral inside the ten loci is image-checked before the run and the corrections are stored in materials/ocr-fixes.json.
  3. L125 and L127 overlap by about thirty words, because the flagged pages are two apart and the emitted spans carry a margin. Not smoothed: the loci are the scored unit.
  4. Arm B spans are wider than arm A spans and may carry neighbouring sentences. Stage 2 is secondary for this reason and its instruction says so.
  5. The canary item I05 and the control item I03 are the same locus, one seeded and one not. A seat may notice the near-duplication. This is declared; it is also what makes the canary a matched control rather than a free-standing one.

4. Procedure

Stage 1 (primary). 11 items — 5 flagged, 5 control, 1 canary — each presented as the French span and one English rendering (arm A). No arm is named, no page number is shown, item ids carry no set information. Seats P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P5 deepseek/deepseek-v4-pro (config/models.md), each in two orderings (forward and reversed) → 6 calls. Per item a seat returns ITEM=Inn ERRORS=k and, for each error, a line carrying a verbatim English quote and the French it misrenders.

Stage 2 (secondary). The same 10 real loci with arm B instead of arm A, 3 seats × 1 ordering → 3 calls.

Stage 3 (control on the instrument, one call per seat). After stages 1–2, each seat is asked whether it can name either translator. S103 found recognition total on a Victorian Russian pair; whether it holds for two obscure 1896 Gounods is a condition on how any preference here is read.

Seats are used as CODERS of factual error, not as jurors. No quality judgement is asked for and none is reported. Tier D is NOT PASSED; every evaluative sentence downstream carries provisional: true.

P3 is mechanical and costs nothing: English tokens per French token, arm A vs arm B, per locus.

5. Registered predictions

6. Failure criteria — what withholds a result

7. Budget

Pre-flight worst case, built from max_tokens and the worst plausible provider (note (abc), and the routing caution in config/models.md): critic 1 × 12,000 → $0.22; stage 1 six calls × 6,000 → $0.42; stage 2 three calls × 6,000 → $0.21; stage 3 three calls × 1,500 → $0.05. Worst case $0.90. Today's headroom at session start: $3.998 of $5.00.

8. What this cannot show

It cannot show that the reviewer was right about the book: five pages of 223 are a sample he chose, not a census, and a page he did not flag may still contain an error he did not notice or did not think worth printing. It cannot rank the two translations. It says nothing about "idiomatic", which is not a located claim, and it does not test the reviewer's actual overall verdict — "Both are well done" — which is not falsifiable in this form.


9. Amendments A1–A10, from the independent pre-run critic pass

runs/critic_kimi-k3_try1.txt, P4 moonshotai/kimi-k3, NEEDS-AMENDMENT, 5 BLOCKING + 5 ADVISORY, all ten accepted, applied before any scoring call was dispatched.