Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260725-svidanie-audit/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260725-svidanie-audit
statusfrozen
created2026-07-25
updated2026-07-25
sensesaccuracy
internal-judgment-onlytrue
linkswiki/base/searches/SR-20260725-parity-evidence.md, workshop/translations/svidanie/R04-v1/translation.md, workshop/experiments/E-20260725-svidanie-audit/landmarks.json, tools/audit_landmarks.py, workshop/experiments/E-20260725-turgenev-filter/design.md, wiki/decisions/resolved/D-20260725-06-heldout-arm-operationalisation.md

Frozen design — auditing Garnett and Hapgood on the one passage the comparative evidence is about

Frozen 2026-07-25 (S024), before any published English rendering of «Свидание» was fetched. Verifiable in git: T-svidanie-R04-v1, landmarks.json and this page are committed before the first retrieval of Garnett or Hapgood.

1. Why this runs, and how it is wired to the study limb

SR-20260725-parity-evidence closed the project's best-configured Tier D candidate — Turgenev in Garnett and Hapgood — under exclusion code X2: comparative evidence exists, at passage level, naming both, and it ranks. The document is soundslikewish.org/?p=7860, and the passage it compares is the opening of «Свидание», not «Бежин луг» — the parity remark it contains ("Her choice of 'not … nor … nor' is as good as Hapgood's 'not … not … not'") is about the four-fold negation, which occurs in that paragraph and nowhere else in the sketch.

That search record's own §9 names the observation that would reopen the candidate:

A second, independent assessment of Garnett vs Hapgood contradicts the blog's ranking. Would put case 1 back to genuinely unresolved rather than excluded.

S023 produced two straws in that direction and did not weigh them: on «Бежин луг» the single factual divergence found ran against the blog's direction (Garnett turns «угловатую» into "square"; Hapgood keeps "angular"), and Hapgood's «лапти» → "lindenbark slippers" was the most materially precise rendering in the run. Neither was a designed test.

The wire, in one sentence: the lead's blind translation of the «Свидание» opening supplies the source-derived landmark list that tests, on the very passage the project's only comparative-reception document assesses, whether that document's ranking of Garnett above Hapgood is corroborated or contradicted on factual accuracy.

What this wire can and cannot deliver, said before the run rather than after. The blog ranks on prose merit; this audit measures factual damage. They are different dimensions, and a flag count cannot refute an aesthetic judgment. What an accuracy result can do is supply one independent, non-aesthetic assessment of the same two renderings of the same paragraph, which is strictly more than the project has now (one unrefereed blog post). If it runs against the ranking it makes case 1 less settled than SR-20260725-parity-evidence recorded; it does not by itself reopen the gate, and this page will not claim it does.

2. Question

On the opening paragraph of «Свидание» — the one passage about which comparative criticism of these two translators exists — does either Constance Garnett's or Isabel Hapgood's rendering carry source-anchored factual damage, and does the distribution of damage run with or against the only published ranking of the pair?

3. Materials

label text provenance
S Turgenev, «Свидание» (1850), opening paragraph, 468 words ru.wikisource cross-checked word-by-word against az.lib.ru/t/turgenew_i_s/text_0080.shtml. The two disagree at three points and each carries a unique typo; the stored text takes wikisource with az.lib.ru's two corrections. workshop/translations/svidanie/R04-v1/source-ru.txt, SHA-256 1607c1a3ec7ba4c2c2b8c01d8bbaed9012e0790b9b2a6147540bff04ad37babe
T1 Constance Garnett, A Sportsman's Sketches — "The Tryst" Project Gutenberg — not yet fetched at freeze time
T2 Isabel F. Hapgood, Memoirs of a Sportsman — "The Rendezvous" (title unverified) archive.org — not yet fetched at freeze time. Per note (aa), two independent scans must agree on every landmark site before any disposition is read
T3 the lead's T-svidanie-R04-v1 frozen before T1/T2 were fetched; 639 English words

4. The instrument

tools/audit_landmarks.py v3 against landmarks.json — 17 landmarks, typed, accept patterns frozen verbatim. Types: one order (C3, the narrator's aspen → settle → sleep sequence), six values (C1 the date, C7 the winter sun, C9 the absent birds, C14 the summer evenings, C15 the low boughs, C16 the hunters' sleep), seven conjunctions (C4 white silk, C5 overripe grapes, C6 new snow, C8 the red-or-gold young birch, C10 the tomtit's steel bell, C12 grey-green metallic, C13 round leaves on long stalks), one cooccur with declared rivals (C11, the pale-lilac bound to the trunk and not to the foliage a few words later), two presences (C2 birch, C17 the dog).

v3 exists because v2 produced two bad observations in S023, both logged and neither fixed then (E-20260725-turgenev-filter/verification.md §7). Both are fixed here and the fixes are demonstrated in runs/instrument-regression.out, which runs the two offending landmarks verbatim from the frozen v2 spec:

defect v2 v3
(a) accept sets punish precision. B11's frozen \bbark\b could not match Hapgood's "lindenbark slippers" — the best rendering of лапти in the run FLAG — first term absent PASS, with compound: true relaxing the boundary and the observation naming the containing word
(b) presence confusables not anchor-scoped. Garnett's B16 reported found instead: ['coat'] from a coat two sentences away names an unrelated word absent (confusables not reported: no anchor declared), or scoped to a declared anchor window

Compound relaxation does not suppress over-matching — it makes it visible: bark inside barking reports 'barking' and is adjudicated by hand. That is deliberate. A filter meant to disqualify must not silently prefer the vocabulary its author happened to think of.

v3 also adds a precision-probe gate. Every presence, value, conjunction and cooccur landmark must carry at least one accept self-test tagged "probe": "precision" — a more specific or less obvious correct rendering than the default one. Writing the lindenbark case is what would have caught B11 before the run. cooccur is included precisely because B11 was a cooccur; a gate that skipped the type the lesson came from would be decorative (note (u)).

What the gate does not establish, stated because the limit is the point (note (z)). The same author writes the landmark and its probe. The gate catches type errors, not world errors: an auditor who cannot imagine the word cannot write the probe for it. It forces the attempt and records whether it was made. The fix that would actually close this is organisational — landmark spec and accept sets written by different parties — and this project has one author.

Self-test state at freeze: 68 cases pass across 17 landmarks, 16 of them precision probes. The v2 spec still passes its own 52 cases unchanged under v3 (spec_version gating), so the matcher changes are non-regressive.

Input preparation. Each text is reduced to the opening paragraph alone — no volume apparatus, no translator's preface — matched at its endpoints (the sentence about sitting in a birch grove in September; the sentence about the sleep known to hunters).

5. Controls and exclusions

6. Predictions, recorded before the texts were fetched

# prediction status
P1 T3 passes 17/17 confirmed before freeze (positive control)
P2 At least one of T1/T2 flags on an odd colour compound — C11 pale-lilac trunk, C12 grey-green metallic, or C5 overripe grapes. These are the phrases whose oddity is the point, and a period translator smooths them. blind
P3 Neither text errs on the date (C1), the order (C3) or the negated presence (C9). Dates, sequences and negations survive translation; this is S022's P5 and S023's P3, both of which held. blind
P4 Neither text is clean — at least one flag each. This prediction FAILED in S023 (both texts came through «Бежин луг» at 16/16) and is repeated verbatim rather than weakened. A prediction that failed once and is put back unchanged is a test; one quietly softened is a hedge. blind
P5 The two texts do not flag on the same landmark set. If they did, the flags would more likely be measuring the landmark list than the translations. blind
P6 The flag counts will not favour Garnett — i.e. Hapgood's flags ≤ Garnett's. This runs against the only external evidence the project has (the blog ranks Garnett above Hapgood on this passage) and with the two unweighed straws from S023. Directional and falsifiable; if Garnett comes out cleaner, the blog's direction gains its first independent corroboration and case 1 closes harder. blind

7. Failure criteria, and what this cannot show