Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260725-gnezdo-audit/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260725-gnezdo-audit
statusfrozen
created2026-07-25
updated2026-07-25
sensesaccuracy
internal-judgment-onlytrue
linksworkshop/translations/dvoryanskoe-gnezdo/R04-v1/translation.md, workshop/experiments/E-20260725-gnezdo-audit/landmarks.json, workshop/experiments/E-20260725-gnezdo-audit/verification.md, wiki/findings/results/RS-20260725-gnezdo-audit.md, wiki/base/searches/SR-20260725b-nation-1904-excerpt.txt, wiki/decisions/resolved/D-20260725-07-athenaeum-1906-condition-ii.md, tools/audit_landmarks.py

E-20260725-gnezdo-audit — does the 1904 Nation's parity claim about Дворянское гнездо survive an independent factual audit?

The question, and the wire to the translation limb

The 1904 Nation joint review makes its strongest parity claim about this novel and not about the cycle the project has audited:

"In 'A Nobleman's Nest' and 'A House of Gentlefolk' we have found almost exactly the same number of errors."

The project has audited Записки охотника twice and this novel never. The study limb of this session established what that review actually says (SR-20260725b-nation-1904-excerpt.txt, rewritten from page geometry and page images). This limb tests its central factual claim on the work it is about, on a passage the reviewer did not quote.

The wire in one sentence: the study limb recovers an external critic's counted claim about this novel, and the translation limb tests that claim by counting independently, 122 years later, from the same Russian.

Materials

role text provenance
source Turgenev, ch. I paras 1–4, 442 words ru.wikisource + az.lib.ru, word-by-word cross-check, one az typo adjudicated; SHA-256 d5b0ff5…
T1 Constance Garnett, A House of Gentlefolk (1894), 595 words Project Gutenberg #5721
T2 Isabel Hapgood, A Nobleman's Nest (1903), vol. IV, 611–612 words three independent archive.org scans: novelsstoriesofi04turg, novelsstoriesofi04turguoft, noblemansnest00turg
T3 the lead, T-dvoryanskoe-gnezdo-R04-v1, 584 words frozen at commit 2a0bcaa, before any of T1/T2 was fetched

All public domain (Turgenev d. 1883, Garnett d. 1946, Hapgood d. 1928).

Passage selection. Chapter I paras 1–4, stopping where the dialogue starts. Chosen to lie clear of every site the 1904 review quotes: the review's ~20 parallel-column extracts are cited to Hapgood vol. iv pp. 8, 41, 56, 67, 69, 71, 83, 109, 115, 141, 144, 145, 153–156, 203, 228, 249, 257, 277, 287, and this passage precedes p. 8.

Three scans, not two. NEXT.md note (aa). The three Hapgood scans agree on all 615 tokens apart from OCR noise, and the noise is not decorative: scan A renders «желчный и упрямый» as s[)lenetic and stub])orn, which flagged landmark G5 on scan A alone. Auditing one scan would have produced a false finding against Hapgood.

Procedure

  1. Translate the passage blind from the Russian; freeze the translation and its log (commit 2a0bcaa).
  2. Fix the two instrument defects S024 found by hand; ship spec_version: 4 with a refusal regression and a positive control.
  3. Derive landmarks from the Russian alone; pass the self-test gate; freeze (commit 67de530).
  4. Only then fetch T1 and T2.
  5. Run the frozen instrument on all five texts (T1, three T2 scans, T3).
  6. Adjudicate every flag by hand against the Russian. A flag is a referral, never a finding.

Predictions, recorded before the run

  1. If the Nation's claim holds, T1 and T2 should carry a similar small number of upheld factual deviations.
  2. Most flags will be instrument defects, not translation errors — the prior on this project's landmark specs is that generous accept sets are still bounded ones (notes (y), (ee)).
  3. T3 should be clean, since its translator wrote the landmarks. A T3 flag is evidence about the instrument, not about T3.

Failure criteria

Gate accounting — what was kept and what was NOT

Stated plainly, because the charter names four gates and this run kept three.

gate kept?
frozen written design before the run NO — see below
independent pre-run critic pass NO
run with raw outputs preserved yes — runs/audit-full.out, all five texts, SHA-256 per text
post-run verification recomputing every reported number yes — verification.md

What was actually frozen before the run, and provable in git: the translation and its log (2a0bcaa), and the landmark spec — the operative instrument — with its 61 self-tests and 16 precision probes (67de530). Both precede the commit that first fetched T1/T2. That is the load-bearing freeze and NEXT.md note (f) was kept.

What was not: this prose page was written after the run, and no independent critic read the design before it executed. The consequence is real and is not argued away — a critic pass is what caught a fatal flaw in E-20260725-heldout-pair before it ran, and this design did not get one. The result is reported at correspondingly lower confidence, and the two upheld findings below should be treated as candidates for an independent check rather than as settled. Both are single-token facts checkable in under a minute by anyone with the three files, which is the only mitigation on offer.