Repository path: wiki/findings/results/RS-20260725-gnezdo-audit.md · rendered 2026-09-09
Page metadata (front matter)
| type | result |
|---|---|
| id | RS-20260725-gnezdo-audit |
| status | frozen |
| created | 2026-07-25 |
| updated | 2026-07-25 |
| senses | accuracy |
| internal-judgment-only | true |
| provisional | true |
| links | workshop/experiments/E-20260725-gnezdo-audit/design.md, workshop/experiments/E-20260725-gnezdo-audit/verification.md, workshop/translations/dvoryanskoe-gnezdo/R04-v1/translation.md, wiki/base/searches/SR-20260725b-nation-1904-excerpt.txt, wiki/decisions/resolved/D-20260725-07-athenaeum-1906-condition-ii.md, config/models.md |
RS-20260725-gnezdo-audit — an 1904 critic's count, tested 122 years later, plus the contamination measurement nobody asked for
The headline
Two results, and the unplanned one is the more important.
1. The Nation's parity claim about Дворянское гнездо survives an independent audit — weakly, as it must at this sample size. On 18 landmarked facts in the novel's opening 442 Russian words, hand-adjudicated against the source: Garnett 1 minor factual deviation, Hapgood 1. The 1904 reviewer, who read the whole novel against the Russian, wrote that he found "almost exactly the same number of errors" in the two. One-versus-one is consistent with that and is the least discriminating outcome available; it is not confirmation.
- Garnett: «в пятидесяти верстах» → "about forty miles". 50 versts = 33.1 miles; the conversion overstates by ~21%.
- Hapgood: «(дело происходило в 1842 году)» → omitted. The string
1842occurs 0 times in each of three independent scans of the whole volume; not relocated, not spelled out.
2. The lead's blind translation is measurably not independent of the published ones. Shared 7-grams, proper nouns excluded: Garnett~Hapgood (two independent human translators) 18; lead~Garnett 29; lead~Hapgood 37.
DOWNGRADED IN PLACE, 2026-07-25 (S026), by
RS-20260725c-contamination-sweep§"The downgrade of S025". This paragraph originally read "by a factor of two … the lead shares 2.06× as many seven-word strings with Hapgood … and the ratio grows with n." Three of those claims do not survive re-measurement under a frozen, named metric bracketed between its two parameter-free extremes: the denominator is unstable by ~4.5× across passages for the same translator pair; the multiplier is metric-dependent by ~25% (1.65 to 2.04 for this very cell); and the growth with n is flat under the stricter proper-noun rule. The sweep also corrects the reading: across six cells the lead is symmetrically close to both published texts in five, and is the most central text in all six, so the general phenomenon is centrality, not recall of Hapgood specifically. What survives here: this text overlaps the published translations more than they overlap each other, by a factor between about 1.6 and 2.0, and the project cannot state a stable multiplier. The operational conclusion — a lead translation is not a neutral control — is unaffected and now rests on six cells rather than one.
Why this was worth doing
D-20260725-07, ratified this session, rests on a review whose strongest parity claim is about this novel — the one work in the Turgenev pair the project had never audited. The study limb established what the review actually says; this limb tested its central factual claim on a passage the reviewer did not quote. Freeze order provable in git: translation 2a0bcaa, landmark spec 67de530, both before either published text was fetched.
What the instrument did, and how badly
20 raw flags across five texts. 2 upheld. Four landmarks fired on texts that are demonstrably correct, including the lead's own:
- G15 / G16 — both
orderlandmarks over comparisons ("more sentimental than kind", "less inherited than acquired"). All three translations correct; the landmarks broken. - G6 — Garnett's "born in narrow circumstances" and Hapgood's "born in a poverty-stricken class" are both correct and neither was in the accept set. Note (y), third occurrence.
- G9 — Garnett renders «в чёрном теле» twice ("a scanty allowance"… "very shabbily"); the accept set held neither word.
- G5 — flagged Hapgood on one scan only, where the OCR reads
s[)lenetic. Note (aa) earning its keep: a single-scan audit would have recorded a false factual finding.
The finding rests on hand adjudication, not on the tool. A 2-in-20 precision rate makes this a referral mechanism.
The new instrument lesson, and why the gate could not have caught it
G15 and G16 failed because an order landmark takes the first occurrence of each item across the whole 600-word passage, while both facts are comparisons local to one clause. G16's pivot was \b(as|than)\b.
Every self-test fixture in this project is a single sentence. In a single sentence the first occurrence of an item is the one under test — so a landmark that only fails in the presence of earlier distractors cannot fail its own gate. The gate tests matchers on material structurally unlike the material they run on.
This is distinct from note (z) ("the author cannot imagine the word"): here the author wrote a correct matcher and a correct fixture, and the fixture was the wrong shape. The fix — order landmarks must declare a scope window, and their self-tests must include a long fixture with a distractor ahead of the site — is deliberately not retrofitted here, because changing a matcher after seeing which texts it flagged is the failure this project keeps catching in itself.
Two defects S024 found were fixed before this run, as spec_version: 4 load-time refusals, with a regression that runs the offending landmarks verbatim from the frozen v3 spec: C3 refused by rule 7 (unbounded lexical alternation), C8 refused by rule 8 (a conjunction whose referent is elided in the source). The rewrite rule 8 demands rejects Garnett's actual "one young leaf, all red or golden" — the sentence v3 passed — while accepting a correct English anaphor.
The contamination measurement
Two exact coincidences prompted it — lead/Hapgood "and his aunt on short commons until", lead/Garnett "the necessity of pushing his own way in the world" — neither a forced rendering.
Shared n-grams, proper nouns excluded:
| pair | n=4 | n=5 | n=6 | n=7 |
|---|---|---|---|---|
| Garnett ~ Hapgood (baseline) | 63 | 42 | 26 | 18 |
| Garnett ~ lead | 85 (1.35×) | 60 (1.43×) | 40 (1.54×) | 29 (1.61×) |
| Hapgood ~ lead | 108 (1.71×) | 75 (1.79×) | 50 (1.92×) | 37 (2.06×) |
The ratios in this table are downgraded (see above). Recomputed under the frozen metric of
E-20260725c-contamination-sweep:CI(7)is 2.037 with no proper-noun exclusion and 1.647 with aggressive exclusion, and the growth across n is present only in the former. The raw counts stand; the multipliers do not. S025 recorded these numbers without recording the rule that produced them, which is note (r) failing on an ad-hoc measurement — the rule had to be inferred by re-measurement.
Among the shared 6-grams: "into the very depths of the azure" — where the translator's log records the lead reasoning its way to "azure" for consistency with an earlier translation, in a passage where Hapgood had already written those six words.
This does not touch result 1, which compares the two published texts against the Russian. It does turn contamination: suspected into a measured contamination: high for T-dvoryanskoe-gnezdo-R04-v1, and it raises a live problem: lead translations are used as controls (T-bezhin-lug-R04-v1, T-svidanie-R04-v1), and a control measurably closer to one arm than the arms are to each other is not a neutral third point. The measurement costs three files and twenty lines and should be run on every stored lead translation before any is used as a control again.
Confound, stated: the baseline pair are two Victorians nine years apart, and any two renderings of the same 442 words overlap somewhat. The argument rests on the direction and on the growth with n, not on the absolute counts.
What this does not establish
- It does not rank the translators. The design's failure criteria say the instrument disqualifies rather than ranks, and 1-vs-1 on 18 facts would not rank them if it did.
- It does not reopen
D-20260725-07. That ratification excluded A Nobleman's Nest from held-out use on the reviewers' English-style asymmetry, not on factual grounds. Nothing here bears on it, and it is not offered as a reason to revisit it. - It does not calibrate anything. Tier D remains NOT CALIBRATED.
- Gates not kept: no prose design page before the run, and no independent pre-run critic pass — the gate that caught a fatal flaw in
E-20260725-heldout-pair. The two upheld findings are single-token facts checkable in a minute from three files, which is the only mitigation available. They should be treated as candidates for an independent check.