Repository path: workshop/experiments/E-20260725-svidanie-audit/verification.md · rendered 2026-09-09
Page metadata (front matter)
| type | result |
|---|---|
| id | RS-20260725-svidanie-audit |
| status | active |
| created | 2026-07-25 |
| updated | 2026-07-25 |
| senses | accuracy |
| internal-judgment-only | true |
| provisional | true |
| links | workshop/experiments/E-20260725-svidanie-audit/design.md, workshop/experiments/E-20260725-svidanie-audit/landmarks.json, workshop/translations/svidanie/R04-v1/translation.md, wiki/base/searches/SR-20260725-parity-evidence.md, workshop/experiments/E-20260725-turgenev-filter/verification.md, tools/audit_landmarks.py |
Run and verification — «Свидание», Garnett vs Hapgood, condition (iii) on the passage the evidence is about
Headline, and it has two halves that pull in opposite directions.
Raw: Garnett 15/17, Hapgood 17/17, the lead 17/17. Adjudicated: no material factual damage in either published text — both Garnett flags resolve to non-damage, one of them to a defect in my own frozen matcher. But the hand reading found, off the landmark list, one place where Garnett's rendering is grammatically impossible in the source and Hapgood's is right — and the instrument passed it.
Every judgment on this page is the lead's, carries internal-judgment-only, and is provisional per charter §2.4. Nothing here is a jury verdict; no jury was called.
1. Materials as actually run
| label | text | provenance and preparation |
|---|---|---|
| S | «Свидание» opening, 468 Russian words | workshop/translations/svidanie/R04-v1/source-ru.txt, SHA-256 1607c1a3ec7ba4c2c2b8c01d8bbaed9012e0790b9b2a6147540bff04ad37babe |
| T1 | Garnett, "The Tryst", A Sportsman's Sketches vol. 2 | Project Gutenberg #8744 (born-digital transcription). 607 words |
| T2 | Hapgood, "The Tryst", Memoirs of a Sportsman vol. II | archive.org, four independent scans compared. 666 words after preparation |
| T3 | the lead's T-svidanie-R04-v1 |
frozen at commit 5330369, before any of the above was fetched. 639 words |
Note (aa) applied and satisfied threefold. Four scans of Hapgood were fetched and compared word-by-word: novelsstoriesofi02turg (UCLA), novelsstoriesofi02turguoft (Toronto), cu31924088425347 (Cornell), novelsstoriesofi0002isab (Internet Archive). After preparation the first three are 666 words and agree exactly; the fourth diverges at eight points, all OCR noise, including one substantive slip — bue for hue — that would have been invisible without a second scan. UCLA is the base text. The Cornell scan carries autmnn, drizzhng; Toronto carries torn-tit for tom-tit. None of the divergences falls on a landmark site in a way the other scans do not resolve.
Preparation, recorded because it is an intervention on an audited text. The Hapgood scans carry line-break hyphenation flattened into the text (inter- spersed, me- tallic, birch-cop- pice) and one embedded running header (130 THE TRYST). Both were removed mechanically: (\w)- (\w) → \1\2, and the header stripped. Garnett needed no such repair. This is a real asymmetry in the materials — one text is a clean digital transcription, the other is OCR — and it runs in Hapgood's disfavour, since every repair is an opportunity to introduce an error the audit would then attribute to the translator.
2. Raw result
T1-garnett 15/17 flagged: C3, C10
T2-hapgood 17/17 flagged: —
T3-lead 17/17 flagged: —
Full output: runs/audit.out. Instrument: tools/audit_landmarks.py v3, 68 self-test cases passing across 17 landmarks, 16 of them precision probes, all run before any text was read.
3. Adjudication of both flags, by hand, against the Russian
C3 (order) — NOT damage. A defect in my own frozen matcher.
Observed: order settled < aspen < slept (expected aspen < settled < slept).
Garnett's actual sequence is correct:
"Before halting in this birch copse I had been through a wood of tall aspen-trees with my dog. … not stopping to rest in the aspen wood, I made my way to the birch-copse, nestled down under one tree … I fell into that sweet untroubled sleep only known to sportsmen."
The flag comes from the frozen pattern settled|nestled|ensconced|…, which has no word boundary, matching inside "the weather was unsettled" — 2,950 characters before the fact under test, in a sentence about the weather.
This is the converse of the defect v3 was built to fix, and it is worth stating in exactly those terms. NEXT.md note (y) says an accept set punishes precision because it is bounded by the auditor's vocabulary; v3's compound: true relaxes boundaries so a more precise rendering cannot be punished. C3 shows the other edge of the same knife: an unbounded pattern over-matches, and here it produced a flag against exactly one of three texts because of a synonym chosen three thousand characters away from the fact. Hapgood wrote "the weather was inconstant" and the lead wrote "the weather was changeable"; neither contains the string. The flag measured vocabulary, not order.
It is also a second instance of NEXT.md note (z) inside one session: 68 self-tests, 16 of them precision probes, and none of them contained the word "unsettled." The gate catches the errors its author thought to write down.
The frozen spec is not amended. Recorded here, and carried to NEXT.md as an instrument defect for the next session.
C10 (conjunction: titmouse + steel + bell) — NOT damage. A generalisation, with a collateral loss.
Observed: missing ['steel']; found ["'tomtit'", "'bell'"].
Source: «лишь изредка звенел стальным колокольчиком насмешливый голосок синицы» — only now and then the mocking little voice of a titmouse rang out with a steel little bell**.
- Garnett: "at times there rang out the metallic, bell-like sound of the jeering tomtit"
- Hapgood: "only now and then did the jeering little voice of the tom-tit ring out like a tiny steel bell"
Steel is a metal, so "metallic" is not false. The disposition is not damage.
The collateral loss is worth recording and is not what the landmark was testing. Turgenev uses two different words twelve lines apart — «стальным» of the bird's voice and «металлической» of the aspen's foliage — and Garnett renders both as "metallic", collapsing a distinction the source makes. Hapgood keeps them apart ("steel" / "metallic"). This is a shading observation, not a factual finding, and it carries no weight in the result.
4. Adjudicated result
Neither published text carries material factual damage on the seventeen landmarks. Garnett 15/17 raw → 17/17 adjudicated; Hapgood 17/17 raw and adjudicated.
5. What the hand reading found that the instrument did not — and this is the session's real finding
Reading both texts against the Russian line by line, outside the frozen landmark list, turned up one divergence that is not a matter of shading:
«Листва на березах была еще почти вся зелена, хотя заметно побледнела; лишь кое-где стояла одна, молоденькая, вся красная или вся золотая…»
- Garnett: "only here and there stood one young leaf, all red or golden"
- Hapgood: "only here and there stood one, some young tree, all scarlet, or all gold"
- the lead: "only here and there stood a single young one, all red or all gold" (the noun left unsupplied, as in the Russian)
Garnett's reading is not available in the source, on two independent grounds:
- Gender. «одна… молоденькая… вся красная… вся золотая» is feminine singular throughout. Russian лист ("leaf") is masculine. The phrase cannot refer to a leaf. It agrees with берёза, feminine — a young birch. (The other feminine candidate, листва, is a collective mass noun; "one foliage, young" is incoherent.)
- The verb. «стоя́ла» — stood. Trees stand; leaves do not.
The following clause supports the same reading: the sun's rays break through «сквозь частую сетку тонких веток» — the close net of slender branches — to light it up, which is a sapling seen through the surrounding branches.
The frozen instrument passed this. C8 tested for [red|crimson|scarlet|ruddy] and [gold] and [young|little|sapling|small]. Garnett's "one young leaf, all red or golden" supplies all three. The landmark tested the attributes and could not see that they were attached to the wrong referent.
That is the same family as the defect the v2 self-test gate caught in S023 — v1's cooccur asked whether "officer" sat near "Gorny" and so could not detect a role swap. A matcher that tests attributes cannot test what the attributes are attached to. In S023 that failure was caught by a self-test before any data was touched. Here it was caught by reading, after the instrument had already said PASS — which is the outcome the self-test gate is not able to deliver, and the honest measure of what the gate is worth.
6. Predictions, scored
| # | prediction | outcome |
|---|---|---|
| P1 | T3 passes 17/17 | confirmed (pre-freeze positive control) |
| P2 | At least one of T1/T2 flags on an odd colour compound (C11, C12, C5) | FAILED. Garnett: "with its pale-lilac trunk and the greyish-green metallic leaves". Hapgood: "with its pale-lilac trunk, and greyish-green, metallic foliage". Both keep the over-ripe grapes. Neither smoothed the oddities. The prediction was that a period translator normalises; two period translators did not |
| P3 | Neither text errs on the date (C1), the order (C3) or the negated presence (C9) | confirmed on the facts. C3 flagged on Garnett but the flag was a matcher defect, not an order error. Dates, sequences and negations survive translation — now 3/3 sessions |
| P4 | Neither text is clean — at least one flag each | FAILED, for the second consecutive sketch. Hapgood was clean raw; Garnett's two flags both resolve to non-damage. Repeated verbatim after failing in S023, and it failed again |
| P5 | The two texts do not flag on the same landmark set | vacuous. Hapgood produced no flags, so there was no set to compare |
| P6 | Flag counts will not favour Garnett (Hapgood ≤ Garnett) | confirmed on the letter, empty in substance. Raw 0 ≤ 2. But after adjudication both are at zero damage, so there is no accuracy difference for the prediction to have a direction about. Reported this way deliberately: the raw differential is entirely explained by one over-matching regex and one hypernym |
7. What this does and does not do to the Turgenev candidate
SR-20260725-parity-evidence §9's revision trigger 4 is: "A second, independent assessment of Garnett vs Hapgood contradicts the blog's ranking. Would put case 1 back to genuinely unresolved rather than excluded."
The trigger does not fire, and this page will not pretend otherwise. The audit found no accuracy difference between the two texts. An assessment that finds no difference on a dimension the blog was not ranking cannot contradict the ranking. Case 1 stays X2.
What it does add is one thing, stated at the strength the evidence carries:
On both Turgenev sketches the project has now audited, the only hard factual divergence found between Garnett and Hapgood runs against the blog's direction. S023, «Бежин луг»: Garnett turns «угловатую» (angular) into "square"; Hapgood keeps "angular". S024, «Свидание»: Garnett attaches «одна, молоденькая, вся красная» to a leaf, which the source's gender forbids; Hapgood attaches it to a tree. Two sketches, two divergences, both the same way.
This is weak and must be read as weak. n = 2 divergences over 2 passages; the dimension is factual accuracy and the blog ranked prose merit; the shading observations in the same reading go both ways — Garnett keeps «холодно» in "the cold rays of the winter sun" where Hapgood substitutes "sparkling"; Garnett keeps «рдеющим» as "reddening beams" where Hapgood generalises to "glowing"; Garnett's "slovenly leaves" is closer to «неопрятных» than Hapgood's "dirty"; and Garnett's «червонным золотом» → "purplish gold" mis-shades a word Hapgood renders exactly, as "the golden hue of ducats" (червонец is a gold coin). Neither translator is coming out of this as the more accurate one. What is accumulating is only that the hard errors found so far are Garnett's.
8. Instrument defects found this session, for NEXT.md
- Unbounded alternation over-matches (C3, new).
settled|nestled|…matched inside "unsettled". Word boundaries must be mandatory on any lexical alternation used as a positional item, andcompound: trueshould be the only route to boundary relaxation — never an accidental one. Candidate fix: have the tool refuse a spec whose pattern contains a bare lexical alternative without\b, unlesscompound: trueis declared. - An attribute conjunction cannot see a wrong referent (C8, new, and it is the important one). C8 passed "one young leaf, all red or golden" where the source's grammar forbids "leaf". A landmark type is needed that binds an attribute set to a named referent and rejects a rival referent — the
cooccur+rivalsmechanism exists and was simply not used for C8, because at spec-writing time the referent did not look contestable. - The precision-probe gate did what it was built for and no more. 16 probes were written; none of them contained "unsettled" and none of them tested a referent. Note (z) held.
9. Reproduction
python3 tools/audit_landmarks.py --spec workshop/experiments/E-20260725-svidanie-audit/landmarks.json \
T1-garnett=.../runs/T1-garnett.txt \
T2-hapgood=.../runs/T2-hapgood-1903.txt \
T3-lead=.../runs/T3-lead-2026.txt
Raw output runs/audit.out; pre-freeze positive control runs/audit-positive-control.out; instrument regression runs/instrument-regression.out. Cost: $0 — no API call was made.