Repository path: workshop/experiments/E-20260725-published-audit/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260725-published-audit |
| status | frozen |
| created | 2026-07-25 |
| updated | 2026-07-25 |
| senses | accuracy |
| internal-judgment-only | true |
| links | wiki/decisions/resolved/D-20260725-06-heldout-arm-operationalisation.md, wiki/base/anchors/A-chekhov-pari/A-chekhov-pari.md, wiki/findings/results/RS-20260725-heldout-pair.md, workshop/translations/posle-teatra/R04-v1/translation.md, tools/audit_landmarks.py, workshop/experiments/E-20260725-published-audit/landmarks.json |
Frozen design — the factual-damage audit, run on a second Chekhov cell
Frozen 2026-07-25 (S022), before either published English text was fetched. The freeze is the whole point of the design and is verifiable in git: T-posle-teatra-R04-v1 was committed at 992c1cc, and this design and its landmark spec are committed before the first curl of Gutenberg #1732 or #55283.
1. Why this runs
D-20260725-06 was ratified this session to Q-A, and its implementation carries two conditions binding on every future held-out candidate pair. The second is a mandatory pre-run factual-damage audit against the source: known material damage disqualifies a pair even if comparative reception evidence exists. That instrument did not exist; this builds it and runs it on the obvious next cell.
It also discharges a named revision trigger on RS-20260725-heldout-pair:
Either of the two remaining Chekhov overlap stories shows K&M clean and Garnett erring. That would reverse point 4's direction and make the error rate look like noise rather than a property of the 1915 text.
2. Question
On a second story translated by the same two translators, does the pattern found on «Пари» hold — Koteliansky & Murry 1915 carrying plain factual error where Garnett 1920 is clean — or was that one text?
This is a question about published human translation as a Tier 1 precedent anchor (charter §4), not about which translator is better. Three errors in one story was never evidence about a translator; the project said so at the time. A second independent cell is the cheapest thing that moves it either way.
3. Materials
| label | text | provenance |
|---|---|---|
| S | Chekhov, «После театра», 1892 | ru.wikisource / ФЭБ, from PSS vol. 8 pp. 32–34 (the author's Marks-edition text). SHA-256 8ff3d1e…f9e594…, stored at workshop/translations/posle-teatra/R04-v1/source-ru.txt |
| T1 | Koteliansky & Murry, 1915 | Project Gutenberg #55283, the volume already identified in A-chekhov-pari — not yet fetched at freeze time |
| T2 | Constance Garnett, 1920 | Project Gutenberg #1732, same — not yet fetched at freeze time |
| T3 | the lead's T-posle-teatra-R04-v1 |
frozen at 992c1cc before T1/T2 were fetched |
4. The instrument
tools/audit_landmarks.py against landmarks.json — 16 landmark facts read off the Russian, with accept patterns frozen verbatim in the spec file.
The design principle, stated because it is what makes this an audit and not a hunt. Accept patterns are deliberately generous: any correct rendering must pass whatever vocabulary it chooses (L9 accepts rogue/rascal/scoundrel/swindler/knave/scamp/cheat; L12 accepts skittles/ninepins/bowls). The instrument therefore cannot flag a translator for word choice — only for a fact that is absent or different. A FLAG is not a finding. It is a referral to hand adjudication against the Russian, and every adjudication is recorded verbatim with the source quoted.
The 16 landmarks span kinds of fact that damage differently: a number (L1 sixteen, L4 two o'clock), proper names (L2 Onegin, L6 Maxim, L11 Gorbiki), role assignments that can swap (L3a Gorny/officer, L3b Gruzdev/student), a species (L7 raven not crow), objects and substances (L5 rubber ball, L13 wormwood, L14 icon), an action's content (L9 the raven's insult, L10 it looks round first), a pair of games (L12), and a count (L15 "Lord!" three times).
Input preparation, specified. Each text is reduced to the story body alone — no front matter, no translator's log, no volume apparatus — because L15 counts within the final 300 characters and would otherwise measure whatever follows the story. (This rule was added after the spec was written and before any published text was fetched: the first self-test run flagged L15 on the lead's own file, whose log follows the story. The spec itself is unchanged; only the input rule. Recorded here rather than folded in silently — NEXT.md note (r).)
5. Controls
- Positive control, already run and passed: the lead's own translation scores 16/16. T3 is source-derived by construction — it was written from the Russian in this session — so an instrument that flagged it would be measuring synonym choice rather than fact. It does not. This is the limb
E-20260725-heldout-pairdid not have until its critic demanded one (NEXT.mdnote (p)): a case the instrument should pass. - Negative control:
A-chekhov-pariitself. The three K&M errors on «Пари» are of exactly the kinds this list contains — a number ("five minutes" for five hours), a proper noun ("the New Testament" for the Gospel), a time ("midnight" for noon). The instrument is not run on «Пари» here; the point is that the landmark kinds were chosen to be the kinds that were previously found to fail, which is a known and declared bias in the instrument's sensitivity.
6. Predictions, recorded before the texts were fetched
| # | prediction | status |
|---|---|---|
| P1 | T3 passes 16/16 | already confirmed (run before freeze; the positive control) |
| P2 | T2 (Garnett) passes ≥ 15/16, with any flag adjudicating to a legitimate rendering rather than an error | blind |
| P3 | T1 (K&M) carries at least one factually altered landmark | blind — this is the directional prediction, and it is the one that can fail |
| P4 | Both published texts flag on L14 (образ → "image" rather than "icon") | blind — the trap I expect a period translator to fall into, predicted for both, so it is not a prediction against either |
| P5 | Neither published text errs on L1, L2, L6 or L11 (age, Onegin, Maxim, Gorbiki) — proper names and round numbers survive translation | blind |
7. Failure criteria, and what this cannot show
- If T1 passes 16/16, P3 fails and the finding is a null: the «Пари» error rate does not reproduce, and
RS-20260725-heldout-pairpoint 4 must be narrowed to one story. That result is written up exactly as readily as the other one. - Two cells are not a rate. Whatever comes out, this is n = 2 stories from one volume-pair by one translator-pair. It cannot support any claim about published translation in general, about either translator's body of work, or about 1915 versus 1920 as periods.
- The instrument sees only what is on the list. Sixteen landmarks over ~1,100 words leave most of the story unaudited; a translation can be flatly wrong somewhere the list does not look. It is a filter, which is what the ratification asked for, not a quality measurement.
- No jury, no senses scored, no evidential weight.
accuracyappears in this page's front matter because that is the sense the landmarks bear on; nothing here is a jury verdict and Tier D remains unpassed.