Repository path: workshop/experiments/E-20260804c-peer-record/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260804c-peer-record |
| status | frozen |
| created | 2026-08-04 |
| updated | 2026-08-04 |
| links | wiki/arms/ARM-tierP.md, wiki/decisions/resolved/D-20260725-07-athenaeum-1906-condition-ii.md, wiki/decisions/resolved/D-20260725-06-heldout-arm-operationalisation.md, workshop/translations/pevtsy/R04-v1/translation.md, wiki/goodness-senses.md, config/models.md, PROJECT.md |
| senses | accuracy, naturalness, affect |
| provisional | true |
E-20260804c — does the 1904 record's split verdict reproduce on the prose?
Frozen 2026-08-04 (S103), before any dispatch. Charter §5, Tier P — peer discrimination. The lead's translation and its log were frozen in an earlier commit; that ordering is provable in git and is what makes §7's registered prediction P4 a prediction.
1. The question
In 1904 an unsigned reviewer in The Nation read A Nobleman's Nest and three sketches of
Memoirs of a Sportsman in the Russian and compared Constance Garnett's rendering with Isabel
Hapgood's, printing parallel columns of errors from each. In 1906 The Athenaeum did the same
thing more briefly. Both split the pair by dimension: Hapgood the more accurate on this cycle,
Garnett the better English, neither better overall. D-20260725-07 ratified that record as
comparative reception evidence for the Memoirs cycle.
Do those per-sense directions reproduce when readers who have never been told whose prose they are reading compare the two renderings against the Russian, passage by passage — and if they do, is the reproduction separable from recognition of the canonical text?
The second half is not a formality. Garnett is the canonical English Turgenev and Hapgood is not,
and the lead has now been measured reproducing Garnett at 12 contiguous tokens on a 315-word
blind rendering of this very story while reproducing Hapgood at 9 with no shared 12-gram at all
(materials/gate-result.json; the translation artifact's contamination block). If a reading
instrument carries the canonical text in its memory and not the rival, a preference for the
canonical text is not evidence about the prose.
2. What the record actually says, and what it does not
Verbatim from D-20260725-07 and its two excerpt files. Scope: the Memoirs of a Sportsman
cycle — the ratifying vote restricted the evidence to it, and «Певцы» is a sketch of that cycle.
| the record's dimension | direction | the words |
|---|---|---|
| accuracy, on this cycle | Hapgood | "decidedly the more accurate"; "a slight advantage over her predecessor" |
| English style / idiom | Garnett | Hapgood "translate[s] too literally, foregoing English idioms"; Athenaeum: Garnett's "version is in elegant English, and perhaps in this respect superior" |
| literary reach | Garnett | Garnett "seems to rise more often than Miss Hapgood to the possibilities of her subject" |
| apparatus, notes | Hapgood | notes "well done, and not overdone" |
| overall | level | "of essentially the same character"; "neither produces work of marked literary excellence" |
Three things this design refuses to take from the record.
- The 40:12 count is not about these materials. It is scoped to A Nobleman's Nest; on Memoirs the reviewer says only that the evidence points "in the same direction, though by no means so emphatic". The English-style direction is used; the magnitude is not.
- Apparatus is excluded as a testable dimension. Its ground is Hapgood's footnotes, which are
not part of anyone's rendering of Turgenev — and they are stripped from the payload (§4),
because they name her. So
cultural-mediationis not tested here and no claim is made about it. - "Level overall" is not a prediction of this design. The record's own overall judgment is
parity, which is what
D-20260725-06needs and what Tier D would test. This experiment tests the dimensional split, which is the opposite structure and is the only part with a direction.
3. Mapping the record onto live senses — declared, with one flagged interpretation
accuracy— direct. Current wording, including the compelled-specification rules ofD-20260727-08.naturalness— direct ("English idiom", "elegant English"). The sense requires an evaluation to name a register anchor; this one names period-idiomatic, since both published arms are Victorian/Edwardian.affect— exploratory, and NOT part of the primary criterion.literary-qualitywas retired 2026-08-01 (D-20260801-11) and its work assigned tonaturalness(threshold competence) andaffect(reader experience). The record's "rise to the possibilities of her subject" is a reader-experience claim, soaffectis the nearest live sense — but that mapping is a lead interpretation of a nineteenth-century phrase and is flagged as one.internal-judgment-only. Reported, never counted in the primary.
4. Materials
Source. Turgenev, «Певцы» (1850), from «Записки охотника», read in Russian. 82 paragraphs,
5,365 words, at materials/source-ru-full.txt.
Locus rule, fixed before the lead translated and unchanged since: the five longest paragraphs of the Russian, plus the story's longest continuous run of dialogue. → L1 ¶68 (518 w, Yakov sings), L2 ¶2 (413, the publican), L3 ¶72 (393, the drunken aftermath), L4 ¶16 (383, the room), L5 ¶49 (383, the Wild Master), L6 ¶26–43 (347, drawing lots). The next longest paragraph is ¶47 at 357 words, so "the five longest" is unambiguous. 2,437 Russian words, 45% of the story.
Arms, three.
| arm | text | provenance |
|---|---|---|
garnett |
Constance Garnett, A Sportsman's Sketches vol. 2, Heinemann 1897 | PG 8744 |
hapgood |
Isabel F. Hapgood, Memoirs of a Sportsman vol. II, Scribner 1903 | two independent archive.org scans, required to agree |
lead |
T-pevtsy-R04-v1, frozen in an earlier commit |
this repository |
Extraction is anchor-exact, not aligned by heuristic. materials/extract.py cuts every span
between a verbatim start and end anchor and asserts each anchor is unique in its arm. The first
DP-alignment attempt was discarded when its boundaries proved wrong at five of six loci; nothing
from it survives.
Note (aa) fired and was worked, not waived. The two Hapgood scans disagreed at eight places
across the six loci on the first run. Every one is a single-token OCR slip and neither scan is
uniformly right — scan A carries Yakoif, CasUlots, prftynny, Ovsydnikoff; scan B carries
longwings, ever}7, YakofY's, factory -hand, greatl3*. All eight adjudications are listed
in SCAN_REPAIRS in extract.py, and the script now asserts word-for-word agreement between
the scans after repair. It does.
Four Hapgood footnotes fall inside the loci and are removed, listed in FOOTNOTES: the Table
of Ranks note, the pritynny note, the hawks squawk note, and the cross-reference to
"Freeholder Ovsyanikoff". They are her apparatus, they end "— TRANSLATOR.", and leaving them in
would break the blind outright. Their two reference markers are removed with them. Removing them is
the only content dropped from any arm, and it is the reason apparatus is not a tested dimension
(§2.3). The pritynny note is set inside a hyphenated word, so particu- larly is rejoined.
Paragraph breaks are flattened to one block per locus, identically for all three arms, and this
is pre-registered rather than incidental. The reason is that Hapgood's paragraphing is recoverable
only from OCR of a scan and is therefore not a property of her translation that this project can
read; comparing paragraphing across arms would be comparing one translator against a scanner. It
also removes a formatting cue: the lead's L1 is six paragraphs against the Russian's one (log D2).
Declared cost: paragraphing is removed from what is judged, which bears on affect.
5. Procedure
Seats: P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5 — three
non-Anthropic labs (config/models.md). The lead never judges (charter §5), and one arm is the
lead's.
Stage 1 — graded ranking, source present. 3 seats × 2 orderings = 6 bodies. Each body sees all six loci: the Russian, then the three arms as A/B/C under a per-locus sha256-derived permutation, so no letter carries meaning across loci. For each locus the seat returns a full ranking of the three arms on each of the three senses, 1 = best. 6 loci × 3 senses = 18 lines. Ranking rather than a pick is what preserves the pairwise Garnett-vs-Hapgood test at a 1/2 null while still placing the lead.
Stage 2 — recognition probe, source absent. 3 seats × 1 = 3 bodies. One whole-call
permutation. The seat is asked to name the translator behind each letter, or UNKNOWN. This
measures the canonicity confound on the instrument that produced stage 1.
Stage 3 — the per-seat memorisation probe. This is the control the S014 Tier P run did not
have. 3 seats × 1 = 3 bodies. Each seat is given ¶4 of the Russian — a paragraph outside
every locus — and asked to translate it, with no English shown. Each output goes through
tools/dependence_check.py against Garnett and Hapgood, giving a per-seat
Δ = (longest run vs Garnett) − (longest run vs Hapgood) and the same for shared 12-grams.
The lead's own value is already on record: Δrun = +3, Δ12gram = +1.
Judgment is not parallelised. Every prompt is written to runs/ before dispatch; raw bytes hit
disk before any parse; max_tokens is set from the payload, not from an assumed answer length
(note (abc)), and reasoning effort is low on the first dispatch to every seat (note (b)).
6. Analysis, and the criterion
Pooled over 6 loci × 3 seats × 2 orderings = 36 pairwise Garnett-vs-Hapgood judgments per sense, read off the rankings.
PRIMARY — the record is reproduced iff BOTH hold:
- (a) Hapgood takes a strict majority of the 36 on
accuracy; and - (b) Garnett takes a strict majority of the 36 on
naturalness.
Exact null probabilities by enumeration under an unbiased-coin null, per sense and per seat.
FAILURE CRITERIA, registered.
- F1 — uniform winner. If one translator takes the majority on both
accuracyandnaturalness, the record is NOT reproduced, whatever the margins. This is the criterion the canonicity confound trips, and it is the reason the primary is a dissociation rather than a ranking. - F2 — order instability. If the two orderings disagree in sign on either primary sense, that sense is reported as unresolved and does not support the primary.
- F3 — seat collapse. Fewer than 6 accepted stage-1 bodies, or any body with fewer than 18 answer lines, is a seat failure; the runner enforces the line count.
- F4 — recognition dominance. If stage 2 identifies Garnett at above 4 of 6 seat-loci and the primary reproduces, the primary is reported confounded, not reproduced. Recognition is not a footnote to this result; it is a condition on it.
- F5 — degenerate rankings. If a seat returns the same ranking at every locus for a sense, that seat contributes no information on that sense and is reported separately.
7. Predictions, registered before dispatch
- P1 — the dissociation reproduces: Hapgood majority on
accuracy, Garnett majority onnaturalness. - P2 — at least one grading seat shows a positive Δ on stage 3, i.e. reproduces Garnett more closely than Hapgood, matching the lead's +3.
- P3 — stage 2 recognises Garnett more often than Hapgood.
- P4 — the wire between the limbs, and it is a real prediction. The lead's translator's log,
frozen before any English was read, records at D4 that the whole of L6 is in the familiar
second person and that the lead could see no way to mark it in English — noting that this is
exactly the site at which the 1904 Nation attacks Hapgood, for solving it with "thou".
P4: L6 will be in the top two of the six loci by Garnett's
naturalnessmargin over Hapgood. Null 1/3; the rank of L6 among six is reported whatever it is. - P5 — the lead arm will be ranked last on
naturalnessmore often than onaccuracy, because its own log's D1 declares an unmarked contemporary register against a period-idiomatic anchor.
A no-information benchmark is computed for P4 and reported beside it: the rank of L6 by locus length and by raw arm-length difference, the two properties available without reading anything. S094's log-prediction result was overturned by exactly this check, and it is registered here rather than added afterwards.
8. What this cannot establish
- Nothing is calibrated by this run. Tier P is a certification, not a gate (charter §5); a failure is data. Tier D remains NOT PASSED and no jury verdict here carries evidential weight.
- The
owedobligation is respected by avoidance.D-20260802-13condition 4 requires Tier D's faileddrop(naturalness)specificity test to be re-run under the revised wording before any revised-sense verdict is called calibrated or compared with prior Tier D results. This design uses the revisednaturalnesswording, calls nothing calibrated, and compares nothing with any Tier D figure. The obligation is scheduled asARM-tierPstep 3. - One story, one pair, one language. The record is about two whole editions; this tests 45% of one sketch of one of them.
- Period is uncontrolled between the published pair and the lead, as it was at S102. The lead arm is secondary throughout for this reason and for its measured Garnett-dependence.
- The lead arm is not an independent third opinion of anything — note (bhb) and the measured
gate. Its
accuracyandnaturalnessresults against Garnett must not be read as independent.
9. Pre-flight budget estimate (written before any dispatch)
Worst case built from max_tokens, not from an assumed answer length — note (abc).
| stage | bodies | input (est. tok) | max_tokens |
worst case |
|---|---|---|---|---|
| critic | 1 | ~9,000 | 12,000 | $0.21 |
| 1 — graded ranking | 6 | ~28,000 | 10,000 | $0.79 |
| 2 — recognition | 3 | ~22,000 | 6,000 | $0.26 |
| 3 — memorisation probe | 3 | ~600 | 4,000 | $0.09 |
Declared worst case: $1.35. Today's headroom before this session is $4.391673436
(config/budget.md, UTC 2026-08-04, after S101 and S102). The run fits with room, and routing can
move a per-call price by ~4× (the S022 caution in config/models.md), which the worst case above
absorbs at the stage level but not four-fold; a stage that overruns is stopped, not continued.
10. Amendments A1–A11 — from the pre-run critic pass, applied before any grading call
NEEDS-AMENDMENT, 6 BLOCKING + 5 ADVISORY, all accepted (one in a weaker form). Full record and
dispositions: critic.md. Where an amendment contradicts §6 or §7 above, the amendment governs;
the superseded text is left in place because a design that quietly rewrites itself is not frozen.
- A1 (F1) — the primary is cluster-respecting. §6's "strict majority of 36" treated 36 judgments
as independent when they are 6 loci × 3 seats × 2 orderings of the same bodies. The primary is
now: for each sense, pool each seat's 12 pairwise judgments into that seat's direction; the
record is reproduced iff all three seats independently give Hapgood
accuracyand all three independently give Garnettnaturalness. Null per sense (1/2)³ = 0.125; joint 0.0156. The 36-count and its naive null are still reported, labelled an upper bound on evidence, not the test. A seat that ties on a sense counts against reproduction for that sense. - A2 (F2) — F4 gets a real denominator. The recognition criterion is now: if 2 or more of the
3 stage-2 seats correctly name Garnett, the primary is reported
confounded. (Stage 2 yields three identifications of the Garnett letter, one per seat, not "6 seat-loci".) - A3 (F3) — a null probe never clears the confound. Registered: if stage 2 names nobody and
stage 3 returns Δ ≈ 0 for every seat, the result is reported as "confound status undetermined"
and never as "separable from recognition". A probe can only raise the confound, never discharge
it. The critic's stronger remedy — a seeded canary establishing positive detection power — is
declined for this run, because it needs materials this design does not have and would change
what stage 3 measures; it is recorded as an un-taken check and is the first candidate for step 2
of
ARM-tierP. - A4 (F4) — the accuracy construct is narrowed, not defended. Seats may grade
accuracyby surface correspondence to the Russian, which rewards literalness — and Hapgood's documented characteristic is literalness ("translate[s] too literally, foregoing English idioms"). So a Hapgood win onaccuracyis consistent with a literalness heuristic and with the 1904 reviewer's accuracy claim, and this design cannot separate them. Every reported sentence says so. Registered diagnostic, computed whatever the outcome: per locus, each published arm's word-count ratio to the Russian and its shared-7-gram count with the lead's independent rendering, as two crude literalness indices; if Hapgood'saccuracywins track her literalness index across loci, the confound is showing and is reported as showing. (S015 found all four non-Anthropic panel seats passing a six-item Russian competence screen at ≥5/6; that is weak prior evidence about reading Russian and is not validation of accuracy grading, and is not used as such.) - A5 (F5) — the circularity on the naturalness limb, named. These seats' standard of idiomatic
literary English was constituted from a corpus in which Garnett's translations are
over-represented — she is the canonical English Turgenev, Chekhov, Dostoevsky and Tolstoy at once.
Measuring "whose English is better" with such an instrument is partly measuring Garnett against a
norm she helped set. This is not the same as recognition and is not discharged by a null
stage-2: recognition is sufficient evidence of the confound and not necessary evidence of it.
Added to §8 as a standing limit on any
naturalnessfinding here, in either direction. - A6 (F6) — the lead is removed from editing its rival's text.
materials/adjudicate.pydecides every OCR divergence by frequency of the minimal differing token across both whole scanned volumes (161,593 tokens), writesocr-decisions.json, andextract.pyapplies that file and nothing else. Two divergences the rule did not settle are handled and declared in that script's header: it chose against the lead onfactory-hand(a line-break hyphen, so hyphen spacing is now normalised identically in all three arms) and tied onOvsydnikoff/Ovsyanikoff, which sits inside a stripped footnote and never reaches a grader. - A7 (F7) — P5 is demoted from a prediction to a manipulation check on the register anchor. It is true by construction and will not be reported as evidence.
- A8 (F8) — P2 and P3 are hardened or labelled. P2 now requires all three seats to show Δrun ≥ +1 (Garnett-ward), not one. P3 is a sanity check, not evidence.
- A9 (F9) — a second no-information benchmark for P4. Besides locus length, the loci are ranked by dialogue fraction — share of Russian words inside quotation marks — and P4 is reported against that ranking too. L6 is the only extended-dialogue locus, so if L6 tops the naturalness margin, the dialogue benchmark predicts it without anyone reading anything, and P4's content collapses to that. Reported either way.
- A10 (F10) — F2 replaced. Instead of a sign rule on pooled orderings: report the order-effect size per sense (Garnett's share under o0 minus under o1), and call a sense unresolved only if two or more seats individually flip sign between orderings.
- A11 (F11) — flattening's asymmetry. Paragraph flattening removes a registered structural property of the lead arm (log D2: six English paragraphs against the Russian's one) as well as Hapgood's unreadable OCR paragraphing. Noted in §8; the caveat is attached to L1 in the per-locus report.