Repository path: workshop/experiments/E-20260725-heldout-pair/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260725-heldout-pair |
| status | frozen |
| created | 2026-07-25 |
| updated | 2026-07-25 |
| senses | accuracy |
| internal-judgment-only | true |
| links | PROJECT.md, config/models.md, wiki/base/anchors/A-chekhov-pari/A-chekhov-pari.md, workshop/translations/pari/R04-v1/translation.md, wiki/base/anchors/A-beowulf-ingeld/A-beowulf-ingeld.md, workshop/experiments/E-20260725-tierD-ladder/design.md, workshop/experiments/README.md |
Frozen design — qualifying a held-out arm for Tier D
This design is frozen before any measurement of the candidate pair is run. The one place where that is not true is stated in full in §4 (gate G4), and that gate is marked as carrying no independent evidential force.
No API calls are made by the run itself. One API call is budgeted for the independent pre-run critic pass (§9).
0. Why this exists
Charter §5 makes a held-out arm — "an independent same-quality translation, chance expected" — a mandatory control on Tier D, which is the gate on the project's evidential authority. E-20260725-tierD-ladder could not supply one: its independent critic established by measurement that the only candidate pair in the repository was Beowulf, where Gummere 1909 is alliterative verse, Kirtlan 1913 is prose with 17 archaic tokens, and the lead's rendering is plain modern prose with none. That design dropped the arm and pre-committed to claiming no Tier D pass. config/models.md has read NOT CALIBRATED ever since, and NEXT.md names finding a genuine pair as the single blocker.
The blocker is a materials problem, and this design attacks it. It also attacks a second thing, which NEXT.md standing note (l) made the headline lesson of the last session: a control has to be built and measured, not asserted. The Beowulf pair was rejected because somebody measured it. A replacement pair that is merely asserted to be matched would be the same defect with a different text.
1. Question
Do two independent published English translations of Chekhov's «Пари» (1889) satisfy a pre-registered, measured qualification test for use as Tier D's held-out arm — and does that test reject the pair the project has already rejected?
2. What qualification can and cannot mean
This has to be said before the gates, because it is the thing most easily overclaimed.
Passing these gates does not establish that the two translations are of equal quality. No measurement in this design could. What the gates do is rule out the confounds that made the Beowulf pair uninterpretable: a difference of form, of period idiom, of length, of sentence architecture, or an outright omission in one text. With those excluded, a jury that then reports a large quality gap is telling us something about the jury. With them present, the same number tells us nothing, because verse-against-prose across a century of English will separate on any sense whatever.
So: the gates make the held-out arm interpretable. The arm then measures the jury. Whether these two translations are in fact of equal quality is not asserted here and must not be asserted by whoever runs the arm; "chance expected" is the charter's null hypothesis for the jury, not a claim about the texts.
3. Materials
All four texts are public domain and freely reachable. Nothing copyrighted is stored.
| id | text | provenance | PD basis |
|---|---|---|---|
| S | Чехов, «Пари» (1889), 2,189 words — wiki/base/anchors/A-chekhov-pari/pari-chekhov-1889-ru.txt |
ru.wikisource, reproducing ПСС vol. 7 pp. 229–235 (FEB) | Chekhov d. 1904 |
| T1 | S. S. Koteliansky & J. M. Murry, "The Bet", in The Bet, and Other Stories (Boston: John W. Luce & Co., 1915) — Gutenberg #55283 | published human translation | US publication 1915, pre-1929 |
| T2 | Constance Garnett, "The Bet", in The Schoolmistress and Other Stories = The Tales of Tchehov vol. IX (London: Chatto & Windus, 1920) — Gutenberg #1732 | published human translation | UK/US publication 1920, pre-1929 |
| T3 | the lead's T-pari-R04-v1 (2026), complete |
lead translation, frozen at commit 37ad54e before T1 and T2 were fetched |
lead-authored |
Imprints for T1 and T2 were verified against library catalogue records (Internet Archive item metadata: betandotherstor00murrgoog, betotherstories00chekiala → Boston: J. W. Luce & Co., 1915; schoolmistressot00chekuoft → London: Chatto & Windus, 1920), not taken from the Gutenberg header alone. NEXT.md note (d): bibliographic priors are wrong often enough to matter.
Negative control — the pair the project already rejected. wiki/base/anchors/A-beowulf-ingeld/: Gummere 1909, Kirtlan 1913, Morris & Wyatt 1895, and the lead's T-beowulf-ingeld-R04-v1. Every pairing drawn from these must fail at least one gate, or the instrument does not reproduce a judgment the project has already made on independent grounds.
Why this source and this pair. Both translators are working in the same decade, in the same publishing milieu, on the same prose short story, from the same author; the gap between imprints is five years. NEXT.md action 1 named exactly this configuration ("multiple PD English Chekhov translators of the same story … same decade, same prose register"). Two further stories overlap between the same two volumes — "After the Theatre" and Припадок (K&M "The Fit" / Garnett "A Nervous Breakdown") — so a qualifying result extends to a second and third cell without new materials.
What T3 is for. It is not part of the candidate pair. It is a third point of known-different period, and it is the only check available on whether the period gate is sensitive — a qualification test that cannot distinguish a 2026 rendering from a 1915 one is not measuring period-matching. Its role was fixed before the gates were written, and its own log records the register decisions independently.
4. The gates
Instrument: tools/pair_match.py, written after this page is frozen, printing every intermediate count. A pair qualifies only if all six gates pass.
| gate | measure | criterion for the pair (T1, T2) |
|---|---|---|
| G1 form | median line length in characters over non-blank lines of the stored file (primary), with the fraction of non-blank lines beginning with an uppercase letter reported as corroboration | both texts median ≥ 55 chars (= prose, not verse lineation) |
| G2a period idiom | archaism rate per 1,000 words from the closed list in §5 plus the third-person -eth pattern |
both ≤ 5.0 per 1,000 and absolute difference ≤ 3.0 per 1,000 |
| G3 undamaged | the 15 landmark facts of §6 present in both; total sentence counts within ±15% of each other | all 15 present in both; any miss triggers a hand read against the Russian, recorded verbatim, and disqualifies the text if it is a real omission |
| G4 length | word-count ratio, larger ÷ smaller | ≤ 1.15 |
| G5 sentence architecture | mean sentence length in words | difference ≤ 20% of the smaller mean |
| G6 register proxy | long-word rate: percentage of word tokens of ≥ 8 characters | difference ≤ 3.0 percentage points |
Where the thresholds come from. G1, G2a and G4 are set so that they reject the Beowulf pair on the figures already on record (Gummere 403 words / 56 verse half-lines / 5 archaic tokens = 12.4 per 1,000; Kirtlan 489 words / 17 archaic tokens = 34.8 per 1,000; the lead's 0; Kirtlan ÷ Gummere = 1.21). That is the only calibration available: the instrument must reproduce a rejection the project made on independent grounds, or it is not measuring what disqualified that pair. G5 and G6 are set from principle — a translation that systematically splits or fuses periods, or that sits at a different Latinate density, is stylistically distant from its partner whatever the dates say.
G4 is not blind, and is marked accordingly. The word counts of T1 (2,691) and T2 (2,829) were computed while the texts were being extracted, before this page was written; the ratio is therefore known, and a ±15% gate that the pair passes cannot count as a prediction. It is retained because dropping the gate that disqualified Beowulf would weaken the instrument, and because it still binds the negative control. It carries no independent evidential force for the candidate pair. Everything measured by G1, G2a, G3, G5 and G6 was unknown when this page was frozen; the only text of T1 and T2 seen at freeze time was the first ~120 and last ~160 characters of each, printed by the extraction script as a sanity check.
A second period measure, reported and not gated (G2b). Period orthography: occurrences of to-day, to-morrow, to-night, shew(n/ed), connexion, and the spaced forms any one, some one, every one, per 1,000 words. This is expected to separate 1915/1920 from 2026 while G2a separates Kirtlan's pseudo-biblical idiom — two different axes, and conflating them is how a "period" gate ends up measuring one thing and being reported as another. It is not gated on the pair, because a small orthographic difference between two texts five years apart is not a period mismatch.
5. The archaism list (closed, frozen here)
Case-insensitive whole-word matches:
thou, thee, thy, thine, hast, hath, doth, dost, didst, shalt, wilt, wouldst, shouldst, couldst, canst, saith, wert, nay, verily, whilst, amongst, betwixt, ere, oft, whither, whence, hither, thither, wherefore, methinks, perchance, forsooth, albeit, unto, o'er, e'er, ne'er, 'tis, 'twas, mayhap, alack, prithee, lo, aught, naught, nought
Plus the pattern \b[a-z]{3,}eth\b, excluding the stoplist death(s), breath(s), teeth, beneath, underneath, wreath(s), sheath(s), heath(s), bequeath(s), twentieth, thirtieth, fortieth, fiftieth, sixtieth, seventieth, eightieth, ninetieth, hundredth, thousandth.
Deliberately excluded as ambiguous, so that both texts are undercounted equally rather than one being penalised by a false positive: art, ye, yea, hence, thence, behold, upon. upon is reported separately as a formality proxy and is not an archaism — 1915 and 2026 formal prose both use it.
6. The fifteen landmark facts (frozen here)
Concrete, countable content from the Russian, chosen so that an omission would be visible. Each is matched by a tolerant alternation to allow for different wording, then hand-checked if it misses.
two million · twenty-five years old (the lawyer's age) · fifteen years (the term) · 14 November 1870 → 14 November 1885 · six languages (the letter) · two shots fired in the garden · about six hundred volumes · the Gospel read for about a year · Byron or Shakespeare · Elbrus and Mont Blanc · frogs and lizards on apple and orange trees · roses smelling of a sweating horse · mice under the floor · forty years old (the prisoner at the end) · five hours before the appointed term
7. Predictions, pre-registered
- P1. T1 and T2 both pass G1 and G2a. (Blind.)
- P2. Every pairing among the Beowulf texts fails at least one gate: Gummere on G1, Kirtlan on G2a, and at least one pairing on G4.
- P3. T3 (2026) is separated from T1 and T2 by G2b, at a rate difference of at least 2.0 per 1,000. If instead all three sit near zero on G2b, the finding is that these measures do not separate 1915/1920 British prose from 2026 prose, the period gate is declared insensitive at this dose, and the qualification is reported as resting on G1/G3/G5/G6 plus the documented five-year imprint gap rather than on a demonstrated period measurement. Both branches are reported; neither is a failure of the run.
- P4. All 15 landmarks are present in T1, T2 and T3. Any miss is read by hand and recorded.
- P5. G5 and G6 pass for the pair. (Blind — this is the prediction most likely to fail, because two translators of the same story may well differ in period length by more than 20%.)
8. Failure criteria — what would make this run a null
- Any gate fails for (T1, T2) → the pair does not qualify. The Tier D blocker stays,
config/models.mdstays NOT CALIBRATED, andNEXT.mdrecords which gate failed and what a replacement pair would have to look like. The next candidates are named in §3 (the two other overlapping stories) and, failing those, a different translator pair. - The negative control passes any Beowulf pairing on all six gates → the instrument is void, whatever it says about Chekhov, and nothing is claimed. This is the check that the gates are not simply loose.
- A real omission is found in T1 or T2 → that text is disqualified as a held-out reference, and this is a finding about the text, recorded on the anchor page.
- P5 fails while P1 and P4 hold → report the pair as matched on form, period and completeness but not on sentence architecture or register, do not call it qualified, and record it as the closest candidate yet with the specific residual mismatch named. A partial pass is not a pass.
9. Procedure
- This page is committed. (Freeze.)
- Independent pre-run critic pass, routed through one non-Anthropic panel model (
config/models.md), which receives this design, the archaism list, the landmark list, and the four texts' provenance, and is asked to attack the gates — in particular to find gates that cannot fire, gates tuned to the answer, and confounds the six gates miss. Its verdict and every disposition are recorded incritic.md.NEXT.mdnote (n): the standing dispositions inworkshop/experiments/README.mdare checked first so that this design does not reintroduce a defect an earlier critic already established.NEXT.mdnote (o): each gate is checked for whether it can fire before dispatch. tools/pair_match.pyis written and run; its full output is preserved atruns/pair_match.out.- Post-run verification recomputes every reported number independently of the analysis prose (
verification.md). - Results: a Tier 1 precedent anchor page for the pair (
A-chekhov-pari) and a result page;config/models.mdupdated to record whether the Tier D blocker is removed. No Tier D claim is made by this run — it qualifies materials; it does not run the arm.
10. Budget
Pre-flight: the run itself is $0.00 (no API calls). The critic pass is one call to one non-Anthropic panel model with a long prompt (design + lists ≈ 4,500 input tokens) and a reasoning-heavy response; per NEXT.md note (m), the estimate is built from per-call maxima: central $0.04, worst case $0.12 on x-ai/grok-4.5 at $2.00/$6.00 per M with up to 12k reasoning+output tokens. Today's remaining headroom is ~$0.85. The worst case fits with a large margin.