Repository path: workshop/experiments/E-20260726-period-control/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260726-period-control |
| status | frozen |
| created | 2026-07-26 |
| updated | 2026-07-26 |
| senses | accuracy, style-correspondence |
| internal-judgment-only | true |
| provisional | true |
| links | workshop/experiments/E-20260725c-contamination-sweep/design.md, wiki/findings/results/RS-20260725c-contamination-sweep.md, workshop/translations/roza/R04-v1/translation.md, tools/ngram_overlap.py, workshop/regimes/R04-lead-close.md |
Frozen design — the period control: is the lead's centrality a property of the lead, or of writing in 2026?
Frozen 2026-07-26 (S027). Nothing in §§1–8 is written after seeing a single number from the run. The freeze is verifiable in git: the commit that adds this file adds no English text of «Роза», creates no runs/ directory, and adds no tool. The lead translation that is one of the four subjects, T-roza-R04-v1, is added by the same commit and was written before any comparator was fetched.
1. Why this runs
NEXT.md action 1, the top action, and it is top because it is the one confound the contamination sweep declared and could not control. E-20260725c-contamination-sweep/design.md §7, verbatim:
Period is confounded with authorship throughout. In six of seven cells the baseline pair are two people writing within twenty years of each other and the lead is writing in 2026. There is no free modern published translation stored for any cell. P3 is the only handle the design has on this, and it is a weak one.
The sweep's principal finding — note (ll) — was that the lead is the most central text in its cell, 6 cells for 6, 12 for 12 across both parameter-free extremes: the largest share of its wording occurs somewhere else in the cell. The sweep read that geometrically rather than psychologically (a text near the centroid is nearer every point than the points are to each other, with no memory involved).
But there is a rival geometric story the sweep could not touch, and it is not about the lead at all. Twentieth-century English translations are individually idiosyncratic; twenty-first-century plain English is the register they each deviate from, in different directions. On that story any modern translator sits near the centroid of a cell of Victorians, and the lead's centrality is a fact about the year rather than about the lead. The sweep had no modern translator to put in any room, so it could not tell the two apart.
This run puts one in the room.
2. Questions
- Q1 — the replication. With a 21st-century human translation of the same complete text present as a fourth voice, is the lead still the most central text in the cell?
- Q2 — the period test proper. Is the 21st-century human comparator also more central than the two Victorians are to each other? If it is, centrality is at least partly a property of modern English; if it is not, it is a property of the lead.
- Q3 — the denominator, attacked properly. Every
CIfigure this project has reported divides by one published pair, and note inNEXT.mdrecords that the Garnett~Hapgood pair on three different Turgenev passages varies by ~4.5×. What does the Garnett~Hapgood overlap rate actually look like over several dozen matched passages? - Q4 — is the modern comparator independent? A Wikisource contributor translation could have been made with Garnett or Hapgood open. This is measurable, and it gates Q1 and Q2.
3. Materials
One cell, roza, four texts, one complete work — not a passage. Turgenev's «Роза» (April 1878), 319 Russian tokens, from Стихотворения в прозе. Choosing a complete short work rather than an excerpt means all four texts translate exactly the same unit, so no hand-alignment is involved and the extent rule is satisfied by construction.
| label | translator | year | how reached |
|---|---|---|---|
lead |
the lead agent | 2026 | workshop/translations/roza/R04-v1/translation.md, frozen in this commit |
garnett1897 |
Constance Garnett | 1897 | Project Gutenberg #8935, Dream Tales and Prose Poems, section POEMS IN PROSE, poem THE ROSE |
hapgood1904 |
Isabel F. Hapgood | 1904 | Project Gutenberg #15994, A Reckless Character, and Other Stories, section POEMS IN PROSE, poem THE ROSE |
davidludi2011 |
Wikisource contributor User:Davidludi |
2011 | en.wikisource.org, Translation:Rose (Turgenieff), page created 2011-03-12, CC BY-SA |
Garnett 1897 and Hapgood 1904 are the same published pair that supplies the baseline in three of the sweep's six cells (svidanie, bezhin-lug, dvoryanskoe-gnezdo). That is the reason this text was chosen over any other with a modern free rendering: the new cell's baseline is directly commensurable with the ones already measured, and the new voice is the only thing that has changed.
What davidludi2011 is and is not. It is a 21st-century human translation into English, made independently of this project, from the Russian, published under a free licence. It is not a professionally published literary translation, and no claim about its quality is made or implied anywhere in this design — the metric counts shared wording and knows nothing about quality. What the design needs from it is exactly one property: that a human, writing English in the 2010s rather than the 1890s, translated the same complete text. Its independence from Garnett and Hapgood is not assumed; it is gated in P5.
Priming, declared, and it is not nil. Seven words of davidludi2011 — its title line and its rendering of the poem's first sentence, truncated mid-word — were returned inside a MediaWiki revision-history edit summary while establishing from page metadata that the translation exists. This is recorded in full on T-roza-R04-v1 per R04 §1, and it is why P6 exists.
The reference set for Q3. Every prose poem in Senilia that appears in the POEMS IN PROSE section of both Gutenberg volumes. Both files were downloaded 2026-07-26 from gutenberg.org/cache/epub/<id>/pg<id>.txt; their SHA-256 digests are recorded in the run output. Only aggregate statistics and per-pair token counts are stored in this repository, not the poem texts, which are two clicks away and would be repo bloat.
Extraction rules, frozen verbatim.
lead: the text oftranslation.mdbetween the line## The translationand the next line consisting of exactly---, with lines beginning##removed. (tools/ngram_overlap.py's existing rule for.mdfiles, unchanged.)garnett1897,hapgood1904: within each Gutenberg file, thePOEMS IN PROSEsection is the text following the second occurrence of a line equal toPOEMS IN PROSE; within it, a heading is a line, preceded and followed by a blank line, consisting only of capital letters, spaces, apostrophes, commas, full stops, hyphens and exclamation or question marks, of length ≤ 60 characters, and not a Roman numeral. A poem is the text from a heading to the next heading.THE ROSEis the heading of the subject poem in both volumes.davidludi2011: theTEXTof the page's rendered wikitext with template markup, the header template, category links and interwiki links removed.- Gutenberg licence boilerplate, footnote markers and endnote apparatus are removed by the rules above or fall outside the extracted poem.
- No manual editing of any subject text. Any character that has to be touched by hand is reported in verification.
4. The metric
Unchanged and already frozen. tools/ngram_overlap.py, frozen verbatim in E-20260725c-contamination-sweep/design.md §4 and not altered since: NFC normalisation, lowercase, non-[a-z0-9'] → space, shared n-gram types at n = 4, 5, 6, 7, per-cell baseline = mean over all published×published pairs, CI(n) = max over published P of shared(lead,P,n) / B(n).
Bracketed, per note (kk). Every figure is reported at both parameter-free extremes of the metric's one free parameter — frozen (the aggressive proper-noun rule) and none (no exclusion at all). No third setting is computed and none is chosen.
Centrality, now frozen verbatim instead of ad hoc. In the sweep, centrality was a post-hoc diagnostic whose script was not committed — note (r)'s failure mode. Here it is part of the frozen design and its implementation is committed with the run:
For a cell of texts S and a given n, C(T, n) = 100 × |{g ∈ ngrams(T,n) : g ∈ ngrams(U,n) for some U ∈ S, U ≠ T}| / |ngrams(T,n)| — the percentage of T's n-gram types that occur somewhere else in the cell. Headline n = 5, matching the sweep's headline, with n = 4, 6, 7 reported alongside.
Reference distribution (Q3). For every prose poem present in both volumes: align the two heading sequences by order; a pair enters the reference set only if the shorter text has ≥ 120 tokens, the extent ratio max/min of tokens is ≤ 1.25, and shared(n=4) > 0. The shared(4) > 0 test is a parameter-free alignment check — two translations of the same short text always share at least one four-word run — and every pair it excludes is listed and counted. It is conservative in the direction of a higher baseline, hence of a smaller CI, and that direction is declared here rather than discovered later. The statistic is the rate: shared n-gram types per 1000 tokens of the shorter text, at each n, under both extremes.
5. Predictions — REGISTERED BEFORE THE RUN
What the predictor had seen when predicting. The sweep's full results. The lead's own «Роза». The seven words of davidludi2011 quoted above. The heading lists of both Gutenberg volumes' POEMS IN PROSE sections, extracted mechanically to confirm that THE ROSE is present in both. No sentence of Garnett's or Hapgood's «Роза», and no further sentence of Davidludi's.
P0 — the lead's self-estimate, scoreable, one more datum for note (jj). CI(5) < 2.0 under both extremes. Reasoning: Poems in Prose are less reprinted than «Бежин луг» or "The Bet", and the sweep's six cells ran 1.26–6.7 with four of six under 2.0.
P1 — the replication, and the point of the run. The lead has the highest C(·, 5) of the four texts, under both extremes.
P2 — the period effect on centrality. C(davidludi2011, 5) > max(C(garnett1897, 5), C(hapgood1904, 5)) under both extremes. Reasoning: the modern-register mechanism in §1. P2 is the prediction this run exists to test, and it is deliberately set against the outcome most flattering to the lead.
P3 — the period effect on the published texts themselves. Of the three published×published pairs, garnett1897~hapgood1904 has the largest shared(5). Reasoning: two translators seven years and one register apart, the later of whom could have read the earlier. If P3 fails, the metric does not detect a period signal even where one must exist, and P2's mechanism is undercut whatever P2 does.
P4. CI(5) > 1.0 under both extremes (the sweep's 6-for-6 pattern reproduces with a fourth voice present).
P5 — the gate on the modern comparator. Neither shared(davidludi2011, garnett1897, 7) nor shared(davidludi2011, hapgood1904, 7) exceeds 2 × shared(garnett1897, hapgood1904, 7).
P6 — the exposure sensitivity. Recomputing every headline with paragraph 1 removed from all four texts does not change the truth value of P1.
P7 — the denominator. Across the reference set (expected ≥ 20 pairs), max/min of the Garnett~Hapgood rate at n = 5 exceeds 3.0. Reasoning: 4.5× was observed on three passages.
P8 — is roza a freak passage? The roza Garnett~Hapgood rate at n = 5 falls inside the interquartile range of the reference distribution.
6. The decision table, written before the run
| P1 | P2 | what the run licenses |
|---|---|---|
| holds | fails | The period confound is retired: a period-matched human is present and is not central, the lead is. Centrality is a property of the lead. This is the strongest available support for note (ll)'s reading. |
| holds | holds | Centrality is partly a modern-register property. The informative quantity becomes the margin C(lead) − C(davidludi2011), and every centrality statement in RS-20260725c-contamination-sweep must be re-worded to say "most central among texts of its own period" is untested there. |
| fails | holds | Most of finding 2 is a period effect, exactly as NEXT.md action 1 anticipated. This is reported as the session's principal finding and note (ll) is amended in place. |
| fails | fails | The cell behaves unlike all six sweep cells and nothing is concluded from it beyond the failure to reproduce; the likeliest cause is that a 319-token complete poem is too short, and the power figures decide. |
7. Failure criteria and stopping rules, written before the run
- P5 fails.
davidludi2011is treated as possibly derived from a Victorian text, is disqualified as an independent modern witness, and Q1/Q2 are reported as unanswerable by this cell. The reference distribution (Q3) still stands, and the session reports a null on its principal question. - Extent. Any pair in the cell differing by more than 25% of tokens: the cell is dropped and said to be dropped (the sweep's rule, which dropped
beowulf-ingeld). - Power floor. If
B(7) < 10,CI(7)andC(·,7)are reported as underpowered and n = 5 carries the headline. At 319 source tokens this is likely and is not a surprise if it happens. - If fewer than 12 pairs survive into the reference set, Q3 is reported as underpowered and P7/P8 are not scored.
- If P1 and P2 disagree between the two extremes — true at one, false at the other — the prediction is scored FAILED, not "mixed". Note (kk) exists because a claim that survives only one setting of a free parameter is not a claim.
8. What this design cannot do
- It is one cell. A single 319-token poem cannot establish a general period effect. It can reproduce the six-cell pattern in the presence of a modern voice, or fail to; both are informative, neither generalises on its own.
- One modern voice is not "the 21st century".
davidludi2011is one person. If its centrality is low, that could be this translator rather than this period — a translator who is unusually idiosyncratic would look exactly the same. The design cannot separate them and does not claim to. The honest statement of a P2 failure is "the one period-matched human available is not central", not "modern translators are not central". - It cannot show the lead consulted anything. Freeze order is provable in git; high overlap under a provable freeze is evidence about memory, not about conduct.
- The reference distribution is one translator pair, one book, one genre. Garnett~Hapgood across the Poems in Prose tells us how much that pair's overlap varies over short lyric prose by one author. It is not a general baseline for translation overlap and must never be quoted as one.
- The seven-word exposure is real and P6 bounds it rather than removing it.
- Nothing here judges any translation, published or lead.
provisional: true,internal-judgment-only; sensesaccuracyandstyle-correspondenceare named only because shared wording touches both.