Repository path: workshop/experiments/E-20260726-genealogie-period/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260726-genealogie-period |
| status | frozen |
| created | 2026-07-26 |
| updated | 2026-07-26 |
| senses | accuracy, style-correspondence |
| internal-judgment-only | true |
| provisional | true |
| links | workshop/experiments/E-20260726-period-control/design.md, workshop/experiments/E-20260725c-contamination-sweep/design.md, workshop/translations/genealogie/R04-v1/translation.md, tools/ngram_overlap.py, workshop/regimes/R04-lead-close.md, config/models.md |
Frozen design — the period control, second attempt: a named modern translator in the room
Frozen 2026-07-26 (S028). Nothing in §§1–9 is written after seeing a single number from the run. The freeze is verifiable in git: the commit that adds this file adds the lead translation T-genealogie-R04-v1 and the German source, creates no runs/ directory, and adds no extraction or analysis tool. No English rendering of the three target sections had been fetched or read when this file was committed.
The wire, in one sentence. The translation limb — three complete sections of Zur Genealogie der Moral rendered blind from the 1887 German — supplies the subject whose overlap the study limb places inside a 78-section distribution of agreement between a 1913 and a 2014 published translator, which is the period control the project has twice failed to build.
1. Why this runs, and what killed the last attempt
NEXT.md action 1, top action, second attempt. E-20260725c-contamination-sweep found that the lead is the most central text in its cell in six cells of six, and could not tell two explanations apart:
- H-lead — centrality is a property of this translator.
- H-period — twentieth-century English translations are individually idiosyncratic; twenty-first-century plain English is the register they each deviate from, in different directions, so any modern translator sits near the centroid of a cell of Victorians.
Separating them needs a modern translator in the room. E-20260726-period-control (S027) tried and returned a null: the one freely reachable post-2000 translation of a text with two public-domain comparators — Translation:Rose (Turgenieff) on en.wikisource, header "Translated from Russian by User:Davidludi", 2011, CC BY-SA — was machine output, and the design's gate did not catch it because the gate tested derivation from the comparators, not human authorship. That produced note (mm): metadata that asserts human authorship is not evidence of human authorship — read enough of the text to see whether a person wrote it, and read it on a different passage from the one the session will translate, so the freeze survives.
This design does that first, and records the result before anything else (§4).
2. Questions
- Q1 — the period test proper. Does the lead sit closer to the 2014 translator than to the 1913 translator? H-period predicts a clear tilt towards the modern voice. H-lead predicts near-symmetry, which is what the sweep and S027 both found against pairs of Victorians, where symmetry was uninformative about period because there was no modern text to be asymmetric towards.
- Q2 — the denominator, on a second pair, a second language and a second century-gap. S027 measured Garnett 1897 ~ Hapgood 1904 on 42 matched prose poems and found the rate varies 17.8× min-to-max, which is why
CIwas downgraded to ordinal. Does the instability reproduce on a cross-period pair (101 years apart), in German→English, on argumentative prose rather than prose poems? - Q3 — the rank statistic, out of sample. Note (nn) replaced division with placement: report the lead's overlap as a rank inside the measured published-versus-published distribution. That statistic was invented and applied in the same session, on one work. Does it behave on a new work, a new pair and a new language pair?
- Q4 — centrality with a modern comparator present. With a 2014 human translation as one of three texts, is the lead still the most central?
3. Materials
One work, three complete sections, three texts per section.
| label | translator | year | status | how reached |
|---|---|---|---|---|
de |
— (source) | 1887 | public domain | Nietzsche Source eKGWB, nietzschesource.org/texts/eKGWB/GM; every word cross-checked against Projekt Gutenberg-DE |
samuel |
Horace B. Samuel, M.A. | 1913 | public domain | Project Gutenberg #52319, gutenberg.org/files/52319/52319-h/52319-h.htm (Levy edition, T.N. Foulis, Edinburgh & London) |
johnston |
Ian Johnston, Prof. Emeritus, Vancouver Island University | 2014 (rev. of 2009) | in copyright, free to read and redistribute non-commercially | web.viu.ca/johnstoi/nietzsche/genealogy{preface,1,2,3}.htm |
lead |
the lead agent | 2026 | project artifact | workshop/translations/genealogie/R04-v1/translation.md |
Note on the host. NEXT.md recorded johnstoniatexts.x10host.com as verified live. It is live but serves an empty directory index and 404s on every content path tried; the working host is web.viu.ca/johnstoi/. Recorded here so the next session does not repeat the search.
Copyright hygiene (charter §7, rule 5). Johnston's text is in copyright. It is fetched to a scratch directory outside the repository, reduced to counts, and never stored whole in the repo. What this experiment commits is derived statistics plus brief attributed excerpts. The consultation is logged in wiki/base/consulted.md.
Target sections and how they were chosen. Excluding the Vorrede, take from each Abhandlung the section whose German whitespace-token count is closest to 350; ties to the lower section number. This yields I §1 (373 German tokens), II §9 (353; tie with II §10 resolved down), III §4 (409). The rule uses length only, was applied mechanically before any German was read, and cannot have selected for passages that suit the translator or the measurement. The Vorrede is excluded because §4 uses it.
4. The translator-verification gate, discharged before the freeze
Per note (mm), and before the target sections were selected or translated:
Johnston's Prologue §1 (a passage excluded from selection by construction) was read in full against the German. It renders «Wir sind uns unbekannt, wir Erkennenden» as "We don't know ourselves, we knowledgeable people"; «wo euer Schatz ist, da ist auch euer Herz» as "Where your treasure is, there shall your heart be also", carrying a numbered endnote to the biblical source; the noon-clock passage accurately; and the inverted proverb «Jeder ist sich selbst der Fernste» as "Each man is furthest from himself". The page carries a signed preliminary note stating an editorial policy on italics, bracketed insertions and the retention of Nietzsche's punctuation; contractions are used systematically as a register decision; a print edition exists from Richer Resources Publications, and a Broadview edition edited by Gregory Maertz. Verdict: human, competent, and a genuinely contemporary English voice. This gate is discharged. It is re-checked once more against the target sections at run time (F4 below), because a gate discharged on one passage is evidence about that passage.
Samuel 1913 was opened only as far as its title page and table of contents, which name the translator, the Levy edition and the publisher. Its provenance is a 1913 print volume; no humanity gate is needed and none is claimed.
5. Procedure
5.1 Extraction, frozen
samuel. From the Project Gutenberg HTML: delete every<div class="footnote">block; delete everything from the<h4>whose text isPEOPLES AND COUNTRIES.onward (the appended Kennedy fragment is a different work by a different translator); split the remainder at each<p class="parnum">, grouping under the preceding<h5>ofPREFACE./FIRST ESSAY./SECOND ESSAY./THIRD ESSAY.johnston. From each of the four pages: delete everything from the centred paragraph whose text isENDNOTESonward (the endnotes carry Nietzsche's foreign-language originals and would pollute overlap); split the remainder at each centred paragraph whose entire text is a small integer.lead. The three### First Essay, §1/### Second Essay, §9/### Third Essay, §4blocks oftranslation.md, up to the next###or---.- Common. Strip tags, unescape entities, collapse whitespace; then delete standalone digit tokens (endnote and footnote reference markers; applied symmetrically to all three texts).
Structure was inspected before the freeze — heading and marker lines only, never prose — which is how the two splitters above could be written down rather than tuned afterwards.
5.2 The statistic, frozen
Exactly the metric of tools/ngram_overlap.py, which implements E-20260725c-contamination-sweep/design.md §4: NFC normalise, map curly quotes to straight and em/en dashes to space, lowercase, delete everything outside [a-z0-9'], whitespace-split. For a pair of texts, count shared n-gram types for n ∈ {4, 5, 6, 7} and express as shared types per 1000 tokens of the shorter text.
5.3 The free parameter, bracketed rather than chosen (note (kk))
The proper-noun exclusion is a free parameter nobody can justify, so every figure is reported under both parameter-free extremes and one declared sensitivity variant:
- V-frozen — the R1 ∪ R2 name rule exactly as
tools/ngram_overlap.pycomputes it. - V-none — no name exclusion at all.
- V-ellipsis — R1 as frozen, plus: a token whose final character is
…counts as sentence-ending. This variant is declared because Nietzsche's prose is full of sentence-final…and?…, and R1's exemption list is.!?:"— so under V-frozen every capitalised word that opens a sentence after an ellipsis is misread as a proper noun. This is note (ii) firing for the fourth time, on a fourth typographic genre: prose, verse, OCR running heads, dash-introduced dialogue, and now ellipsis-terminated sentences. It is declared in advance rather than discovered in the results, and it is a sensitivity check, not the headline.
5.4 Cell size is held constant (note (oo))
S027 found that CI and centrality depend on how many texts are in the cell — 2.125 with three published texts, 1.500 with two, same lead translation. So:
- Every number entering a rank comparison is computed in a 2-text cell. The reference distribution is
samuel~johnstonper section, cell = {samuel, johnston}. The lead's placements arelead~samuel(cell = {lead, samuel}) andlead~johnston(cell = {lead, johnston}). Under V-none cell composition is irrelevant by construction, which makes V-none the cleaner of the two extremes here. - Centrality is computed in 3-text cells ({lead, samuel, johnston}, target sections only) and is labelled as cell size 3 wherever it is reported. It is never compared against a 2-text figure.
5.5 The controls
- Matched-section reference distribution.
samuel~johnstonon all 78 aligned sections. This is a cross-period published pair — 101 years — which is precisely the comparison the project has never had. - Mismatched-section floor control (note (p): a case the instrument should pass). For each section i, pair
samuelsection i withjohnstonsection i+1 (cyclic). Two competent English translations of different Nietzsche sections should share far less than two translations of the same section. If they do not, the statistic is measuring English rather than translation, and every rank below is meaningless.
6. Predictions, registered
Registered before any English of the target sections was read. P4–P6 are the lead's third registered self-estimate; the first two both missed, in the flattering direction (note (jj)).
- P1 (alignment).
samuelyields exactly 8 / 17 / 25 / 28 numbered sections andjohnstonyields exactly 8 / 17 / 25 / 28. - P2 (floor control). Under V-none at n = 5, the median mismatched-section rate is below 20% of the median matched-section rate.
- P3 (denominator instability generalises). Under V-none at n = 5, max/min of the matched
samuel~johnstonrate across usable sections is ≥ 5×. - P4 (the period test — the one that matters). At n = 5,
lead~johnstonexceedslead~samuelin at least 2 of the 3 target sections, under both V-frozen and V-none. - P5 (magnitude). Under V-none at n = 5, the mean over the three sections of
lead~johnston/lead~samuellies in [1.0, 1.6]. - P6 (rank). Under V-none at n = 5,
max(lead~samuel, lead~johnston)falls in the top 15% of the matchedsamuel~johnstondistribution in at least 2 of the 3 sections. - P7 (centrality, cell size 3). At n = 5, the lead is the most central of the three texts in at least 2 of the 3 sections, under both V-frozen and V-none.
- P8 (contamination
high, operationalised). Under V-none, in at least 2 of the 3 sections the lead shares at least one 7-gram with at least one published text.
7. Failure criteria, pre-committed
- F1. If either published text fails to yield exactly 8 / 17 / 25 / 28 sections, alignment is not established: the cell is void, and the session reports the alignment failure rather than any overlap number. (S027's alignment defect — a heading rule that read speaker labels as titles — cost 27 of 42 pairs and was only recovered by a declared post-hoc realignment. Here the counts are pre-registered.)
- F2. Any aligned section whose
samuelandjohnstontoken counts differ by more than 2× is excluded from the reference distribution; the exclusions are counted and named. - F3. If fewer than 60 sections survive F2, the distribution is still reported, but every rank statement is written as "k of N" with N stated, and no percentile language is used anywhere.
- F4. If, when the target sections are finally read,
johnstonshows any sign of machine origin — the failure mode that voided S027 — the cell is void and the run is reported as a second null. - F5. If the floor control (P2) fails, no rank statement is made at all, in this session or by citation later.
8. What each outcome licenses, and what it does not
| result | licensed reading |
|---|---|
P4 holds, clear tilt to johnston |
H-period gains real support: the lead's closeness tracks the century, not the translator. Every centrality result in the sweep is then partly an artifact of comparing 2026 prose with pre-1920 prose, and must be re-stated. |
| P4 fails, near-symmetry | H-period loses its most direct test on this work and this pair. It does not establish H-lead: symmetry between a 1913 and a 2014 comparator is evidence that the lead is not period-tracking here, on argumentative German prose, and nothing more. |
P4 fails, tilt to samuel |
The most interesting outcome and the one no hypothesis predicts. It would be reported as an anomaly, not explained, and it would need replication before it meant anything. |
Nothing here is a quality claim about any of the three translations. Overlap is not merit, in either direction. The lead does not judge its own translation (charter §5); this experiment measures wording overlap and nothing else, and every figure it produces is internal-judgment-only and provisional until Tier D passes.
Extent, stated honestly. One work, one language pair, three sections, one published pair. The reference distribution is 78 sections, which is the strong part; the lead's placement rests on three measurements, which is the weak part, and no amount of precision on the distribution repairs that. Any claim from this run is a claim about Zur Genealogie der Moral as rendered by Samuel and Johnston, until it is replicated.
9. Pre-run critic
Per charter §8, an independent pre-run critic pass on this frozen design, routed to a non-Anthropic panel model (config/models.md), before any measurement is run. Critic findings and the response to each are recorded in critic.md in this directory. Findings that require a design change are applied by amending this file with the amendment marked and dated, never by silent edit; findings declined are recorded with the reason.
Executed 2026-07-26, openai/gpt-5.6-terra (P1), $0.080616, raw in runs/critic-P1.json, findings and responses in critic.md. Five defects, eleven confound mechanisms, eight under-specified phrases, and two predictions ruled unfalsifiable as written. All applied except one declined as unfixable; the amendments are §10. §§1–9 above are the text as frozen before the critic ran and are not edited.
10. Amendments — added 2026-07-26 after the critic pass, before any measurement
Every amendment below was written before a single overlap number existed. runs/ did not exist when this section was committed.
A1 — what a tilt towards johnston is allowed to license (critic C1)
§8's first row over-claimed. There is one translator per period. A tilt towards johnston is not separable from "this lead's wording resembles Ian Johnston's", and "Ian Johnston" is not a sample of 21st-century English. The row is replaced by:
| result | licensed reading |
|---|---|
P4 holds, clear tilt to johnston |
Consistent with H-period and equally consistent with "the lead resembles this one translator". The design cannot separate them, and no wording in the result may imply otherwise. What it would license is a next experiment — several independently chosen translators per period — not a conclusion about period. |
The other two rows of §8 stand. Separating period from translator identity needs ≥2 translators per period on the same work and goes to NEXT.md as an action, not into this run's conclusions.
A2 — the contraction channel, tested rather than named (critic C3)
Johnston's preliminary note documents systematic contractions; Samuel is Edwardian and expands. Overlap could therefore tilt modern for a purely orthographic reason.
- Diagnostic: report, per text per section, the count of contraction tokens (tokens matching
\b\w+'\w+\bafter normalisation) per 1000 tokens. - Fourth variant, V-expand: before tokenising, expand a fixed list in all three texts —
don't→do not,doesn't→does not,didn't→did not,isn't→is not,aren't→are not,wasn't→was not,weren't→were not,can't→cannot,couldn't→could not,wouldn't→would not,shouldn't→should not,won't→will not,hasn't→has not,haven't→have not,hadn't→had not,it's→it is,that's→that is,there's→there is,he's→he is,she's→she is,what's→what is,who's→who is,let's→let us,we've→we have,they've→they have,I've→I have,you've→you have,we're→we are,they're→they are,you're→you are,we'd→we would,I'd→I would,we'll→we will,I'm→I am. Possessive'sis untouched. The list is fixed here, before the run. - Rule: if a tilt reported under V-none does not survive V-expand, it is reported as a contraction artifact and P4 is treated as failed.
A3 — the reading rule for three sections (critic C4)
P4 stands exactly as registered. Its interpretation is pinned now:
- 2 of 3 in the same direction is not evidence. Exact binomial P(≥2 of 3) under a symmetric null is 0.5. If P4 holds at 2 of 3 it is reported as "P4 holds as written, and the result is uninformative", in those terms.
- 3 of 3 under both parameter-free extremes is P = 0.25 under the same null — suggestive at best, and it will be called that and nothing stronger.
- "Clear tilt" (§8, undefined) is now: same direction in all three sections, under both V-none and V-frozen and surviving V-expand, and mean |log₂(
lead~johnston/lead~samuel)| ≥ 0.5, i.e. a ratio of at least 1.41× or at most 0.71×. Anything short of all four conditions is "no clear tilt". - Power is a property of this run and is stated in the result, not left for a reader to notice: three sections, one work, one pair.
A4 — P8 demoted, and the null it needed (critic C2)
P8 is not a contamination test and is not reported as one. A shared 7-gram between independent translations of the same German sentence is ordinary; contamination by memory can be paraphrastic and leave no long string. P8 is retained, as registered, as a check on whether the declared high has any long-string consequence at all.
The null the critic asked for is already in the materials and is now computed:
- Independent-pair null: the distribution of shared 7-gram counts between
samuelandjohnstonon the same section, across all 78. Two translations known to be independent, 101 years apart. - Pure-English null: the same, on mismatched sections (§5.5), where any shared 7-gram is a fact about English, not about translation.
A lead 7-gram count is reported against both. Neither null makes the lead's declaration none; the declaration stays high on provenance grounds regardless of what the counts do.
A5 — reproducibility of a mutable source (critic C5)
SHA-256 and retrieval timestamp of every fetched file — the four Johnston pages, the Project Gutenberg HTML, and both German witnesses — are recorded in runs/manifest.json. The copyright rule forbids storing Johnston's text in the repo, so the hash is the only reproducibility handle available.
A6 — the eight under-specified phrases, pinned (critic item 3)
- "clear tilt" — defined in A3.
- "any sign of machine origin" (F4) — three concrete tests, applied to the target sections when they are finally read: (a) a word rendered as a homograph or near-homograph of the German rather than by sense; (b) a fluent English clause that is referentially absurd against the German; (c) a clause present in the German and silently absent or garbled. One confirmed instance voids the cell. These are the three failure classes the S027 machine text actually exhibited.
- "usable sections" — aligned and surviving F2. Sections whose rate is zero are retained; dropping zeros would bias every distribution upward and every rank against the lead.
- "top 15%" —
rank= the number of reference values strictly greater than the lead's value;percentile=rank / N. Ties count as not greater, which is the direction conservative against the lead. "Top 15%" meansrank / N ≤ 0.15. - "the most central" — centrality of a text in the 3-text cell = the mean of its two pairwise per-1000 rates. Highest wins; an exact tie means no winner is declared and the section counts against P7.
- "all 78 aligned sections" — alignment is positional within each division, and F1's counts are joined by a structural check: Pearson r between German section token count and English section token count must be ≥ 0.90 for both
samuelandjohnstonacross the 78 sections. Below that, alignment is not established and F1 voids the cell. This is the check that would have caught S027's 13-position drift, which passed a count test and failed the texts. - The humanity gate's qualitative basis — declined as unfixable; no rubric identifies machine output, reading does. The evidence is quoted in §4 so a reader can disagree, and it stays
internal-judgment-only. - The Johnston version — A5.
A7 — confounds recorded as untested
The critic's eleven mechanisms stand as the confound list for this and any future period work (critic.md). Two were absent from the design and are recorded as untested: training-exposure asymmetry between the two published texts, and the fact that a 2014 revision of a 2009 translation by one academic is not a sample of contemporary English.