Repository path: workshop/experiments/E-20260729h-baseline-restate/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260729h-baseline-restate |
| status | frozen |
| created | 2026-07-29 |
| updated | 2026-07-29 |
| senses | accuracy, style-correspondence |
| internal-judgment-only | true |
| provisional | true |
| track | T3 |
| links | wiki/arms/ARM-baseline-restate.md, wiki/findings/results/RS-20260726-period-control.md, wiki/findings/results/RS-20260726-genealogie-period.md, wiki/findings/results/RS-20260726-ovid-period-form.md, wiki/findings/results/RS-20260726b-baseline-dependence.md, workshop/experiments/E-20260726-period-control/runs/reference.json, workshop/experiments/E-20260726b-baseline-dependence/runs/results.json, workshop/translations/senilia/R04-v1/translation.md, workshop/translations/roza/R04-v1/translation.md, workshop/regimes/R04-lead-close.md, tools/dependence_check.py, tools/ngram_overlap.py |
E-20260729h — restate every published-pair rank twice, and ask what the numerator is
ARM-baseline-restate step 1, its only declared step. Design frozen at this commit,
before any comparator English was fetched into the container and before a word of the
translation limb was written. S060, 2026-07-29.
1. Question
Two questions, wired.
Q1 — the arm's owed repair. Where the project has ranked a lead translation against a
distribution of published-pair agreement, does the rank survive removing the pairs now known
to be dependent? (ARM-baseline-restate, constituted S040 as the review-or-retire discharge of
a backlog item opened S030.)
Q2 — what preparing the repair exposed. The rank places one lead observation inside a distribution whose measured spread is 17.8×. Is the lead's overlap rate a stable property of the lead, or one draw from a spread as wide as the one it is being ranked in — and does it track the published pair's rate unit by unit?
The wire, in one sentence. The study limb restates a rank whose denominator is a measured distribution; the translation limb tests whether its numerator is a point or a distribution, by translating — for the first time in this project — prose poems of the same book on which the two published translators are independent, which is the one stratum every lead translation the project holds has missed.
2. What the stored data already shows, before this experiment runs
Five facts, all recomputed from reference.json and results.json at design time. They are
not predictions and are not scored; they are what makes Q2 worth asking, and three of them
are defects in the figures the arm exists to repair.
- The split is 13 flagged / 28 clean / 1 unchecked of 42 — not "13 flagged / 29 clean of 42"
as
wiki/arms/ARM-baseline-restate.mdstates, and not "13 / 41" asRS-20260726b-baseline-dependencestates without saying which unit is missing. - One unit was silently dropped by the dependence check.
tools/dependence_check.py'srun()skips any unit whose text is falsy (if a in texts and b in texts and texts[a] and texts[b]), andE-20260726b'sbuild.pyre-fetched and re-segmented the Gutenberg volumes rather than using stored bodies. Index 1, Garnett's A CONVERSATION, is absent from the 41 checked units, and nothing anywhere records that it is absent. - The dropped unit is the reference distribution's minimum. Its
frozenn=5 rate is 8.798, the smallest of the 42, andRS-20260726-period-control§2 names it as the least-agreeing passage. So the published max/min = 17.8× divides by an observation whose dependence status has never been checked. - The maximum is a flagged unit. THE MONK (index 39), rate 156.463, carries 4 shared 12-grams and a 14-token run. The published 17.8× is therefore max(flagged) / min(unchecked).
- «Роза» — the unit the lead's rank was measured on — is itself flagged (index 15, n12 = 6,
n15 = 3, longest run 17). And all nine lead prose-poem translations this project holds are
on flagged units: the eight of
T-senilia-R04-v1(indices 3, 4, 11, 17, 20, 27, 33, 39, selected byE-20260726cbecause they carried shared runs) and «Роза». Zero are on clean units. The project has never measured its own overlap on a unit of this book where the two Victorians are independent of each other.
Fact 5 is the reason this session has a translation limb rather than being an afternoon of arithmetic.
3. Materials
Stored, unmodified.
workshop/experiments/E-20260726-period-control/runs/reference.json— the 42 LCS-aligned Garnett 1897 ~ Hapgood 1904 unit pairs, per-unit rates at n = 4, 5, 6, 7 under both variants.workshop/experiments/E-20260726b-baseline-dependence/runs/results.json— per-unit n7 / n12 / n15 / longest-run for 41 of them.workshop/translations/senilia/R04-v1/translation.md— eight lead translations, frozen S030.workshop/translations/roza/R04-v1/translation.md— one lead translation, frozen S027.
Source for the translation limb. Six complete prose poems, runs/sources-ru.md, 898
Russian words, committed in this same commit. Two witnesses — ru.wikisource.org and
rvb.ru (PSS, Наука 1982, т. 10) — collated word-for-word: identical on all six, no
emendation needed (F1 met and recorded).
| # | idx | Russian | Garnett title | Hapgood title | ru words | published n=5 frozen |
n12 |
|---|---|---|---|---|---|---|---|
| 1 | 24 | Корреспондент | THE REPORTER | THE CORRESPONDENT | 108 | 12.903 | 0 |
| 2 | 21 | Щи | CABBAGE SOUP | CABBAGE-SOUP | 207 | 23.891 | 0 |
| 3 | 14 | Черепа | THE SKULLS | THE SKULLS | 148 | 31.414 | 0 |
| 4 | 41 | Молитва | PRAYER | PRAYER | 130 | 46.980 | 0 |
| 5 | 13 | Воробей | THE SPARROW | THE SPARROW | 199 | 56.604 | 0 |
| 6 | 34 | Завтра, завтра | TO-MORROW! TO-MORROW! | TO-MORROW! TO-MORROW! | 106 | 103.704 | 0 |
The Garnett↔Russian mapping was established from titles only — Hapgood's title is a second independent English witness to it in all six rows — and no poem body in any English was fetched, opened or printed before the translations are committed.
Comparators, deliberately not yet present. Project Gutenberg 8935 (Constance Garnett, Dream Tales and Prose Poems, 1897) and 15994 (Isabel F. Hapgood, A Reckless Character and Other Stories, 1904). Both public domain. Fetched only at step 4 below, after the six translations and their logs are committed.
4. The selection rule, frozen and mechanical
Applied to the 42 rows of reference.json['lcs/frozen']['rows'], in this order:
- Exclude the 9 units the lead has already translated — indices 3, 4, 11, 15, 17, 20, 27, 33, 39.
- Exclude flagged units (n12 > 0). The target stratum is the one the project has never had.
- Exclude index 0 (THE COUNTRY). Its 9-token Garnett~Hapgood run was printed into this session's own context while the split was being tabulated. Recorded rather than ignored.
- Exclude index 1 (A CONVERSATION). Its dependence status is unchecked at design time and it is a subject of the study limb; it may not also be a subject of the translation limb.
- Length filter:
min(tokens) ≤ 450, so six units fit one session honestly. - Sort the survivors ascending by n=5
frozenrate; take the units at positionsround(k·(N−1)/5)for k = 0…5.
N = 21 after step 5. The six selected span 12.903 → 103.704, against a clean-stratum range of 8.902 → 103.704: the rule reaches the top of the clean stratum and stops one unit short of its bottom, because that unit is index 0.
Exposure register, kept because step 3 only works if it is complete. The only Garnett or
Hapgood word-sequences printed into this session before the translations are committed are: the
9-token run of index 0, and the 18-token run of index 4, quoted in
RS-20260726b-baseline-dependence §Where the runs sit. Both indices are excluded. No sequence
from any of the six selected units, or from any of the 42 units' bodies, has been printed.
5. Statistic — inherited, not chosen here
Frozen at E-20260725c-contamination-sweep §4 and used unchanged by S026–S028:
shared n-gram types per 1000 tokens of the shorter text.
n = 4, 5, 6, 7; primary n = 5. Two parameter-free extremes, both reported: frozen (the
aggressive proper-name rule of tools/ngram_overlap.name_tokens) and none (no exclusion).
Tokenisation is tools/ngram_overlap.tokenise, unmodified.
Inherited reading rule (RS-20260726-period-control §7): a prediction true at one extreme
and false at the other is FAILED, not mixed.
Rank = how many of the reference observations a rate strictly exceeds, computed by the
construction E-20260726-period-control/percentile.py already uses, tolerance 1e-6, reported as
k/42 (or k/N_clean) and never as a percentage, so self-membership stays visible.
6. Procedure
- This design, the six Russian sources and the source manifest committed. No English present.
- Independent pre-run critic pass (§9). Amendments recorded in §A of this file before any translation is written and before any comparator is fetched.
- Translate the six from the Russian alone, regime
R04-lead-close: draft, frozen and committed as its ownR06-v1artifact (R04 §2a), then self-revision committed asR04-v1, each with a translator's log frozen at translation time. - Fetch PG 8935 and 15994. Extract all 42 units with the unmodified S027 extractor. Reproduction gate F3.
- Compute rates and ranks at n = 4…7, both variants, for: the six new units, the eight
T-senilia-R04-v1units, and «Роза» re-derived. - Check index 1's dependence with
tools/dependence_check.py, unmodified. Finalise the clean subset. Recompute every published rank over it. - Amend the three result pages — both figures side by side, erratum block naming what
changed, nothing discarded (
ARM-baseline-restate§Constraints: amend, do not rewrite). - Independent verifier recomputing every reported number from the stored bodies, plus a mutation test.
7. Predictions — registered before any comparator English exists
This table is post-critic and is the scored one. The pre-critic table had nine predictions, of which the critic established that two could not fail and one was confounded with the selection rule; §A records what each became. Eight are scored; two demoted items are sanity checks in §8. The count went down, not up.
| prediction | scored on | |
|---|---|---|
| P1 | Across the 15 lead units, the lead~Garnett n=5 rate varies by ≥ 3.0× max/min, under both variants | is the numerator a point? |
| P2′ | The median ratio (lead~Garnett rate ÷ published Garnett~Hapgood rate) at n=5 is LOWER on the 9 flagged units than on the 6 clean units, under both variants | the dependence reading, tested on a quantity the selection rule does not fix (A3) |
| P3 | The lead~Garnett rank on the 6 clean units, against the all-units distribution, has median < 41/42 — «Роза»'s figure is not reproduced on clean material | is 41/42 typical? |
| P4′ | Spearman ρ between lead~Garnett rate and published Garnett~Hapgood rate across the 15 lead units is ≥ 0.5 and p < 0.05 by permutation (100,000 permutations, seed 20260729), under both variants | does overlap track the material? (A4) |
| P6′ | On the clean subset, «Роза»'s lead~Garnett n=5 rate exceeds every clean observation (k = N_clean), under both variants | can fail; the fraction version could not (A1) |
| P7 | Median published n=5 frozen rate of the 13 flagged units exceeds that of the 28 clean units |
if so the clean subset is not a random subsample and the prescribed repair cannot be neutral |
| P8′ | Lead self-estimate (note (jj), 1 of 3 to date): on ≥ 4 of 6 clean units the lead's rank is still ≥ 30/42 | the lead's own guess, scored — causal clause struck (A6) |
| P9 | Index 1 (A CONVERSATION) is clean (n12 = 0) | the unchecked minimum |
Control C1, added by the critic (A3), reported descriptively. Within the overlapping published-rate band — the interval covered by both strata's lead-translated units — compare the median lead~Garnett n=5 rate on flagged against clean units. The band and both n are computed and printed. With single-digit n on at least one side this is descriptive and is not a result, in either direction; it exists so that a difference between strata cannot be read as a dependence effect without the matched-band figure beside it.
8. Failure criteria
- F1 — source. Two witnesses must agree word-for-word on a unit or it is dropped. Met: 6 of 6 identical.
- F2 — extraction. The six lead translations must be extractable by unmodified
tools/ngram_overlap.extract. If not, the rates are not commensurable with S027's and nothing is reported. - F3 — reproduction gate. The re-fetched volumes must reproduce
reference.json's stored per-unit token counts exactly for every unit a rank is reported on. Any mismatch → no rank is reported for that unit, and the mismatch is reported instead. - F4 — self-membership. Each unit's own Garnett~Hapgood observation sits inside the
reference distribution. Every rank declares this and is reported as
k/N, never as a percentage. - F5 — power. Six clean and nine flagged units is not a sample of the book. A prediction firing under one variant only is FAILED. Any stratum effect is reported with n stated and no claim of generality beyond this translator pair and this book.
- F6 — enumeration completeness (added by the critic, A5). The dependence results must
contain exactly 42 keyed units, enumerated against
reference.json, before the clean subset is finalised. Any missing unit blocks every downstream rank and is reported as the finding. This gate exists because §2 fact 2 is precisely the defect of a check that skips silently, and the design's step 6 had left the same door open.
Sanity checks — registered, not scored (demoted from the prediction table by the critic, A2):
- S1. On the clean subset the published max/min at n=5
frozenis > 3.0. Already true at freeze from figures printed in §4 (8.902 → 103.704 = 11.65×), hence not a prediction. - S2. The two source witnesses agree word-for-word on all six units. Already met (F1).
9. Critic pass
One independent non-Anthropic panel model, role per config/models.md. No model is a subject
in this design — every measurement is arithmetic over texts plus the lead's own translation —
so the role-collision constraint that has governed S053–S059 does not bind, and the critic is
chosen for depth of structural criticism instead. Verdict, findings and the disposition of each
finding recorded in critic.md; amendments recorded in §A below before step 3.
10. What this cannot license, stated in advance
- Nothing about memory or independence. The lead is measurably contaminated on
Turgenev–Garnett — a 21-token verbatim run,
RS-20260728b-forced-run-ru. This design measures overlap; overlap is its measurand, not a confound it controls for.CLAUDE.md's standing selection gate is therefore inapplicable in its usual role, and that is declared here rather than waived silently. - Nothing about quality, in any direction, for any of the eleven texts involved. Overlap is not merit. The lead does not judge its own translation (charter §5).
- No general baseline. One translator pair, one book, one genre, one language pair.
- Not a random sample. The 15 lead units are stratified on dependence status and on published rate by construction. The marginal distribution of the lead's ranks is not an estimate of what a random 15 would give; the primary statistics are the paired comparison (P4) and the stratum contrast (P2).
- Length. Units run 135–434 tokens. Short texts give noisier rates, and the rate normalises by length without removing that. Reported alongside every figure.
- Every number here is
internal-judgment-onlyandprovisional: Tier D has not passed, no jury is involved, and this is arithmetic over texts.
A. Amendments after the critic pass
moonshotai/kimi-k3 (P4), NEEDS-AMENDMENT, six findings — three BLOCKING, three MANDATORY,
all six accepted, five by changing the design and one on the first of its two offered fixes.
Full record and disposition: critic.md. Written before step 3; no translation existed and no
comparator English was present in the container when this section was written.
- A1 (finding 1, BLOCKING). P6 struck — it could not fail. Removing 13 flagged units from a 42-unit denominator, when the lead's rate already exceeds 41 of them, can only leave the fraction equal or higher whatever the truth is. Replaced by P6′, which asks whether the lead's rate exceeds every clean observation and can therefore come back no. The fraction change is restated as an arithmetic identity in the result, not as a finding, and the arm's "the direction of the error is conservative" claim is scored on P2′ instead.
- A2 (finding 2, BLOCKING). P5 struck and re-registered as sanity check S1: its answer is printed in §4 of this design. Not replaced — restoring the count to nine would be the tally-inflation the finding objects to.
- A3 (finding 3, BLOCKING). P2 replaced by P2′. The nine flagged units were selected by
E-20260726cbecause their published agreement was high, so a raw lead-rate comparison between strata passes on the selection rule alone. The scored quantity becomes the ratio lead ÷ published, which the selection rule does not fix: if dependence inflates the published rate at flagged units while the lead — which read neither translation — is driven only by the source, the ratio must come out lower on flagged units. Control C1 (matched published-rate band) added and marked descriptive. - A4 (finding 4, MANDATORY). P4 → P4′: ρ ≥ 0.5 and permutation p < 0.05, 100,000 permutations, seed 20260729, under both variants. The stratified construction of the 15 is reported beside it so the inflated range stays visible.
- A5 (finding 5, MANDATORY). New gate F6, §8: exactly 42 enumerated units or every downstream rank is blocked.
- A6 (finding 6, MANDATORY). P8 → P8′: the causal clause "because convergence on plain Turgenev prose is mostly forced by the source" is struck, leaving the bare numerical prediction. The finding's second offered fix — add a control limb bounding the memory contribution — is declined, in writing and with a reason. It is a second experiment; the arm's declared budget is one session; and a limb on other Russian prose would have a different comparator and so would not be commensurable with this reference distribution, so it could not settle the memory question either. §10's refusal stands and is now the only place the question is addressed.