Repository path: workshop/experiments/E-20260726c-forced-or-borrowed-ru/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260726c-forced-or-borrowed-ru |
| status | frozen |
| created | 2026-07-26 |
| updated | 2026-07-26 |
| senses | accuracy, style-correspondence |
| internal-judgment-only | true |
| provisional | true |
| links | workshop/experiments/E-20260726b-forced-or-borrowed/design.md, wiki/findings/results/RS-20260726b-forced-or-borrowed.md, wiki/findings/results/RS-20260726b-baseline-dependence.md, workshop/experiments/E-20260726-period-control/design.md, tools/dependence_check.py, tools/ngram_overlap.py, workshop/regimes/R04-lead-close.md |
Frozen design — forced or borrowed, second language pair: Turgenev's prose poems
Frozen 2026-07-26 (S031). Nothing in §§1–9 was written after seeing a single recall number, and nothing in §§1–9 was written after reading any English rendering of any prose poem this design will have the lead translate. §10 holds dated amendments; §11 is written after the run.
The wire, in one sentence. S030 found, on Latin→English, that where Riley 1851 and Brookes More 1922 share a verbatim run the Latin does not push an independent translator toward that wording — every FORCED prediction failed; this session's translation limb runs the same instrument on a second language pair, Russian→English prose, by having the lead translate eight complete Turgenev prose poems from the Russian alone and asking whether it lands on Garnett's and Hapgood's shared wording more than on Garnett's adjacent wording in the same poem.
1. Why this runs
NEXT.md action 4. S030's finding is one pair, one poem, one language pair, and its §7 said so: "The mechanism, if found, is general only as far as the argument for its generality goes, and that argument is not measurement." Two of the three baselines RS-20260726b-baseline-dependence flagged are not Latin→English. Garnett 1897 ~ Hapgood 1904 on the prose poems is the second-largest flagged cell (51 shared 12-grams over 41 comparable units, longest run 18 tokens) and is the cheapest replication available, for four reasons stated before any locus existed:
- Both English texts are public domain (Garnett 1897 PG #8935; Hapgood 1904 PG #15994), so runs may be quoted in the result page and the journal.
- The Russian is free and the project has already used it (
rvb.ru, reproducing ПСС vol. 10;ru.wikisource). - The unit is a complete short work. This is the design gain over S030 and the reason this replication is worth running rather than merely repeating: S030's largest residual was interpolation error inside a 29-line marker segment (its A22 recorded the asymmetry that F3(b) existed to catch). Here there is no alignment step at all — the lead translates the whole poem, so the Latin-side question "does the window contain the material the target run renders?" is answered yes by construction. S030's F3 is not weakened here; it is unnecessary.
- The same-translators-other-work control already exists and is already clean: Garnett~Hapgood share 51 runs across the prose poems and zero on «Свидание» and on A House of Gentlefolk (
RS-20260726b-baseline-dependence). So whatever is happening is about this book.
2. The question
Q. At the loci where Garnett (1897) and Hapgood (1904) share a verbatim run of ≥12 tokens, does an independent translator of the same Russian reproduce that wording more than it reproduces the wording elsewhere in the same poem by the same translator?
Read exactly as S030 §2 and A19 read it, and no wider: FORCED does not strictly entail a high rank, BORROWED does not strictly entail 0.5. If the shared-run loci were systematically ones the Russian constrains, an independent translator should reproduce their wording measurably more than adjacent wording; the absence of that elevation is evidence against the FORCED reading and is not proof of borrowing.
3. This is a replication, and its thresholds are not re-tuned
Every threshold in §5.5 is S030's, verbatim, scaled only by S030's own F1 rule for n < 10. No threshold was chosen after looking at this pair. That is the whole point of running a replication rather than a new experiment, and it is stated here so that a later session can check it against E-20260726b-forced-or-borrowed/design.md §5.5 line by line.
The statistic (§5.4), the stoplist, the tokenisation, the control-stride bracketing (S030 A23), the |C_i| ≥ 15 gate (A24), the covariate reporting (A20) and the reading rules (§5.6) are likewise carried over unchanged. What differs is the material, the unit definition (§4.2), and the absence of an alignment step.
4. Materials
| label | translator | year | rights | source |
|---|---|---|---|---|
ru |
— (source) | 1878–82 | public domain (Turgenev d. 1883) | rvb.ru/turgenev/01text/vol_10/02senilia/, reproducing ПСС в 30 томах, М.: Наука, 1982, т. 10; cross-checked against ru.wikisource |
garnett |
Constance Garnett | 1897 | public domain | PG #8935, Dream Tales and Prose Poems |
hapgood |
Isabel F. Hapgood | 1904 | public domain | PG #15994, A Reckless Character and Other Stories |
lead |
the lead agent | 2026 | project artifact | workshop/translations/senilia/R04-v1/translation.md (does not exist at freeze) |
4.1 The unit set, and the alignment that is not cross-lingual guessing
S027 (E-20260726-period-control) aligned Garnett's headings to Hapgood's by LCS and froze the result as runs/reference.json key lcs/frozen: 42 rows, of which one (A CONVERSATION) has an empty Hapgood body, leaving 41 comparable units. This design uses those 42 rows exactly as frozen — no fresh alignment — so the units are the same units S027's reference distribution was built on.
The Russian side is added by an explicit, committed table mapping each Garnett row index to one rvb.ru page (§4.3). The table was written from the rvb.ru page titles only — Russian titles, fetched and printed before this design was written, with no English and no poem body displayed. It is not asserted; it is gated:
- G1 — monotone. The table must be strictly increasing in rvb page number as the Garnett row index increases. Senilia has a fixed canonical order and Garnett follows it, so any transposition of two poems breaks monotonicity. This is the check that a title-by-title guess cannot pass by luck.
- G2 — length band. For every selected unit,
garnett_tokens / russian_tokens ∈ [1.05, 1.90]. Russian is inflected and drops articles, so English runs longer; the band is written before any unit was measured and a unit outside it is dropped and reported, not adjusted. - G3 — reported, not gated. Paragraph counts on both sides, per unit, printed. rvb.ru's paragraph markup is not known to be faithful, so this is a covariate.
- G4 — lead audit, post-freeze, mandatory. After the translation is committed, the lead reads each
S_ibeside its own rendering of that poem and records yes/no whether that poem's Russian contains the materialS_irenders. A locus recording no is dropped and the drop reported. If more than 2 loci are dropped the mapping table of §4.3 is judged unsound and the experiment reports a method null. This is S030's F3(b) kept even though §1.3 argues it cannot fail here — a check that the design predicts will pass is exactly the check worth keeping (note (p)), and it is the only thing standing between the mapping table and an unverified bibliographic prior (note (d), which has already fired once in building this design: the lead's prior for THE MONK was «Чернец» and the correct page is «Монах»).
G4 is a lead judgment and is internal-judgment-only. It is a judgment about which Russian a passage renders, not about quality.
4.2 The unit, and why the whole poem
The unit is one complete prose poem. The translated window is the whole poem. The measurement region M_i is Garnett's whole body for that poem, excluding the title line. There is no interpolation, no collar, and no window-versus-region distinction — §1.3.
Titles are excluded everywhere, from S_i, from C_i, and from L_i. The reason is a declared exposure: the lead has seen Garnett's 42 English titles, because the probe that selected the flagged units printed them (E-20260726c-name-tokens-repair/probe_prose_poems.py). That probe printed no poem body and no shared-run text, by construction. Title exposure is bounded, stated here, and neutralised for the statistic by excluding titles from every token stream. It is not neutralised for the translation, and §7 records that.
4.3 The mapping table (frozen)
Garnett row index → rvb.ru page, for the 42 frozen rows. Written from Russian titles only.
0 0217 1 0218 2 0219 3 0220 4 0221 5 0222 6 0224 7 0225 8 0226 9 0227
10 0228 11 0229 12 0230 13 0231 14 0232 15 0234 16 0235 17 0236 18 0239 19 0240
20 0241 21 0242 22 0243 23 0245 24 0246 25 0247 26 0248 27 0249 28 0250 29 0251
30 0252 31 0253 32 0254 33 0255 34 0256 35 0257 36 0259 37 0261 38 0263 39 0264
40 0265 41 0266
5. Procedure
5.1 Locus selection (select.py, written after this freeze)
- Rebuild the 42 frozen units exactly as
E-20260726b-baseline-dependence/build.pydoes, by importingE-20260726-period-control/extract.pyandruns/reference.json'slcs/frozenrows. - Compute, per unit, every maximal shared run of ≥12 consecutive Garnett tokens occurring contiguously in Hapgood's body of the same unit. Maximal and contiguity are S030 A10's rule verbatim: extend while the next 12-gram also occurs in Hapgood, then verify the extended run occurs contiguously in Hapgood and shorten from the right one token at a time until it does. An assertion fails loudly if any retained
S_idoes not occur contiguously in both bodies. - A unit is flagged if it has at least one such run. Expected: 13 of 42, with total 51 shared 12-grams and maximal run 18 — asserted, and a mismatch fails loudly (note (vv)).
- Exclude row 15, THE ROSE. The lead translated «Роза» in S027 (
T-roza-R04-v1) and that experiment fetched and scored Garnett's and Hapgood's renderings of it. Its contamination is not merely suspected but recorded. Excluding it is not optional. - Order the remaining flagged units by descending longest shared run, ties broken by ascending Garnett row index. Take the first 8. Deterministic; no randomness anywhere in this design.
S_iis that unit's longest maximal shared run; ties broken by earliest Garnett start position.C_iper §5.4, overM_i= Garnett's body minus the title.- Emit
runs/sources-ru.md— the Russian only, one section per unit, with the rvb page id — andruns/key.json, holdingS_i,M_i, the control spans and both published bodies.key.jsonis committed one commit beforetranslation.mdexists, so the record shows the answers predated the attempt (note (yy)). The lead does not openkey.json, or any file containing Garnett's or Hapgood's English, until the translation and its log are committed.
Why 8 and not 12. Twelve flagged units survive step 4 and the cap is a budget on lead translation effort, not on evidence. It is set here, before selection, at the number the session can translate at the depth R04 requires — 8 units, ~2,400 Garnett tokens, comparable to S030's 140 hexameters. The cost is power, and §7 states it: eight loci with S030's thresholds is a weaker test than ten, and the four units dropped (THE FOOL, THE TWO BROTHERS, THE EGOIST, WHAT SHALL I THINK?) are dropped by a pre-committed rule, not by inspection. Per S030 A21(iv) this is not a random sample of the pair's shared runs — it is biased toward the longest ones, and §7 records what that does.
5.2 Translation (regime R04, T-senilia-R04-v1)
The lead translates all 8 poems from runs/sources-ru.md — the Russian alone — in session, at no API cost, per workshop/regimes/R04-lead-close.md.
- Prose. The source is prose; no form decision arises. (S030 had to declare prose against a verse source and defend it; that does not apply.)
- No published rendering of these poems is read before the freeze, in any language. Russian dictionaries and Russian-side commentary are permitted.
- "In session" is S030 A18's rule verbatim: one drafting pass and one self-revision against the Russian, no other rendering consulted, and the lead may see its own earlier renderings because they are in the same file.
- The translator's log is written at translation time and frozen with it.
5.3 Predictions, registered — S030's, scaled by S030's F1
n is the number of loci surviving §5.1 and G4. n = 8 is expected, so ceil(0.9n) = 8 and ceil(0.3n) = 3.
- P1 — instrument resolution (gate). (a) In ≥8 of 8 units the control recalls
{recall(C, L_i) : C ∈ C_i}have interquartile range > 0. (b) In ≥8 of 8 units,median{recall(C, L_i) : C ∈ C_i}strictly exceedsmedian_{j≠i}( median{recall(C, L_j) : C ∈ C_i} ). If P1 fails, no verdict on P2–P7. What P1(b) licenses, and all it licenses (S030 A6): recall responds to which passage was translated. - P2 — FORCED, sign.
r_i > 0.5in ≥8 of 8 loci. Exactly 0.5 counts as a failure. - P3 — FORCED, central magnitude.
median(r_i) ≥ 0.75. - P4 — FORCED, absolute magnitude.
mean_i recall(S_i, L_i) − mean_i median(recall(C, L_i)) ≥ 0.15. - P5 — FORCED, verbatim landing.
v_i ≥ 4in ≥3 of 8 loci. - P6 — BORROWED.
median(r_i) ∈ [0.35, 0.65]and|P4's margin| < 0.05. P2, P3, P5, P7 are reported alongside but do not enter P6. - P7 — FORCED, phrasing.
median(r_i^tri) ≥ 0.75. Not scored, and reported as uninformative, if more than 80% of all target and controltrivalues across all loci are exactly 0.
The case each prediction should fail (note (p)). P2–P5, P7 should fail if the shared runs are ordinary wording. P1 should fail if the loci are so short, or the renderings so divergent, that recall carries no information. P1(a) is the guard against note (uu)'s saturation; note (ww) applies and is satisfied by construction — S_i is 12–18 tokens and every control is exactly |S_i| tokens from the same poem by the same translator.
Registered ex ante, before any locus was selected: this design's authors expect P2–P5 and P7 to fail, because S030 found they failed on Latin→English. A replication that expects the null must say so in advance, and must state what would change its mind: P3 ≥ 0.75 here, on Russian prose, would mean S030's result does not generalise and that the FORCED reading is alive on at least one pair.
5.4 The statistic, tokenisation and stoplist — carried over verbatim
recall, tri, recall_nonames, v_i, the mid-rank percentile r_i, the three control sets (C_i^all primary, C_i^nomore, C_i^matched), the stride bracketing at 1 (primary) and 3, the |C_i^all| ≥ 15 gate on the stride-1 set, and the 100-item stoplist are exactly as E-20260726b-forced-or-borrowed/design.md §5.3, §5.4, A1, A2, A3, A4, A7, A8, A23, A24 define them, with More read as Hapgood and Riley as Garnett throughout.
Two substitutions are forced by the material and are stated rather than left implicit:
C_i^nohapgoodreplacesC_i^nomore: control spans sharing no 8-gram with Hapgood's body of that unit. S030 A1's finding stands and is repeated: this filter selects on lexical similarity to the second translator, biases toward FORCED, and is therefore a sensitivity check and never the primary.recall_nonamesresolves names bytools/ngram_overlap.py::name_tokens([garnett_body, hapgood_body])— the lead's rendering excluded from the resolution corpus (S030 A3). This session repaired that function (three defects;wiki/findings/results/RS-20260726c-name-tokens-repair.md), so the figure here is computed with the repaired rule and is not comparable to S029's or S030'srecall_nonamesfigures. Stated here because it would otherwise look like a discrepancy.
Tokeniser version freeze (S030 A11's practice). tools/ngram_overlap.py SHA-256 b9bb8cec89e4abd26d87015ee7fbb031b23f0bbfd50800fd66e43bb86a141757 (post-repair); tools/dependence_check.py SHA-256 4af970fc9d62b5cb037db6bf7d39f18100f63e97077c8b6c14a257631152a4b1. Both re-checked by verify.py; a mismatch fails loudly.
L_i extraction from translation.md (S030 A12, adapted): the body of the section whose heading matches exactly ## Unit <i> — <rvb page> <Russian title>, from the line after the heading to the next line beginning ## or a line equal to ---, excluding any line beginning with >, excluding the first non-blank line if it is the lead's English title, and excluding everything from ## Translator's log onward.
5.5 Failure criteria
- F1 — fewer than 8 loci. Thresholds become
≥ ceil(0.9n)and≥ ceil(0.3n). Ifn < 6the experiment reports a materials null and no prediction is scored. - F2 — P1 fails. No verdict on P2–P7. Per-locus values still reported.
- F3 — G1/G2 fail. A unit outside G2's band is dropped and reported. G1 failing at all is a method null for the whole run: the mapping table is then not trustworthy anywhere.
- F4 — a unit's rendering is missing, empty, or shorter than 20 tokens. Dropped, reported.
- F5 — contamination. If
L_ishares a ≥12-gram with Garnett's or Hapgood's body of that unit, the locus is retained and flagged prominently. That is data about the lead, not a defect in the instrument. - F6 — G4 drops more than 2 loci. Method null (§4.1).
5.6 Reading rules, registered
S030 §5.6 verbatim, with one addition for the replication:
- P1 fails → method null; nothing claimed about FORCED or BORROWED.
- P1 passes, P2–P5 and P7 pass → evidence that the shared runs sit where the Russian constrains an independent translator, which would weaken the dependence reading for this pair.
- P1 passes, P6 holds → evidence that the shared runs are ordinary wording one translator reproduced from the other. This does not prove borrowing, and the result page must say so in those words.
- Mixed → per-locus values, no aggregate verdict.
- Replication reading (new). Whatever the outcome, the result page states the comparison to S030 as a two-point comparison between two language pairs, not as a general finding. Two agreeing points are two points. If the two disagree, neither is discarded and the difference is reported as an open question, not resolved by preferring the newer.
5.7 Verification
An independent verify.py, written from §§4–5 and not importing select.py or analyse.py, recomputes: the 42-unit rebuild and its token counts; the flagged set; the maximal shared runs; every L, M_i, S_i, C_i; every recall, r_i, tri, v_i; and every prediction verdict.
Assertions that must fail loudly (note (vv)):
- exactly 42 rows rebuilt, exactly 41 with both bodies non-empty;
- exactly 13 flagged units, total 51 shared 12-grams, maximal run 18 tokens;
- exactly 8 loci selected, row 15 not among them;
- every
S_ioccurs verbatim and contiguously in both Garnett's and Hapgood's body of its unit; |C_i^all| ≥ 15at stride 1 for every retained locus;- the mapping table is strictly monotone (G1) and has 42 distinct pages;
- the 8 section headings are present and unique in
translation.md.
Note (xx) applied as a standing requirement, not an afterthought: verify.py prints the resolved name set for every unit and the first and last control span of every locus. A second implementation cannot see a defect in a function both implementations call; printing the intermediate objects is what caught the name_tokens defect, and this session repaired that function, so its output is exactly what needs eyes on it.
6. Pre-flight cost
One call: the independent pre-run critic pass on this design (P1, per config/models.md). Estimate built from per-call maxima (note (m)): in ~8,500 / out ≤ 4,500 → central $0.088, worst case $0.090 at list. Today's ledger stands at $0.232256 of $5.00; the worst case fits with ~$4.68 to spare. Everything else — the translation, the fetching, the selection, the analysis and the verification — is lead work at $0.00.
Per note (zz), the key-usage cross-check snapshot is taken at session end, not immediately after the call, and an empty response body is retried with every attempt written to its own file.
7. What this design cannot establish
- Direction. Nothing here says who copied whom. Hapgood 1904 came after Garnett 1897; the words say nothing.
- That Garnett~Hapgood is dependent. The same limit S030 stated. This measures whether the source forces the shared wording; it does not measure who read whom.
- Generality. Two language pairs. §5.6's replication reading rule exists to stop this being written as more.
- A random sample. Selection is the 8 longest shared runs among 12 eligible units. Longest runs are the least likely to be forced and the most likely to be memorable, and the bias could run either way. Nothing is claimed about the class of all 51 runs.
- Contamination-freedom of the instrument. Declared high on provenance: Garnett's Turgenev is among the most reproduced English translations of the nineteenth century and any language model's training corpus contains it. If the lead remembers Garnett, recall is inflated — and inflated toward FORCED, i.e. against the outcome §5.3 says this design's authors expect. That direction is stated here so it cannot be claimed later as a convenience. F5 measures it.
- Title exposure is not neutralised for the translation. §4.2 removes titles from the statistic. It cannot remove them from the lead's memory of having read forty-two of Garnett's title choices. The effect on a 12–18-token run in a body is expected to be small and is not measured.
- Anything about quality. No sense of
wiki/goodness-senses.mdis scored.senses:recordsaccuracyandstyle-correspondencebecause those are the senses a claim about "the Russian admits one rendering" would touch. Every judgment on this page isinternal-judgment-only.
8. Artifacts
design.md(this file, frozen),critic.mdselect.py,analyse.py,verify.pyruns/sources-ru.md(Russian only),runs/key.json,runs/results.json,runs/verification.jsonworkshop/translations/senilia/R04-v1/translation.md(the translation limb, with its frozen log)wiki/findings/results/RS-20260726c-forced-or-borrowed-ru.md
9. Freeze
Frozen 2026-07-26 before select.py existed, before any locus was chosen, and before any English rendering of any candidate prose poem had been read by the lead. What had been read when this was written: the rvb.ru Russian titles of 89 senilia pages, and the probe output of E-20260726c-name-tokens-repair/probe_prose_poems.py — Garnett titles, token counts, n-gram counts and run lengths, no text.
10. Amendments
(dated; each records what forced it. All of A1–A16 were applied 2026-07-26, before select.py was run, before any locus, target run, control span or recall number existed. The critic pass that forced them is critic.md; the raw response is runs/critic-P1.json.)
A1 (2026-07-26, critic TASK A — the (rr)-class finding, and it invalidates §3 as written). §3 said the thresholds are S030's "verbatim, scaled only by S030's own F1 rule", and treated that as sufficient for a replication. It is not, and the claim is struck. S030's control spans came from a measurement region of ~91 translator tokens (its ten regions were 57–108, median 91); §4.2 as frozen replaced that with the whole poem, 146–623 tokens. The critic's point is exact: r_i = 0.5 does not have the same null meaning under a different control population, and the direction of the change is indeterminate — distant, topically unrelated spans usually depress control recall and inflate r_i, while Turgenev's refrains and aphoristic closings can raise control recall and deflate it. Keeping the number the same does not keep the test the same.
Consequently the control geometry is bracketed, not chosen (note (kk)), and the primary is the one under which the carried-over thresholds retain their meaning:
M_i^local— the registered primary. The 91 consecutive Garnett body tokens centred onS_i(clipped to the body; if clipping shortens it, extended at the other end where the body allows). 91 is not a free parameter: it is S030's median measurement-region size in translator tokens, and it is chosen for exactly that reason — to reproduce S030's control geometry so that P2–P7 mean here what they meant there.M_i^poem— reported as a co-primary of equal standing, not a sensitivity check. The whole body minus the title, i.e. §4.2 as frozen.
Every r_i, and every prediction verdict, is reported twice, once per geometry. Where the two disagree, the run reports the disagreement and assigns no aggregate verdict — the same rule §5.6 already applies to mixed outcomes. The comparison to S030 is made on M_i^local only, and the result page must say so wherever it compares.
A2 (2026-07-26, critic TASK A). P1(b)'s licence is narrowed, because as written it was false. With every control inside the poem the rendering is of, P1(b) can pass merely because L_i shares that poem's topic, names and diction with every span in it. What P1(b) licenses, and all it licenses: recall responds to which POEM was translated. It does not establish that recall resolves the target locus against its neighbours, and no such claim is made. This is a weaker licence than S030's A6 gave, and it is weaker for a reason the material forces.
A3 (2026-07-26, critic TASK A). §1.3's "the Latin-side question ... is answered yes by construction" is struck. It is answered only if the mapping table of §4.3 is correct. G4 is the check, and a check is not a construction. What §1.3 may claim, and now claims: there is no interpolation step, so S030's interpolation error cannot arise; mapping error can, and G1/G1b/G2/G4 are what stand against it.
A4 (2026-07-26, critic TASK A). §1.4's "So whatever is happening is about this book" is struck as unsupported. Zero shared runs on two other works is consistent with a book-specific mechanism and does not establish one: opportunity counts, edition, extraction, unit matching and the ≥12-token threshold all differ between works. Replacement wording: the same two translators show zero shared 12-grams on two other works, which is what a book-specific mechanism would look like and is not by itself evidence of one.
A5 (2026-07-26, critic TASK A/G). §4.2's "neutralised for the statistic by excluding titles from every token stream" is struck. Excluding title tokens removes them from S_i, C_i and L_i; it does not neutralise title-induced choices in L_i, and openings, motifs and closing refrains are exactly where a title would act. §4.2 now says: titles are removed from the token streams; their effect on the lead's diction is not removed and is not measured.
A6 (2026-07-26, critic TASK B). Position covariates, added to S030's A20 set and reported per locus: the relative offset of S_i in the body (0 = first token, 1 = last), whether S_i lies in the first or last 15% of the body, and the same two statistics summarised over C_i. The critic's mechanism is specific and plausible — openings introduce the title motif and closings are aphoristic or refrain-like — and a covariate that is reported cannot be assumed away. If S_i's mean relative offset differs from C_i's by more than 0.15 at a majority of loci, that goes in the result page's headline, on the same rule A20 sets.
A7 (2026-07-26, critic TASK C — a selection confound §7 missed). Longer poems contain more 12-gram positions and therefore more chance of an extreme longest run, so ordering by descending longest run selects on length and opportunity as well as on run length. Reported, not corrected: (i) per flagged unit, the number of Garnett 12-gram positions (the opportunity count) and the number of maximal shared runs; (ii) the rank correlation between body length and longest run over all 13 flagged units; (iii) the token-length distribution of the 8 selected units against the 4 dropped ones; (iv) whether row 15 would have entered the top eight had it not been excluded, and what its exclusion did to that distribution. §7's "the bias could run either way" was too vague and is replaced by these four numbers.
A8 (2026-07-26, critic TASK C). Taking one run per unit discards multiplicity, and a unit with many moderate shared runs may say more about the pair than one with a single extreme run. The design is not changed — one target per unit keeps loci independent — but n_runs per unit is reported for all 13 flagged units, and the result page states that the selection rule is blind to multiplicity by construction.
A9 (2026-07-26, critic TASK D). G1 is insufficient: an omitted or inserted page, or one early error propagated consistently, preserves monotonicity. Two additions, both machine-checkable:
- G1b — paragraph structure. Per selected unit,
|ru_paras − ga_paras| ≤ max(2, ceil(0.30 × ga_paras)). Turgenev's prose poems are short and heavily paragraphed, and paragraph count is a cross-lingual structural fingerprint that a wrong poem is unlikely to match. A unit failing G1b is dropped and reported.ga_parascounts blank-line-separated blocks of the Garnett body. - G1c — content correspondence, recorded with its evidence. G4 is a lead judgment and the critic is right that it is circular as stated. It cannot be made non-circular this session, so it is made checkable instead: for each locus the lead records, post-freeze, one concrete correspondence between the Russian and Garnett's body — a proper name, a numeral, or a distinctive concrete image — quoting both sides. A reader who knows no Russian can check a name or a number; a reader who does can check the image. The judgment stays
internal-judgment-only; the evidence for it is on the page.
A10 (2026-07-26, critic TASK D). G2's band [1.05, 1.90] is far too broad to identify a poem, and it is no longer described as if it could. Its only job is to catch a gross mis-mapping or a truncated body. A G2 failure is not read as a mis-mapping without G1b and G1c also failing; the critic's list of innocent causes (compressed Russian, expansive Garnett, direct speech, lists, extraction differences) is recorded on the page.
A11 (2026-07-26, critic TASK E — the shared-dependency defect, which this project has already been bitten by). verify.py "not importing select.py or analyse.py" is insufficient while both call tools/ngram_overlap.py: that is precisely how S026–S030's name_tokens defect survived a 218-check verification. So verify.py implements its own tokeniser, its own name rule and its own contiguity search from the spec text, importing nothing from tools/, and asserts token-stream equality with select.py's output on every body. A mismatch is reported whether or not it changes a verdict.
A12 (2026-07-26, critic TASK E). The critic is right that hashing a repaired implementation freezes it without validating it. verify.py therefore runs tools/tests/test_name_tokens.py and fails loudly if any of its ten cases fails. Those fixtures were written before the repair, and three of them exist specifically to reject the repair that was not adopted.
A13 (2026-07-26, critic TASK G). §5.6's BORROWED reading is over-claimed and is replaced. A null elevation is also compatible with: conventional translationese of the period; a shared intermediary text; editorial intervention at the publisher; source constraint that this recall metric cannot see; an inaccurate lead rendering; and a statistic too insensitive to detect a real elevation. The result page must list these where it reports P6, and must not write "ordinary wording one translator reproduced from the other" as though the alternatives had been excluded.
A14 (2026-07-26, critic TASK G). F5 is strengthened and its directional claim softened. Reported: the longest shared token run between L_i and Garnett's body, and between L_i and Hapgood's body, at every locus, as a length — not a ≥12 flag. A ≥12-gram threshold misses remembered 4–11-token fragments and syntactic imitation entirely. And §7's "inflated toward FORCED" is downgraded: memory can act on targets and controls differently, and title exposure can prime particular poems, so the direction is argued, not guaranteed.
A15 (2026-07-26, critic TASK G). §7 gains four items: (i) one translator, one regime, one session — nothing here establishes how independent human translators generally behave; (ii) no independent accuracy assessment is performed, and a claim about what the Russian constrains depends on the lead's rendering being competent, which declaring quality unscored does not remove; (iii) the design cannot separate source constraint from constraint imposed by English literary convention, genre, period style or translator norms — S030's A21(i) said this about Riley's Victorian collocations and it applies at least as strongly to Garnett, who largely made the English convention for translated Russian prose; (iv) the ≥12-token threshold is itself a free choice inherited from note (ss), and the shared-run population it defines is not the population of all agreements.
A17 (2026-07-26, post-run, found by verify.py). §5.4 carries S030's stoplist by reference, and S030 §5.4 labels it "Stoplist, frozen verbatim, 100 items". The list printed there contains 116 items. Every implementation that has ever used it — S030's analyse.py, this session's select.py, analyse.py and verify.py — reads the printed list, so no reported number depends on the count and nothing is recomputed. What is wrong is the count in the frozen prose, and it was caught by verify.py asserting it (note (vv): assert the count of every unit you expect to extract). The correct figure is 116 and both designs should be read that way.
A16 (2026-07-26, critic TASK E/F). L_i extraction is made deterministic rather than conditional. The translation file's template mandates that each unit section's first non-blank line is the lead's English title, and that no unit body contains an internal ## heading or any line beginning >. verify.py asserts all three. The conditional rule "excluding the first non-blank line if it is the lead's English title" is struck; the first non-blank line is dropped unconditionally, and the mandate is what makes that safe.
11. Run record
Run 2026-07-26 (S031). Full reading: wiki/findings/results/RS-20260726c-forced-or-borrowed-ru.md.
- Selection. 42 rows rebuilt, 41 comparable, 13 flagged, 51 shared 12-grams, maximal run 18 — all four asserted and all four matched. Row 15 excluded (contamination on record); 12 candidates; 8 loci selected across rows 4, 3, 11, 20, 17, 27, 33, 39. 1,871 Russian words emitted.
- Gates. G1 PASS (strictly monotone, 42 distinct pages). G1b PASS (paragraph counts within one on 7 of 8). G2 PASS (ratios 1.17–1.37 in a 1.05–1.90 band).
|C_all| ≥ 15PASS on both geometries. Every target contiguous in both published bodies. G4/G1c PASS 8 of 8 with a quotable correspondence recorded per locus. - F1. Rows 3 and 17 dropped by the registered rule (
|content(S_i)| < 4, sorecallundefined): their 15- and 14-token targets carry 3 and 2 distinct content tokens.n = 6, exactly at F1's materials-null floor, so thresholds are≥ 6 of 6and≥ 2 of 6. - P1(a) PASS 6/6. P1(b) PASS 6/6. P2 PASS 6/6. P3 PASS (median
r= 0.907, needed 0.75). P4 PASS (margin +0.225, needed 0.15). P5 PASS 6/6 (needed 2). P7 PASS (median tri-rank 0.958). P6 FAIL. - F5 fires at every locus and it is the headline. The lead's rendering shares a verbatim run of 11–21 tokens with Garnett at 8 of 8 loci and 13–18 with Hapgood at 8 of 8. Design §7 registered in advance that contamination inflates toward FORCED. The run is therefore reported as a contamination null: every FORCED prediction passed and not one of those passes carries evidential weight about source constraint.
- A7. The critic's predicted length/opportunity confound is absent: Spearman(body length, longest run) = −0.245 over the 13 flagged units. Row 15 ranked 2nd and would have entered the top 8.
- A17. The stoplist count carried from S030 ("100 items") is wrong; the printed list has 116. No number depends on it;
verify.py's assertion caught it. - Verification: 172 checks, 0 failures, by a verifier importing nothing from
tools/(A11) and running the name-rule fixtures inside itself (A12). - Cost: $0.091925 (critic). Translation, fetching, selection, analysis and verification $0.00.