Repository path: workshop/experiments/E-20260811d-title-network/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260811d-title-network |
| status | frozen |
| created | 2026-08-11 |
| updated | 2026-08-11 |
| senses | consistency, style-correspondence, voice, accuracy |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-recurrence.md, wiki/base/anchors/A-shinel-significant-person/A-shinel-significant-person.md, wiki/findings/results/RS-20260810w-recurrence.md, workshop/experiments/E-20260810w-recurrence-census/design.md, workshop/translations/shinel-kapot/R06-v1/translation.md, wiki/base/sources/S-berman-tendances.md, wiki/goodness-senses.md, config/models.md |
E-20260811d-title-network — the word that titles the story, and the one place the source refuses it
Frozen before dispatch. ARM-recurrence step 2, study limb.
1. Question
Gogol's «Шинель» repeats its title word 72 times. Once, early, it stops — and says so:
«от нее отнимали даже благородное имя шинели и называли ее капотом» (they even stripped it of the honourable name of overcoat and called it a wrapper)
A капот in 1840s Russian is a woman's loose housecoat. The clerks' joke is a demotion, from a
military greatcoat to something worn indoors by a woman, and Gogol keeps the demotion running: the
narrator himself calls the old coat a капот at four later places, including the one where Akaky
has just called it a шинель in the preceding sentence.
So the source both repeats and, deliberately, does not. Two questions follow, and the second is the one the project has never asked:
- Do published translators hold the repetition when the repeated word is the story's title?
- Do they hold the refusal — do they keep a second, distinct expression for
капот, or do they let the source's one departure collapse back into the network?
And a third, about one translator in particular: RS-20260810w §7 recorded that Field's
dissolution of the значительн- network is confounded with his text being hard to align with at
all, and could not separate them. Three new regions, one of which Field renders at 1.09 words
of English per word of Russian, can.
Subject-rule sentence (continue-prompt.md §4.5): this unit teaches what four published
literary translators do with the word that titles the story they are translating, and with the one
place where the source deliberately refuses that word. The objects measured are four published
renderings of Gogol. No statistic, rater, verifier or published figure of this project's own is the
subject; Q3 is a question about a translator's practice, tested on new material, not about a
prior number's reproducibility.
2. Why the question is owed
ARM-recurrence step 2 owes three things, and this design is the third:
wiki/goodness-senses.md'sconsistencyentry andS-berman-tendances§9 rewritten to what the project now knows — done in the same session, fromRS-20260810wand from this;- the claim step 1 states and cannot separate: Field's low
fixagainst his text's alignability. The anchor already holds the material — «шинель» is a second network in the same four hands — and step 1's own control recovered it at 8 of 8 in three cells of four.
It is also the first time the project has looked at a negative site: a place where the source
declines to use the word it has been repeating. Every measurement of recurrence the project owns
(S152's drifted arms, RS-20260810w's census, S-berman-tendances §9) counts whether the target
repeats. None asks whether the target repeats where the source stops.
3. Materials, all frozen before dispatch
Three contiguous Russian regions, cut from the anchor's stored source text
(shinel-gogol-1842-ru.txt) and frozen as materials/ru_D.txt, ru_E.txt, ru_F.txt:
| region | what it is | Russian words | шинел- |
капот- |
|---|---|---|---|---|
| D | the coat examined and stripped by the clerks of "the honourable name"; Petrovich introduced | 257 | 5 | 1 |
| E | the new coat arrives at the department; the old one dragged out for comparison | 350 | 9 | 3 |
| F | Akaky calls it a шинель and the narrator, in the next sentence, calls it a капот |
167 | 1 | 2 |
| total | 774 | 15 | 6 |
21 network sites, classified before dispatch in sites.json as NET (шинел-, 15) or
CONTRAST (капот-, 6). Regions D and E are the source unit of T-shinel-kapot-R06-v1, whose
translator's log was frozen and committed at ed89623 before this design was written.
The two words are numbered by one combined pattern (build_items.py, NET), so the marks a
seat sees do not tell it which of the two Russian words any given mark stands on.
Four published hands, each read only in the corresponding region, cut by explicit start/end
substrings written into build_materials.py before any answer existed:
| id | hand | target | date | provenance |
|---|---|---|---|---|
H1 |
Isabel F. Hapgood, "The Cloak" | English | 1886 | Project Gutenberg 1197 |
H2 |
Constance Garnett, "The Overcoat" | English | 1923 | Internet Archive overcoatothersto0000niko |
H3 |
Claud Field, "The Mantle" | English | 1916 | Project Gutenberg 36238 |
H4 |
Rudolf Kassner, "Der Mantel" | German | 1911 | Project Gutenberg 27973 |
3.1 The length covariate, declared here because Q3 turns on it
Target words per Russian word, computed from the frozen materials before dispatch, with step 1's two regions alongside:
| region | H1 Hapgood |
H2 Garnett |
H3 Field |
H4 Kassner |
|---|---|---|---|---|
| C (step 1, the portrait) | 1.23 | 1.44 | 2.20 | 1.27 |
| AB (step 1, the coda) | 1.17 | 1.33 | 1.09 | 1.52 |
| D (here) | 1.27 | 1.37 | 1.30 | 1.23 |
| E (here) | 1.24 | 1.39 | 1.09 | 1.17 |
| F (here) | 1.35 | 1.56 | 1.20 | 1.30 |
Field's region C is the one outlier in the table, at 2.20 against a 1.09–1.56 band everywhere
else — and region C is exactly where his fix was 0.250 and three seats could not agree what stood
for the epithet at 9 of 16 sites. In the three regions used here he sits at the bottom of the
band. This is stated before dispatch so that whichever way Q3 comes out, the covariate was on
the page first.
4. Procedure
The alignment is not made by the lead, exactly as at E-20260810w. For each ⟨hand, region,
task⟩ the seat receives the Russian region with the relevant occurrences numbered in place
(⟦1⟧…⟦n⟧) and the hand's corresponding target region, unlabelled and unattributed, and returns
for each numbered site the exact substring of the target that stands for it, or NONE.
- 3 seats,
config/models.mdpanel v1, non-Anthropic:J1=P1,J2=P2,J3=P5. The prompt, the system message and the call parameters are byte-identical toE-20260810w's, which returned 0 dead bodies on this task; note (bmb)'s remedy is therefore already in force (reasoning disabled where the endpoint allows it,effort: lowwhere it does not, cap 4,000). - 20 items per seat = 4 hands × (D-net, D-ctl, E-net, E-ctl, F-net); 60 calls.
- Blind to the hypothesis: the prompt names no theory, no hand, no date, does not say that repetition is the object, and does not distinguish the two Russian words.
- Control task. The same alignment for other recurring items in the same region —
портн-/чиновни-in D (5 sites),чиновни-/швейцар-in E (5 sites). Region F carries no control task;F1pools D and E, as step 1's amendmentA3requires it to pool.
5. Quantities
Per hand, from the majority answer of surviving seats (a site's expression is the string ≥ 2
seats quote, after case-folding and article-stripping), computed over NET, over CONTRAST and
over all:
fix— share of sites rendered by that hand's single most frequent expression.fix_stem— the granularity-robust variant defined atRS-20260810w§5 and imported here as a registered quantity rather than a post-hoc repair: each majority answer is reduced to the set of 6-character prefixes of its words of length ≥ 5; the hand's modal stem is the one in most answers;fix_stemis its share. The rule is mechanical and identical for every hand and language. Amended before dispatch by critic findingA2: Unicode NFD plus combining-mark deletion is applied before the prefix is taken, so thatMantelandMäntelchenreduce to the same stem; the fold is mechanical, language-independent and applied to every hand. Both the folded and the unfolded figure are reported.var, with its valid-site denominator, andttr = var / valid.del— share of sites whose majority answer isNONE.nomaj— share of sites at which no two seats agree.collapse(new, and the point ofCONTRAST) — share ofCONTRASTsites whose majority answer shares the hand's modalNETstem. A collapse is not a deletion: the hand has said something, and what it said is the same word it uses forшинель.
6. Registered predictions and failure criteria
Q1. The title network is held by every hand —fix_stemover the 15NETsites is ≥ 0.75 for all four hands. This is the positive control that can fail, and the arm's completion criterion requires one. A hand below 0.75 on the story's own title word is the finding, not a defect.Q2. The source's refusal costs more than its repetition — amended by the critic,A1below; the registered form is the amended one. For every hand,keep(the share of the 6CONTRASTsites whose majority answer is neitherNONEnor shares the hand's modalNETstem) is strictly lower than that hand'sfix_stemover the 15NETsites. Fails if any hand'skeepis ≥ itsfix_stem. Holding a word is one thing; withholding it at the one place the source withholds it is another. This is a within-hand comparison and is independent ofQ1's threshold: a hand that carried the demotion perfectly while varying the coat would break it.Q3. Field's alignability is a property of region C, not of Field —H3'snomajshare over the 21 network sites here is ≤ 0.25, against 0.5625 (9 of 16) in region C at step 1. Reported twice, over all 21 sites and over the 18 sites of D+E alone (criticA4); the registered verdict is read on all 21. Fails if it exceeds 0.25. If it passes, "this hand's text is hard to align with" does not survive as a general property of the hand, and step 1's Field figures are a statement about that region and that network. If it fails, the confound stands, andRS-20260810w'sP1must be re-read with Field's arm withheld — which would leaveP1's spread at 0.889 − 0.778 = 0.111 and below its own 0.30 bar. Both outcomes are consequential and both are written here.
Failure criteria, registered:
F1. A seat recovering < 0.75 of the 10 pooled control marks actually presented for a hand — 5 in item<hand>-D-ctland 5 in<hand>-E-ctl, no others anywhere (criticA3) — i.e. fewer than 8 of 10, has that hand's cells voided. If fewer than 2 seats survive for a hand, no quantity is read for it and the cell is reported empty.F2. If three-seat agreement on the quoted expression at network sites is < 0.70, the alignment is not reproducible: quantities are reported as descriptive counts with the disagreement rate attached, andQ1–Q3are withheld.F3. If any seat's answers contain an expression not present as a substring of the target passage it was given, that seat's cell is voided as fabricated. De-hyphenation of the passage and of the quoted string (analyse.norm_full) is applied first, identically to every hand and independently of any answer — the materials repairRS-20260810w§6 established.
7. The contamination gate, run after the translation freeze and before this design was written
tools/dependence_check.py over the three lead/published pairs and the three published English
pairs, in regions D and E (dependence.json):
| pair | region | 7-grams | 12-grams | 15-grams | longest run | verdict |
|---|---|---|---|---|---|---|
Garnett 1923 ~ lead R06 |
E | 47 | 15 | 6 | 20 | DEPENDENT? |
Garnett 1923 ~ lead R06 |
D | 12 | 1 | 0 | 12 | DEPENDENT? |
| Hapgood 1886 ~ lead | D / E | 4 / 10 | 0 / 0 | 0 / 0 | 10 / 10 | clean |
| Field 1916 ~ lead | D / E | 1 / 0 | 0 / 0 | 0 / 0 | 7 / 6 | clean |
| Garnett ~ Hapgood | D / E | 2 / 13 | 0 / 0 | 0 / 0 | 7 / 11 | clean |
| Hapgood ~ Field | D / E | 8 / 2 | 0 / 0 | 0 / 0 | 10 / 7 | clean |
| Garnett ~ Field | D / E | 1 / 0 | 0 / 0 | 0 / 0 | 7 / 6 | clean |
Two consequences, both binding.
- The lead is not a hand in this census, as it was not at step 1, and the case is stronger
here: 20 contiguous tokens against Garnett in region E — "he returned home in the happiest
frame of mind took off the overcoat and hung it carefully on the wall" — with 15 shared
12-grams and 6 shared 15-grams. That is the largest overlap this project has measured between a
lead rendering and a published human translation, against a prior record of 24 tokens on a
different work (S079).
contamination: highwas declared on the artifact, with the specific priming named, before the gate ran. - All three published English pairs are clean in every region used here (longest run 11
tokens, 0 shared 12-grams). This is the condition step 1 did not have in region AB, where
Hapgood and Garnett shared 18 contiguous tokens and forced
P1onto region C alone.Q1–Q3may therefore be read across all three regions, and the reason is measured, not assumed.
8. Amendments after the pre-run critic, all made before dispatch
qwen/qwen3.7-max (reserve slug, outside the jury), NEEDS-AMENDMENT, 4 findings — 1 BLOCKING,
2 SERIOUS, 1 ADVISORY — critic.md, $0.01485177. Three accepted, one accepted in part with the
overrule written.
A1(from BLOCKING 1, accepted in full).Q2as first written compared a count of hands againstQ1's count of hands: "a prediction that cannot fail does not test the hypothesis", and the critic is right that the two counts are coupled and that with N = 4 the integer arithmetic nearly guarantees the predicted direction.Q2is restated as a within-hand comparison —keepagainst that hand's ownfix_stem— which is decoupled fromQ1's threshold and fails on any hand that carries the demotion at least as reliably as it carries the coat. The full per-handkeeptable is reported descriptively whichever way the prediction goes.A2(from SERIOUS 3, accepted in full, with a mechanical repair rather than the critic's language-specific one). The stem rule — 6-character prefixes of words of length ≥ 5 — was imported fromRS-20260810w§5, where every hand it was applied to had been read in English or German without the case being examined. The critic's instance is real:MantelandMäntelchenreduce tomantelandmäntel, which the rule counts as different stems, deflatingH4's fixity for a reason that is about the umlaut and not about Kassner. The repair is Unicode NFD plus combining-mark deletion before stemming — mechanical, language-independent, applied identically to every hand and every answer, and written here before any answer exists. The critic's alternative (restrict the metric to English) is overruled: dropping the German hand would deleteRS-20260810w's only cross-language control. Both figures are reported, folded and unfolded, so the repair's effect on every hand is visible.A3(from SERIOUS 2, accepted as a clarification; the substantive claim is overruled). The critic reads §4 as though one call carried all three regions and asks how a seat knows where the controls live. It does not arise: each ⟨hand, region, task⟩ is a separate call whose prompt contains that region's text and nothing else, and the control marks exist only inside the two-ctlitems, which contain only region D or region E.F1's denominator is therefore exactly the 10 marks presented — 5 in<hand>-D-ctl, 5 in<hand>-E-ctl— and no region-F call carries a control mark. The wording of §4 is amended to say so; the procedure is unchanged.A4(from ADVISORY 4, accepted as a written limitation, not repaired). Region F is 167 words carrying 3 marks, and its transparency is higher than D's or E's. Adding distractor marks would change the task after the freeze, so it is not done.Q3is reported over all 21 sites and, separately, over the 18 sites of D+E, and F'snomajis read as an upper bound on alignability.
8.1 One thing the critic did not raise, and it is a real limitation
The lead is not blind to the target passages. It read parts of all four hands while locating the
region boundaries, and translated regions D and E itself; §7 records the specific renderings of
капот it saw. The predictions in §6 were registered after that reading. What protects them is
that the lead makes no alignment: every quantity is computed from the majority answer of three
independent seats, and a lead expectation about a passage cannot move a seat that never sees the
lead. The parts the lead has not seen include every hand's rendering of the two капот sites in
region E's second half, and no hand's fix_stem is known to it at all. This is stated rather than
claimed away.
9. What this cannot establish
- No jury, no scoring, no quality claim. This measures what is on the page. Whether a collapsed contrast costs a reader anything is not asked here and is not answered by it.
- Four hands, one story, two target languages, 21 sites, 6 of them
CONTRAST. Six sites is a small number andQ2is a count over hands, not a test with a null; it is reported as what four hands did, and no significance is claimed for it. Q3is a one-sided repair of one confound. Passing it removes "the hand's English is hard to align with in general" as an explanation of step 1's Field figures. It does not establish that Field'sзначительн-figures are free of every region-specific cause, and it says nothing about the other three hands' step-1 numbers.- The blinding is structural, not semantic (
RS-20260810w§7, critic findingA4there): twenty-one marks on two roots tell a competent reader that these words are the object. It is a common-mode influence — every hand is marked identically — so it cannot manufacture a difference between hands, which is what all three predictions are about. It could inflate every hand'sfixtogether, and no absolute level is claimed on that account. - Tier D is NOT PASSED; the page carries
provisional: true. No recommendation may rest on it.
10. Pre-flight cost estimate, written before dispatch
Worst case built from max_tokens and not from an expected answer length (note (abc)): 60 calls at
a 4,000-token cap with ~1,500-token prompts — J1 $0.51, J2 $0.65, J3 $0.08 — plus one critic
pass at a 12,000-token cap, ~$0.08. Declared ceiling $1.40. The identical task at
E-20260810w cost $0.182 for 48 calls, so the expectation is ≈ $0.25 and the ceiling is the cap
arithmetic, not the forecast. Today's UTC headroom before this session: $1.863005039.