Repository path: workshop/experiments/E-20260730g-nonlead-decisions/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260730g-nonlead-decisions |
| status | frozen |
| created | 2026-07-30 |
| updated | 2026-07-30 |
| senses | style-correspondence, accuracy |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-nonlead-log.md, workshop/regimes/R09-futabatei-form.md, workshop/canon/svidanie/manifest.md, workshop/canon/aibiki/manifest.md, workshop/translations/futabatei-honyaku-hyojun/R04-v1/translation.md, workshop/translations/yanfu-yili-yan/R04-v1/translation.md, framework/closure.md, config/models.md, config/budget.md |
E-20260730g — is there a site-level decision list in a published translator's own account, and can one be recovered where there is not?
ARM-nonlead-log step 1. Frozen 2026-07-30 (S066) before the translation limb was written and before 「あいびき」 was read.
1. Question
framework/closure.md §6 item 5: a release needs a log from a translator who is not the lead. Every coverage number the project has is the lead's self-report about the lead. ARM-nonlead-log asks whether the project can reach one, on material it already holds:
Do 嚴復's 《天演論》譯例言 (1898) and 二葉亭四迷's 「余が翻訳の標準」 (1906) yield site-level decisions — this word, these alternatives, this choice — or only a stated standard?
And one question the arm did not ask, which this design adds because the material forced it. 二葉亭's essay states a rule that is countable: preserve the source's commas and full stops. His execution of it — 「あいびき」, 1888 — is freely readable. So even if the essay reports no decisions, the rule can be applied to his own translation and the sites where it broke can be recovered. Whether those recovered sites are decisions (he could have complied and did not) or artefacts (Japanese could not comply) is the thing the control arms exist to separate.
2. Materials
| id | text | extent | provenance |
|---|---|---|---|
| YF | 嚴復《天演論》譯例言 | 7 numbered sections, 1,247 characters | workshop/translations/yanfu-yili-yan/R04-v1/source-zh.txt (in repo since S044) |
| FT | 二葉亭四迷「余が翻訳の標準」 | 16 paragraphs | workshop/translations/futabatei-honyaku-hyojun/R04-v1/source-ja.txt (in repo) |
| PC | positive control — an excerpt of a frozen lead translator's log, relabelled and stripped of context | 8 entries | drawn from an existing workshop/translations/*/translation.md §Translator's log |
| RU | Turgenev, «Свидание» (1850), complete | 69 ¶, 2,719 words | workshop/canon/svidanie/, two witnesses, SHA-256 in the manifest |
| JA | 二葉亭四迷訳「あいびき」(1888), complete | 76 ¶ (¶0 is a translator's note), 9,031 chars | workshop/canon/aibiki/, one witness |
| EN | T-svidanie-R09-v1 — the lead's English of RU ¶1–¶19 under R09 |
1,074 source words | written in session, frozen before JA is read |
Passage selection for EN, frozen before the source was read for translation: the contiguous run of paragraphs beginning at the first paragraph the project has not already translated (¶1) and extending to the end of the first dialogue movement (¶19). T-svidanie-R04-v1 (S026) holds ¶0. The run is 1,074 words, 73 sentence-final marks, 164 commas in the established Russian.
3. Conditions
A — Classification (does an account contain a decision?)
Every statement in YF and FT is assigned exactly one label from a scheme frozen here:
| label | definition |
|---|---|
| STANDARD | a criterion, aim or principle, stated without reference to a particular place in a particular text |
| ILLUSTRATION | a particular word or construction cited as an example of a general phenomenon, not as something the author chose in a translation of his own |
| DECISION | a particular place in a particular text of the author's own, with the rendering he chose |
| DECISION+ALT | a DECISION that also names at least one rendering considered and rejected |
| OTHER-TRANSLATOR | a particular feature of somebody else's translation |
| NONE | neither — biography, publication history, thanks, dating |
OTHER-TRANSLATOR exists because FT's most site-specific passages are about Zhukovsky rendering Byron, not about Futabatei rendering Turgenev. A scheme without it forces a reader to call those DECISION or STANDARD, and either is wrong. Per note (ber) the realised take-up of every label is reported, and a zero is a finding about the instrument.
Amendment A0, made before dispatch and before the critic saw the design. The unit of classification is a section (YF), a paragraph (FT) or a log entry (PC) — 7 + 16 + 8 = 31 units — and a unit can carry more than one kind of statement. The design as first frozen did not say how to resolve that, which is a gap, not a judgement call. Fixed precedence, highest label present wins: DECISION+ALT > DECISION > OTHER-TRANSLATOR > ILLUSTRATION > STANDARD > NONE. Each reader also quotes the span that earns the label, ≤ 30 characters, in the original, so a count can be checked against evidence rather than taken.
PC is a positive control (note (bdr)): a text that is a decision list by construction, presented in the same call with no marking. If a reader returns no DECISION on PC, that reader's zeros elsewhere are uninterpretable and are reported as such.
Readers: P1, P3, P5, independently, one call each, texts in the original with no lead gloss, blind to the arm and to this design. The lead's own classification is frozen in git before any reader's answer is opened.
B — Recovery (can the rule find what the account did not report?)
R09 J1 and J2 applied to JA against RU, over the whole work.
- Alignment. A monotone paragraph alignment RU↔JA, computed by dynamic programming on paragraph length (Russian words against Japanese non-whitespace characters, ratio fitted on the whole work), permitting 1:1, 1:2, 2:1, 1:0, 0:1. Every non-1:1 block is hand-verified and recorded.
- J1 (paragraph level). For each 1:1 aligned pair, does the Japanese carry the same number of sentence-final marks as the Russian?
- J2 (paragraph level). For each 1:1 aligned pair, the same number of commas?
- J2 (sentence level), secondary. Within pairs where J1 held, sentences are aligned 1:1 by position and comma counts compared.
Counting rules, frozen. Sentence-final mark = a maximal run of [.!?…] in Russian or English; 。, !, ?, or a maximal run of …/‥ in Japanese. Comma = , in Russian and English, 、 in Japanese. Dashes, semicolons and colons are not counted in any language, because Futabatei's rule names コンマ and ピリオド and nothing else. Quotation marks are not counted.
Excluded sites. The six paragraphs where the two Russian witnesses disagree about punctuation (workshop/canon/svidanie/manifest.md) are excluded from B and C — ¶0, 21, 22, 37, 49, 67.
C — The lead control (was the breach elective or forced?)
The lead translates RU ¶1–¶19 under R09 and its J1/J2 compliance is computed by the same program on the same definitions. A rule the lead cannot keep in English is not evidence that Futabatei elected to break it in Japanese — but a rule the lead can keep bounds how much of the breach is language-general.
D — The Japanese forcing probe (the control C cannot be)
C compares two different target languages, so it bounds language-general forcing only. D asks the Japanese-specific question directly.
Twelve sites, selected by a frozen rule: the first twelve aligned paragraph pairs, in source order, at which J1 or J2 failed, excluding the six variant paragraphs. For each, two readers (P2, P5) are given the Russian paragraph, its comma and full-stop counts, and Futabatei's Japanese, and asked — with no mention of Futabatei, of the essay, or of this study — whether a natural Japanese rendering with exactly those counts exists, and to write one if it does.
Reading. A site where at least one reader produces a compliant natural rendering is ELECTIVE. A site where both decline with a stated structural reason is FORCED. Anything else is UNDECIDED.
4. What is blind and what is not — stated rather than implied
Not blind. The lead read YF and FT in full, in the original, earlier in this session, before this design was written. Condition A's outcome is therefore known to the lead in outline, and the lead's own count is not evidence. A is run because the readers are the evidence; the lead's frozen count exists only so that a disagreement is visible.
Blind. B, C and D. workshop/canon/aibiki/manifest.md §Exposure records exactly what of JA was seen before the translation limb was frozen — paragraph count, per-paragraph character lengths, the first 30 characters of eight paragraphs, and the translator's note. No punctuation count was computed and no sentence of the body was read. Provable in git.
5. Predictions (registered)
| # | prediction |
|---|---|
| PA1 | Both readers' majority label for FT contains zero DECISION or DECISION+ALT. |
| PA2 | YF returns at least one DECISION+ALT (§4 of 譯例言, the 卮言 / 懸談 / 導言 sequence). |
| PA3 | PC returns DECISION or DECISION+ALT on a majority of its entries from every reader. |
| PB1 | Futabatei's J1 compliance over aligned 1:1 pairs is below 50%. |
| PB2 | Futabatei's J2 compliance, computed on pairs where J1 held, is below 30%. |
| PC1 | The lead's J1 compliance on ¶1–¶19 is above 80%. |
| PC2 | The lead's J2 compliance is below its own J1 compliance and above Futabatei's J2. |
| PD1 | ELECTIVE at a majority of the twelve sites. |
6. Failure criteria (registered)
- F1 — If the 1:1 aligned blocks cover < 90% of RU's words, B is void and only the alignment is reported.
- F2 — Any measured site falling in one of the six witness-variant paragraphs is excluded; if exclusion removes more than 15% of pairs, B's figures are reported as bounds.
- F3 — If the lead's own J1 compliance is at or below Futabatei's, PD1's inference is withdrawn and no site is called ELECTIVE on C's evidence. D still runs and still speaks.
- F4 — If a D reader declines or evades at more than 25% of its cells, that reader's D column is void; if both, D is void.
- F5 — If the three A readers disagree on the existence of a DECISION in an essay (one returns ≥1 where another returns 0), A's count for that essay is reported as a range, not a number, and PA1/PA2 are scored on the majority with the range stated.
- F6 — If any reader returns zero DECISION on PC, that reader's A column is uninterpretable and is excluded from the majority (note (bdr)).
- F7 — The realised distribution of all six A labels is printed. A label used zero times across all three readers is reported as a defect in the scheme, not as tidy data (note (ber), note (beb)).
7. What this experiment cannot show
- One translator, one pair, writing retrospectively for publication is a different self-report bias, not the absence of one —
ARM-nonlead-log§Constraints says so and it is not softened here. - A punctuation-count breach is a decision only under R09's own terms. It is not a decision in the sense a translator's log records, and B does not claim to have recovered a log. It claims to have recovered a site list with a pass/fail at each site, which is strictly less.
- Neither source text is the one Futabatei used. Both canon manifests say so. The between-witness noise floor (5 commas in 398, 1 ender in 202) bounds transcription disagreement in 2026, not distance from an 1888 printing.
- Nothing here bears on Tier D. No quality judgement is made or elicited, by anyone, about any translation.
8. Cost
Pre-flight in config/budget.md. Worst case built from max_tokens at list out-price plus prompt at list in-price (note (abc)), with a 1.5× provider premium on the critic seat because S065's identical seat billed 1.5× list (note (x), note (b)).
| stage | calls | max_tokens |
worst case |
|---|---|---|---|
| pre-run critic (P4, reserve P2) | 1 (+1) | 12,000 | $0.30 |
| A — classification (P1, P3, P5) | 3 | 8,000 | $0.20 |
| D — forcing probe (P2, P5) | 2 | 8,000 | $0.12 |
| total | 6 (+1) | $0.62 |
Declared role collision, in advance. P4 is the critic seat; its declared reserve is P2, which is also a D reader. If the reserve fires, one model both critiques the design and answers D, and that is recorded on the result page rather than discovered by a later reader. P1/P3/P5 are A readers and cannot critique.
The translation limb costs nothing and is never ledgered (charter §3, A4).
9. Amendments after the independent pre-run critic — ALL NINE FINDINGS ACCEPTED
P4 moonshotai/kimi-k3, provider Together, accepted first call, in 11,617 / out 3,916, stop, $0.0934182 — 28% of the declared worst case; the reserve did not fire, so the declared P4/P2 role collision did not occur. Verdict NEEDS-AMENDMENT: 9 findings, 3 BLOCKING, 5 MANDATORY, 1 ADVISORY. Raw at runs/critic.txt. Note (rr), twenty-fourth consecutive session in which a critic changed a design.
Two of the findings would have made this session report an artefact as its headline, and they are the two the design was most confident about.
A1 — B does not measure what §1 said it measured (finding 1, MANDATORY)
A 1906 rule scored against an 1888 translation measures the distance between a translator's later account of his method and his earlier execution, not "the sites where his rule broke". And the essay itself predicts a low score: FT7 says he often could not make practice meet the standard (「なかなか思うように行かぬ」); FT15 disowns the recent work altogether.
Accepted in full. Every claim about B is reworded to the distance between the 1906 stated standard and the 1888 execution. The headline "Futabatei broke his own rule" is not available and is withdrawn before it was written. FT7 and FT15 are registered here as a prior that low compliance is expected, so a low figure confirms the essay rather than discovering anything.
A2 — C is a control that can only confirm (finding 2, BLOCKING)
The lead translates under R09 trying to comply, knowing PC1 predicts >80%. PC1 cannot fail except by incompetence. C therefore measures achievability under motivated effort and bounds nothing.
Accepted. Two repairs, both taken.
- A second English limb,
T-svidanie-R06-v1, on the same passage, written first, in one pass, without consulting R09's rules — an unforced English baseline. J1/J2 are computed on it by the same program. - The word "control" is struck from C. C is relabelled C-R09 and is a cost measurement: what complying costs in English. The comparison that carries an inference is C-R06 against B, not C-R09 against B.
Declared, because it cannot be repaired: the lead wrote R09 before translating and cannot unsee it, so C-R06 is not blind to the rules. Direction: knowing the rule can only raise C-R06's compliance, which raises the English baseline, which makes Futabatei's figure look worse against it. That is the anti-conservative direction and it is why C-R06 is not the primary null.
A3 — the primary null is a permutation, not a threshold (finding 7, MANDATORY — the finding that would have shipped an artefact as a headline)
PB1 (<50%) and PB2 (<30%) require exact integer matches of per-paragraph counts. Under any no-effort null the exact-match rate is low anyway, so both predictions pass even if Futabatei never tried, and the session would have reported chance as a finding.
Accepted. The primary statistic becomes a departure from a permutation null: Russian paragraph counts are paired with Japanese paragraph counts at random, 10,000 times, and the exact-match rate of the real pairing is scored against that distribution. PB1 and PB2 are rewritten (§10). The permutation null is preferred to C-R06 as the primary null precisely because it is untouched by anything the lead knows.
A4 — B/C confounds, enumerated and conceded (finding 3, BLOCKING)
The critic's list is adopted verbatim into §7 and one item of it is new to this project: 「あいびき」 is a foundational 言文一致 text, so its punctuation conventions were being invented in the text being measured. Nothing in the design had noticed that. Also: the JA witness is modernised and one-witness (§7.3 conceded this for Russian and then ignored it for Japanese); translator, era, and demand characteristics all differ; C covers ¶1–19 and B the whole work.
Accepted. C is scored on the same aligned-pair definition as B, and §7 now states that C bounds nothing whatever about Japanese.
A5 — D was an editing puzzle (finding 4, MANDATORY)
Readers were to be shown Futabatei's Japanese and the target counts, which degenerates to inserting or deleting 、 until the numbers match.
Accepted, all four repairs. Readers get the Russian and the counts only — never Futabatei's Japanese — and must write a Japanese rendering from scratch. A third reader (P3) rates naturality blind to the counts and to the study, with Futabatei's own renderings mixed in unlabelled as anchors. The twelve sites are drawn by a seeded random draw over all failing pairs (random.Random(20260730)), not the first twelve in order.
A6 — the classification prompt leaked its own prediction (finding 6, BLOCKING)
"TEXT 3 — a modern translator's working notes" against "TEXT 2 — an essay about his own method" hands a reader the predicted pattern before a character of any text is read.
Accepted. All genre, date, author and language descriptors are stripped; headers become TEXT n — k units. Text order is permuted per reader by a frozen seeded schedule.
A7 — the positive control could not fail in the languages that matter (finding 5, MANDATORY)
PC is English, says "Rejected:" verbatim in four of eight entries, and is about a different literature. A reader can pass PC and still be unable to recognise a DECISION in classical Chinese, which is exactly what PA2 needs.
Accepted. Two mini-controls are added: CC — three units of classical Chinese — and MJ — three units of Meiji-register Japanese, each set containing one plain DECISION+ALT. The English PC is kept.
Declared leak: the lead wrote CC and MJ. That is the item-authorship leak RS-20260727e names. It is tolerated here and only here, because a positive control asks whether the label is reachable in a language, not whether a reader agrees with the lead — and a reader that returns zero DECISION on a unit reading 「甲」と「乙」との二つを考へて、遂に「甲」と書けり has told us something about itself, not about the lead's prose.
A8 — failure criteria that interact (finding 8, MANDATORY)
- F5 + F6 tie rule, registered: if exclusion under F6 leaves two readers who disagree on the existence of a DECISION, the count is reported as a range and the prediction is scored FAILED.
- F7 + PA1 exemption, registered: a label whose zero is predicted by a registered prediction is exempt from F7. Without this, the scheme flags itself as broken exactly when the headline prediction is confirmed.
A9 — the smaller artefacts (finding 9, ADVISORY, all four taken)
- J2 is reported twice — unconditional over all 1:1 pairs, and conditional on J1 holding — because the conditional subset is selected and easier.
- 1:0 and 0:1 alignment blocks are counted and reported separately. They are where the largest departures live and excluding them biases compliance upward.
- Registered: the statistic counts marks, not mark identity. A Russian
…matches a Japanese。if both are one sentence-final mark. This is intended; Futabatei's rule is about how many, not which. - Reserve critic seat: an A reader may serve as reserve critic only if it has not yet been dispatched as a reader. It had not been, the reserve did not fire, and the question is moot this session.
10. The registered predictions, as amended
| # | prediction (amended) |
|---|---|
| PA1 | The majority label for every FT unit contains zero DECISION and zero DECISION+ALT. |
| PA2 | YF returns at least one DECISION+ALT on a majority of readers. |
| PA3 | Every reader returns DECISION or DECISION+ALT on a majority of PC entries and on at least one unit of CC and at least one of MJ. |
| PB1 | Futabatei's J1 exact-match rate over 1:1 pairs is not distinguishable from the permutation null at p > 0.05 — i.e. the 1888 execution carries no detectable trace of the 1906 rule. (This is now the interesting prediction and it can fail in both directions.) |
| PB2 | Futabatei's J2 exact-match rate is likewise not distinguishable from its permutation null. |
| PC1 | T-svidanie-R09-v1 J1 exact-match rate > 80%. (Registered as a manipulation check, not as evidence.) |
| PC2 | T-svidanie-R06-v1 — the unforced English — has a J1 exact-match rate above the permutation null and below R09's. |
| PD1 | ELECTIVE at a majority of the twelve sites, where ELECTIVE now requires a from-scratch compliant rendering rated ≥ 3 of 5 for naturality by the blind rater. |
11. Cost, as amended
| stage | calls | max_tokens |
worst case |
|---|---|---|---|
| pre-run critic (P4) — DONE, $0.0934182 | 1 | 12,000 | $0.33 declared |
| A — classification (P1, P3, P5) | 3 | 8,000 | $0.20 |
| D — from-scratch renderings (P2, P5) | 2 | 8,000 | $0.12 |
| D — blind naturality rating (P3) | 1 | 4,000 | $0.05 |
| total | 7 | $0.70 |
Both English limbs cost nothing and are never ledgered (charter §3, A4).