Repository path: workshop/experiments/E-20260807b-botchan-register/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260807b-botchan-register |
| status | frozen |
| created | 2026-08-07 |
| updated | 2026-08-07 |
| links | wiki/arms/ARM-ja-register.md, wiki/base/anchors/A-morri-botchan/A-morri-botchan.md, wiki/base/sources/S-botchan.md, workshop/translations/botchan/R04-v1/translation.md, workshop/translations/botchan/R21-v1/translation.md, workshop/regimes/R21-vulgarisation.md, wiki/findings/results/RS-20260806g-negative-pole.md, wiki/goodness-senses.md, config/models.md |
| senses | style-correspondence, voice, naturalness |
| provisional | true |
E-20260807b — does the low pole of a Japanese narrator's register cross into English?
Frozen 2026-08-07 (S126), before any call was dispatched. Every lead rendering named here was
written and committed before a line of the comparator's chapter 1 was read (commit 72bef11).
1. The question, and why it is asked now
Two consecutive results have found the same thing on the same language pair and named it the same
way. RS-20260806f (S123) measured the largest loss in a four-language narrator-crossing run on the
Japanese item, on diction temperature; RS-20260807 (S125) repeated the run on four obscure
works and again found the largest single movement on the Japanese item, this time register
−1.75. Both times the lead's translator's log, frozen before the instrument existed, had named
register as the property being spent, and both times it said the same thing about why: English has
no slot for a register the Japanese carries in its pronouns and verb endings.
Independently, RS-20260806g (S124) produced a craft finding on Italian→English: English's
raising shelf is deep and available in narration, while its lowering shelf is almost entirely
dialogue-shaped — a translator can ennoble narration by lexis alone and cannot do the reverse by
lexis alone.
Put together, those give a claim about the language pair that no run has yet tested against published practice: if the lowering shelf really is dialogue-shaped, then a Japanese narrator whose register is low should arrive in English raised, in any competent hand, and the deferential speech around him should arrive intact. This design tests that on a published human translation, read whole against its source.
One sentence on what it teaches about literature (subject rule, wiki/tracks.md): it measures
what happens to a novel's register system — a low-register narrator set against a deferential
servant — when the novel is translated into a language with one pronoun and no deference morphology.
2. Materials
Source. 夏目漱石「坊っちゃん」(1906), chapter 1, 7,499 non-space characters, 22 paragraphs,
read whole in Japanese. Aozora Bunko 000148/752_14964, 底本『ちくま日本文学全集 夏目漱石』筑摩書房,
fetched 2026-08-07 (materials/src-botchan-ch1.txt). PD: Sōseki d. 1916.
Published comparator. Yasotarō Morri, Botchan (Master Darling), Ogawa Seibundo, Tokyo, 1918
— the earliest published English Botchan (S-botchan §Identity). PD on the US publication basis
(1918, pre-1929) and separately on the translator's life (Morri 1882–1959, d. 1959; the publication
basis is the one relied on). Internet Archive botchanmasterdar1918nats, djvu text, chapter 1
extracted whole: 4,461 words raw, materials/morri-ch1-raw.txt; furniture-stripped working copy
materials/morri-ch1.txt; the five corresponding paragraphs, 882 words, materials/morri-spans.txt.
Lead arms, all $0, all frozen and committed before the comparator's chapter 1 was opened:
T-botchan-R06-v1 (draft, 934 words), T-botchan-R04-v1 (close, 929 words),
T-botchan-R21-v1 (vulgarisation control, 769 words), over the same five whole source
paragraphs — 1,507 non-space characters, chosen because they are the five in chapter 1 that
carry both poles of the register system at once.
Contamination, measured before this design was written and after the renderings were frozen
(materials/contamination.json; tools/dependence_check.py):
| cell | shared 7-grams | 12-grams | longest run | verdict |
|---|---|---|---|---|
subject — lead R04 × Morri 1918 |
4 | 0 | 9 | clean |
| ref — an independent published pair (K&M 1915 × Garnett 1920, «Пари») | 160 | 27 | 24 | DEPENDENT? |
ref — the lead against itself, two regimes, one session (R04 × R21) |
23 | 1 | 12 | DEPENDENT? |
ref — the lead against itself, draft and revision (R06 × R04) |
558 | 420 | 75 | DEPENDENT? |
The subject cell sits below the independent-published-pair reference on every column, which is the comparison that matters: the lead's English is further from Morri's than two independent published translators are from each other. (Limit: the reference pair's texts are whole stories and longer than the subject's 882/929 words, which gives the reference more opportunities to match; the subject's zero twelve-grams is not a length artifact, since a single one would have shown.)
Two exposures are declared rather than left to the number. The lead read Morri's 1918 translator's preface — his statement of register programme, not his text — before translating, and knew his title. No line of his chapter 1 was read until the renderings were committed.
Order of work, as at S123 and S125. Chapter 1 read in Japanese → the 18 loci fixed on
source-side grounds and verified verbatim → R06 frozen → R04 and its log frozen → R21 frozen →
all committed → comparator's chapter 1 opened → loci aligned → contamination measured → this design
written. Every figure in §2 existed before a line of §§3–6 was specified.
3. The instrument
One 0–6 integer per item: how high or low the register of this passage is. 0 = coarse, low,
slangy; 6 = elevated, formal, literary. One scale, no other axes. The wording of the poles, the
system prompt and the JSON shape are given in run.py and are identical for the Japanese and
English stages except for the language note.
18 loci × 4 arms. The loci are single sentences or clauses, fixed from the Japanese alone and sorted into three sets of six:
| set | what it is | loci |
|---|---|---|
| LOW | narration the source marks below its own neutral: 無鉄砲 / もんだ / てやった / つらまえて / なんて / 持てあました | L1–L6 |
| NEU | narration the source leaves unmarked | N1–N6 |
| HIGH | deferential speech: 下さい, 遊ばせ, お好き, お持ちなさいます, ご機嫌よう, and the narrator's own ます to his father | H1–H6 |
Arms: SRC (Japanese), MORRI (1918), R04 (lead close), R21 (lead vulgarisation).
All 72 strings are verified to occur in their arm's full text by materials/build_loci.py, which
prints all 18 x 4 = 72 locus strings verified. For MORRI the check is on the letter-and-digit
skeleton, because the scan breaks words across lines and mis-sets quotation marks; the raw scan
window around each match is stored in loci.json so every restoration is checkable, and the one
single-letter repair (DECAUSE → BECAUSE, a drop cap) is listed in the script.
Dispatch. Four seats (config/models.md panel v1: P1, P2, P3, P5), no Anthropic model.
Three calls per seat: one source call rating all 18 Japanese loci, and two English blocks of
27 items each. Block A holds L1–L3, N1–N3, H1–H3 in all three English arms; block B holds L4–L6,
N4–N6, H4–H6. Each block is balanced across arms and sets by construction, so every comparison
this design makes is within a single call and therefore on a single scale.
Items are presented under opaque ids and shuffled by a per-(seat, block) SHA-256, so the three renderings of one locus are not adjacent and nothing in the prompt names an arm, a set, a translator, a date or the work. Seats are told the passages are English (or Japanese) and nothing else.
4. Predictions, registered
Write C_low(arm) = mean(arm, LOW) − mean(arm, NEU) and C_high(arm) = mean(arm, HIGH) − mean(arm,
NEU), each computed within block and then averaged, and pooled over seats. The comparison of
these contrasts across arms is what holds content fixed: the same six sentences carry the LOW set
in every arm.
- G1 — source-side gate. In SRC,
mean(HIGH) > mean(NEU) > mean(LOW)for at least 3 of 4 seats, and pooledC_low(SRC) ≤ −0.75andC_high(SRC) ≥ +0.75. If G1 fails, the source's register system is not externally visible on this instrument and every prediction below is withheld. - P1 (primary) — the low pole does not cross.
C_low(MORRI) > 0.5 × C_low(SRC)— Morri retains less than half the source's lowering — andC_low(MORRI) ≥ −0.25in absolute terms, at the pooled level and at ≥ 3 of 4 seats. Reported with an exact paired permutation over the six LOW loci (2⁶ = 64 sign assignments) on the seat-averaged per-locus quantityd_i(arm) = r(arm, L_i) − mean_j r(arm, N_j), testingd(SRC)againstd(MORRI). - P2 — the high pole does cross.
C_high(MORRI) > 0andC_high(MORRI) ≥ 0.5 × C_high(SRC). - P3 (the payload) — asymmetric retention.
C_high(MORRI)/C_high(SRC) > C_low(MORRI)/C_low(SRC). - P4 — it is the pair, not the translator.
C_low(R04) ≥ −0.25as well. If instead the lead's close rendering carries the low pole while Morri's does not, the finding is about Morri, and the design says so.
5. Failure criteria, registered
- FC1 — the instrument must be able to see lowering in English narration. Pooled over the 12
narration loci (LOW ∪ NEU),
mean(R21) ≤ mean(MORRI) − 0.5. If it does not, this instrument cannot detect a lowered English narration at all and P1 and P4 are withheld, because their content is a null. - FC2 — seat agreement. Mean pairwise Spearman ρ across the four seats' 18 source-side ratings must be ≥ 0.40. Below that the seats are not rating a shared property and everything is withheld.
- FC3 — completeness. A call returning fewer than all its items, a non-integer, or an out-of-range value is re-dispatched once; a second failure drops that seat, and the drop is reported in the result with the arithmetic recomputed on the survivors.
- FC4 — R21 is not a rival.
R21is a control and is never reported as a translation recommendation. IfC_low(R21)is used for anything beyond FC1, that is a design violation.
What this design cannot show, stated in advance. (a) One published translator, so a MORRI result
is about Morri 1918; the generality claim rests on P4 and on RS-20260806g's independent
Italian evidence, not on this run alone. (b) The instrument is four language models rating register;
Tier D is NOT PASSED and nothing here is a jury verdict — these are readers of register, not
judges of quality, and no arm is scored for goodness. (c) The LOW and NEU sets differ in content as
well as register; only the contrast-of-contrasts across arms controls that, and the absolute
value of any single arm's C_low does not. (d) Morri's English is a 1918 Japanese translator's
English, and "raised" here cannot be separated from "period".
5b. Amendments after the pre-run critic pass (declared before any rating call)
The critic seat (nvidia/nemotron-3-ultra-550b-a55b) returned finish_reason: length with EMPTY
content, spending all 16,000 completion tokens on hidden reasoning — the failure notes (bhf) and
(bhq) describe, on the same slug S123 recorded it for. The findings below were recovered from the
response's reasoning field (critic.md), which is repetitive and has broken numbering. That is
a limit on how much adversarial coverage this run received and it is reported as one; a second
critic pass on a different, non-panel seat was dispatched for that reason (critic2.md). Both
were read; every BLOCKING finding is applied below.
- A1 — H1 leaves the HIGH set (critic BLOCKING). H1 is the narrator's own ます-form to his
father; H2–H6 are deference addressed to the narrator. Those are different phenomena and
averaging them confounds
C_high. The HIGH set is H2–H6, five loci. H1 is still rated, still reported, and reported alone, as a single-locus observation that enters no arithmetic. - A2 — FC1 is split, because the critic is right that it could not tell a blind instrument from a
failed control (critic BLOCKING). FC1a, arm-independent: across all 54 English items the
ratings must span ≥ 3 points and at least 4 items must be rated ≤ 2 — i.e. the seats use the low
end of the scale at all. P1 and P4 are withheld only if FC1a fails. FC1b, the original
R21 ≤ MORRI − 0.5on narration, is retained but demoted to a reported figure: if FC1a passes and FC1b fails, that is a fact aboutR21, not about the instrument, and the result says so.R21's per-locus ratings are printed either way. - A3 — the absolute thresholds are demoted (critic ADVISORY, accepted). LOW and NEU loci differ
in content, so an absolute
C_lowis content-confounded. P1 now holds on two content-matched criteria only:C_low(MORRI) > 0.5 × C_low(SRC), and the exact paired permutation over the six LOW loci comparingd(SRC)withd(MORRI). The −0.25 figure is reported and is no longer a criterion. Same demotion for P4. - A4 — a length-matched contamination reference is added (critic ADVISORY, accepted). The published-pair reference is two whole stories against the subject's ~900 words. The same pair, truncated to the first 900 words of each, is run as a fifth cell and reported beside it.
- Checked, no change: every LOW locus is narration (critic BLOCKING 1). L1–L6 were re-read against the source: all six are narratorial, none is dialogue or reported speech. The verification is the change the finding asked for.
- Noted, no change: the three English arms of one locus fall in the same block. The critic granted that the per-seat hash shuffle mitigates it; it is recorded as a limit.
Second critic pass (qwen/qwen3.7-max, non-panel) — verdict NEEDS-AMENDMENT, all six applied
- A5 — G1 loses its NEU-baselined absolute threshold (BLOCKING 1, and it is right). Requiring
C_low(SRC) ≤ −0.75re-imports on the source side exactly the content confound A3 demoted on the English side. G1 is now: the ordering holds at ≥ 3 of 4 seats, and pooledmean(SRC, HIGH) − mean(SRC, LOW) ≥ 1.0.C_low(SRC)andC_high(SRC)are still reported and still used as the denominators of P1/P2/P4 — which is legitimate where the numerator is the same six sentences in another language, and is not legitimate as a standalone gate. That distinction is the amendment. - A6 — the dispatched Morri strings are shown, word by word (BLOCKING 2). The critic's worry is
that scan damage inside a locus would reach the raters as malformed English and depress its
rating.
materials/build_loci.py --tokensprints all 168 distinct tokens actually dispatched for MORRI; they were read, and every one is a well-formed English word or a transliteration (sasa-ame,rikishas,Kojimachi-ku,Azabu-ku). One printing artifact survives and is declared:house keepingin H5 is set as two words. The mechanical part of the check — no one-letter token outside {a, A, I} — passes. - A7 — the permutation test's power floor is stated, not assumed away (ADVISORY 3). With six
paired loci the two-sided exact P cannot go below 2/64 = 0.03125, and that floor is reached
only when the observed mean difference exceeds every one of the 62 other sign assignments. The
test is therefore reported beside two coarser statistics that do not share its ceiling: how
many of the four seats put
C_low(MORRI)above half ofC_low(SRC), and how many of the six LOW loci individually sit higher in MORRI than in SRC. - A8 — FC1a becomes symmetric (ADVISORY 4). Also require ≥ 4 English items rated ≥ 5. A
scale used only at the bottom would compress
C_highand make P2 fail for an instrument reason. - A9 — P3 is reported ratio-free as well (ADVISORY 5). Beside the two retention ratios, the
result reports points lost on each pole:
|C(SRC)| − |C(MORRI)|for high and for low. The claim "one pole survived better than the other" is then readable without dividing by anything. - A10 — cross-language scale calibration is added to the limits (ADVISORY 6). The 0–6 scale is not shown to mean the same thing when a seat reads Japanese and when it reads English, and this run contains no language-neutral calibration item. Every SRC↔English comparison here — including the P1 denominator — therefore rests on an assumption the run does not test. It is limit (e) in §5 and travels with every figure.
6. Budget
Pre-flight, built from max_tokens and the worst plausible provider (note (abc), and the routing
caution in config/models.md): 12 rating calls $0.22, one pre-run critic call at 16,000
max_tokens $0.17, one re-dispatch allowance $0.11. Declared worst case $0.50. Today's
headroom at session start: $4.574866 of $5.00.