Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260807b-botchan-register/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260807b-botchan-register
statusfrozen
created2026-08-07
updated2026-08-07
linkswiki/arms/ARM-ja-register.md, wiki/base/anchors/A-morri-botchan/A-morri-botchan.md, wiki/base/sources/S-botchan.md, workshop/translations/botchan/R04-v1/translation.md, workshop/translations/botchan/R21-v1/translation.md, workshop/regimes/R21-vulgarisation.md, wiki/findings/results/RS-20260806g-negative-pole.md, wiki/goodness-senses.md, config/models.md
sensesstyle-correspondence, voice, naturalness
provisionaltrue

E-20260807b — does the low pole of a Japanese narrator's register cross into English?

Frozen 2026-08-07 (S126), before any call was dispatched. Every lead rendering named here was written and committed before a line of the comparator's chapter 1 was read (commit 72bef11).

1. The question, and why it is asked now

Two consecutive results have found the same thing on the same language pair and named it the same way. RS-20260806f (S123) measured the largest loss in a four-language narrator-crossing run on the Japanese item, on diction temperature; RS-20260807 (S125) repeated the run on four obscure works and again found the largest single movement on the Japanese item, this time register −1.75. Both times the lead's translator's log, frozen before the instrument existed, had named register as the property being spent, and both times it said the same thing about why: English has no slot for a register the Japanese carries in its pronouns and verb endings.

Independently, RS-20260806g (S124) produced a craft finding on Italian→English: English's raising shelf is deep and available in narration, while its lowering shelf is almost entirely dialogue-shaped — a translator can ennoble narration by lexis alone and cannot do the reverse by lexis alone.

Put together, those give a claim about the language pair that no run has yet tested against published practice: if the lowering shelf really is dialogue-shaped, then a Japanese narrator whose register is low should arrive in English raised, in any competent hand, and the deferential speech around him should arrive intact. This design tests that on a published human translation, read whole against its source.

One sentence on what it teaches about literature (subject rule, wiki/tracks.md): it measures what happens to a novel's register system — a low-register narrator set against a deferential servant — when the novel is translated into a language with one pronoun and no deference morphology.

2. Materials

Source. 夏目漱石「坊っちゃん」(1906), chapter 1, 7,499 non-space characters, 22 paragraphs, read whole in Japanese. Aozora Bunko 000148/752_14964, 底本『ちくま日本文学全集 夏目漱石』筑摩書房, fetched 2026-08-07 (materials/src-botchan-ch1.txt). PD: Sōseki d. 1916.

Published comparator. Yasotarō Morri, Botchan (Master Darling), Ogawa Seibundo, Tokyo, 1918 — the earliest published English Botchan (S-botchan §Identity). PD on the US publication basis (1918, pre-1929) and separately on the translator's life (Morri 1882–1959, d. 1959; the publication basis is the one relied on). Internet Archive botchanmasterdar1918nats, djvu text, chapter 1 extracted whole: 4,461 words raw, materials/morri-ch1-raw.txt; furniture-stripped working copy materials/morri-ch1.txt; the five corresponding paragraphs, 882 words, materials/morri-spans.txt.

Lead arms, all $0, all frozen and committed before the comparator's chapter 1 was opened: T-botchan-R06-v1 (draft, 934 words), T-botchan-R04-v1 (close, 929 words), T-botchan-R21-v1 (vulgarisation control, 769 words), over the same five whole source paragraphs — 1,507 non-space characters, chosen because they are the five in chapter 1 that carry both poles of the register system at once.

Contamination, measured before this design was written and after the renderings were frozen (materials/contamination.json; tools/dependence_check.py):

cell shared 7-grams 12-grams longest run verdict
subject — lead R04 × Morri 1918 4 0 9 clean
ref — an independent published pair (K&M 1915 × Garnett 1920, «Пари») 160 27 24 DEPENDENT?
ref — the lead against itself, two regimes, one session (R04 × R21) 23 1 12 DEPENDENT?
ref — the lead against itself, draft and revision (R06 × R04) 558 420 75 DEPENDENT?

The subject cell sits below the independent-published-pair reference on every column, which is the comparison that matters: the lead's English is further from Morri's than two independent published translators are from each other. (Limit: the reference pair's texts are whole stories and longer than the subject's 882/929 words, which gives the reference more opportunities to match; the subject's zero twelve-grams is not a length artifact, since a single one would have shown.)

Two exposures are declared rather than left to the number. The lead read Morri's 1918 translator's preface — his statement of register programme, not his text — before translating, and knew his title. No line of his chapter 1 was read until the renderings were committed.

Order of work, as at S123 and S125. Chapter 1 read in Japanese → the 18 loci fixed on source-side grounds and verified verbatim → R06 frozen → R04 and its log frozen → R21 frozen → all committed → comparator's chapter 1 opened → loci aligned → contamination measured → this design written. Every figure in §2 existed before a line of §§3–6 was specified.

3. The instrument

One 0–6 integer per item: how high or low the register of this passage is. 0 = coarse, low, slangy; 6 = elevated, formal, literary. One scale, no other axes. The wording of the poles, the system prompt and the JSON shape are given in run.py and are identical for the Japanese and English stages except for the language note.

18 loci × 4 arms. The loci are single sentences or clauses, fixed from the Japanese alone and sorted into three sets of six:

set what it is loci
LOW narration the source marks below its own neutral: 無鉄砲 / もんだ / てやった / つらまえて / なんて / 持てあました L1–L6
NEU narration the source leaves unmarked N1–N6
HIGH deferential speech: 下さい, 遊ばせ, お好き, お持ちなさいます, ご機嫌よう, and the narrator's own ます to his father H1–H6

Arms: SRC (Japanese), MORRI (1918), R04 (lead close), R21 (lead vulgarisation). All 72 strings are verified to occur in their arm's full text by materials/build_loci.py, which prints all 18 x 4 = 72 locus strings verified. For MORRI the check is on the letter-and-digit skeleton, because the scan breaks words across lines and mis-sets quotation marks; the raw scan window around each match is stored in loci.json so every restoration is checkable, and the one single-letter repair (DECAUSE → BECAUSE, a drop cap) is listed in the script.

Dispatch. Four seats (config/models.md panel v1: P1, P2, P3, P5), no Anthropic model. Three calls per seat: one source call rating all 18 Japanese loci, and two English blocks of 27 items each. Block A holds L1–L3, N1–N3, H1–H3 in all three English arms; block B holds L4–L6, N4–N6, H4–H6. Each block is balanced across arms and sets by construction, so every comparison this design makes is within a single call and therefore on a single scale.

Items are presented under opaque ids and shuffled by a per-(seat, block) SHA-256, so the three renderings of one locus are not adjacent and nothing in the prompt names an arm, a set, a translator, a date or the work. Seats are told the passages are English (or Japanese) and nothing else.

4. Predictions, registered

Write C_low(arm) = mean(arm, LOW) − mean(arm, NEU) and C_high(arm) = mean(arm, HIGH) − mean(arm, NEU), each computed within block and then averaged, and pooled over seats. The comparison of these contrasts across arms is what holds content fixed: the same six sentences carry the LOW set in every arm.

5. Failure criteria, registered

What this design cannot show, stated in advance. (a) One published translator, so a MORRI result is about Morri 1918; the generality claim rests on P4 and on RS-20260806g's independent Italian evidence, not on this run alone. (b) The instrument is four language models rating register; Tier D is NOT PASSED and nothing here is a jury verdict — these are readers of register, not judges of quality, and no arm is scored for goodness. (c) The LOW and NEU sets differ in content as well as register; only the contrast-of-contrasts across arms controls that, and the absolute value of any single arm's C_low does not. (d) Morri's English is a 1918 Japanese translator's English, and "raised" here cannot be separated from "period".

5b. Amendments after the pre-run critic pass (declared before any rating call)

The critic seat (nvidia/nemotron-3-ultra-550b-a55b) returned finish_reason: length with EMPTY content, spending all 16,000 completion tokens on hidden reasoning — the failure notes (bhf) and (bhq) describe, on the same slug S123 recorded it for. The findings below were recovered from the response's reasoning field (critic.md), which is repetitive and has broken numbering. That is a limit on how much adversarial coverage this run received and it is reported as one; a second critic pass on a different, non-panel seat was dispatched for that reason (critic2.md). Both were read; every BLOCKING finding is applied below.

Second critic pass (qwen/qwen3.7-max, non-panel) — verdict NEEDS-AMENDMENT, all six applied

6. Budget

Pre-flight, built from max_tokens and the worst plausible provider (note (abc), and the routing caution in config/models.md): 12 rating calls $0.22, one pre-run critic call at 16,000 max_tokens $0.17, one re-dispatch allowance $0.11. Declared worst case $0.50. Today's headroom at session start: $4.574866 of $5.00.