Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260813e-slot-typology-ja/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260813e-slot-typology-ja
statusfrozen
created2026-08-13
updated2026-08-13
sensesstyle-correspondence, voice, cultural-mediation
internal-judgment-onlytrue
provisionaltrue
linkswiki/arms/ARM-footing.md, wiki/findings/results/RS-20260812i-footing-channel.md, wiki/findings/results/RS-20260812h-dakghar-grade.md, framework/v0.2/README.md, workshop/canon/kiseru/manifest.md, workshop/translations/kiseru/R04-v1/translation.md, config/models.md, config/budget.md

E-20260813e — where a source's footing mark lands in English, and whether position alone decides

ARM-footing step 2 (T5), the arm's closing step. Frozen before any English of 「煙管」 sections 三–八 was written, before any hand was dispatched, and before Glenn Shaw's 1930 English of this story was opened by the lead.

1. The question, and why it is not the question step 1 asked

RS-20260812i §3.1 offered, post hoc, an account of why Russian's deference clitic -с reaches English at 42 of 43 while the ты/вы choice reaches it at 0 of 11. Both are bound morphology; neither has an English equivalent. The proposed difference was positional:

The clitic sits at the end of an utterance, and English's utterance-final position is optional and open: sir drops straight into it. The ты/вы choice sits in the subject pronoun, and English's subject pronoun is compulsory and has exactly one form.

That account was written after the numbers were seen, and its own result page says so (§6.2). It has never predicted anything in advance. This experiment makes it predict, prospectively, in a third language family — and it is built around the one cell that can tell the account apart from its obvious competitor.

The two accounts agree everywhere Russian and Bengali have been measured. They disagree on exactly one thing this project can reach: the Japanese honorific prefix 御/お. It attaches at the left edge of a noun; the English counterpart position — a pre-nominal modifier — is free by Account A's test, since the pipe → your gracious pipe adds no propositional content. Account A therefore predicts it is carried. Account B predicts it is lost, because no conventional English device sits there.

Whichever way that cell goes, framework/v0.2 §10 has something to say that is not a restatement of what was already believed.

2. Materials

source 芥川龍之介 「煙管」 (1916), whole, sections 一–八. Sections 一–二 frozen at S061 (workshop/canon/kiseru/source.txt, 1,734 non-space chars); sections 三–八 added this session (source-3-8.txt, 4,617 non-space chars, 69 paragraphs). Aozora card 80, file 80_15183.html, cp932
comparator Glenn W. Shaw, "The Pipe", in Tales Grotesque and Curious (Hokuseido, 1930), Project Gutenberg #78105. A published human hand, 96 years old, blind to this hypothesis by construction. Original and translation both free — the materials rule of CLAUDE.md §Materials satisfied without an excerpt limit
census scope direct speech only — every span inside 「 」. Footing is a property of what people say to each other; narratorial honorifics are a different phenomenon and are excluded by rule. Four quoted phrases that are not speech (加賀の煙管, 天保六歌仙, 莫迦め, 金箔) are dropped

Why this story. Selected on measured source-side signal before anything was designed: of six freely-available Akutagawa stories that Shaw also translated, 「煙管」 has the highest density of graded dialogue per source character. It also carries the single best site this project has met — the same word 手前 used by one speaker as a contemptuous second person to an equal and as a humble first person to a daimyō (C1-06/C1-11/C1-12 against C1-08/C1-13).

3. The four channels

Cut by where the mark sits in the source and by what English has at the counterpart position.

ch source device English counterpart position free? conventional English device there?
C1 personal-pronoun tier — 貴公 お前 手前(2nd) / 己 わし 私 手前(1st) こちとら the pronoun no — compulsory, one form —
C2 utterance-final predicate level — ございまする まする ましょう やす / じゃ ぞよ the utterance end yes yes — sir, my lord, your lordship
C3 pre-nominal honorific prefix 御/お the pre-nominal modifier yes no
C4 lexical honorific/humble — 拝領 仰せ とらす まいれ 殿様 先様 御前(title) 恐れ 有難う 頂く the word itself — —

C4 is a positive control and is declared as one in advance, not a finding. A lexical honorific is also a thing that was said; predicting that it survives is predicting that translators translate words. This is step 1's pre-run critic's finding against its own P1, imported here rather than re-learned.

4. The census, and how sites were adjudicated

census.py emits candidates from patterns committed in this file. adjudicate.py applies an explicit include/drop table to those candidates and writes sites.json. Both were run, and the site list frozen, before any English was written, opened or dispatched.

Adjudication rules, in order:

  1. A candidate is a site only if its form grades the speaker against the addressee — up (deference, self-lowering) or down (contempt, condescension).
  2. C2 counts one site per speech, classified by the level of that speech's final predicate, and only where the level is marked. The plain forms peers use with each other (だ さ よ ぜ) are the default and are not sites; ございます-class forms (up) and the lord's じゃ/ぞよ (down) are.
  3. C3 admits only true honorific prefixes. お in お前, おい, おお is not a prefix and is dropped.
  4. C4 admits only marks whose removal changes what the sentence says. 居りまする (the humble auxiliary おる) is contentless suppletion and is dropped from C4 rather than inflating a control.
  5. Direction (up / down) is recorded per site from the speaker–addressee table in sites.json.

Frozen counts: C1 13 · C2 12 · C3 7 · C4 19 — 51 sites in 44 speeches.

5. Hands

hand text blind to this hypothesis? in the registered numbers?
SHAW Shaw 1930, whole story yes, by construction (1930) yes
LEADb the lead's T-kiseru-R04-v1, sections 一–二, written 2026-07-30 (S061) yes — it predates ARM-footing and predates the free-slot account by thirteen days yes, on its 14 sites
LEADa the lead's sections 三–八, written this session no NO — excluded from every registered number
P1 openai/gpt-5.6-terra, whole story yes — the prompt names no channel, no hypothesis, and no footing yes
P3 x-ai/grok-4.5, whole story yes, same prompt yes

LEADb/LEADa is the internal contrast this design gets for free: one translator, one story, one half written before the hypothesis existed and one after.

6. Coding

Carriage is coded by two panel seats that produced no hand: P2 google/gemini-3.6-flash and P5 deepseek/deepseek-v4-pro. Plain-text line output, not structured output (NEXT.md's standing finding on P5). Each coder sees, per hand: the source utterance, the marked form, a literal gloss, and the hand's English for that utterance — and never the hand's identity, never the channel labels, and never any statement of the hypothesis.

The verdict set, fixed here:

The lead codes all sites as well. The lead's coding is a third reading and enters no primary.

7. Registered predictions

Pooled over hypothesis-blind hands (SHAW, LEADb, P1, P3), NA excluded from denominators.

# prediction bar
G1 C1 is lost. Replicates ты/вы at 0 of 11 in a third language family carriage ≤ 0.15
G2 C2 is carried. Replicates the -с clitic's 42 of 43 carriage ≥ 0.50
G3 C3 is lost — Account B. The English position is free and the mark dies anyway carriage ≤ 0.15
G3′ C3 is carried — Account A. Positional freedom is the whole story carriage ≥ 0.50
G4 C4 is carried — positive control, not a finding carriage ≥ 0.85
G5 No insertion. English footing devices at speech units the Japanese leaves unmarked in all four channels, pooled over blind hands. Follows Russian's 0 of 40 against Bengali's 2 hands at 1 site ≤ 2
G6 Ordering C1 ≈ C3 < C2 < C4 holds in each blind hand separately 4 of 4 hands

G3 and G3′ are the two accounts. Both are registered. A C3 carriage between 0.15 and 0.50 decides neither and will be reported as deciding neither.

8. Failure criteria

# fires when consequence
F1 any channel holds fewer than 6 sites that channel is descriptive only. C3 at 7 is the tightest and is declared so here
F2 a machine hand shares ≥ 3 twelve-grams or a run ≥ 12 tokens with SHAW its agreement with SHAW is not reported as replication anywhere
F3 raw P2/P5 agreement on CARRIED-vs-LOST falls below 0.75 in a channel that channel's number carries the disagreement rate and no ordering claim rests on it
F4 always LEADa is excluded from every registered number
F5 a hand returns NA at more than 0.20 of its sites that hand is descriptive only
F6 the lead's translation of 三–八 shares ≥ 3 twelve-grams or a ≥ 12-token run with SHAW recorded on the artifact; the lead is excluded from the primaries already, so nothing else follows

9. Procedure

  1. Freeze this design, census.py, adjudicate.py, sites.json. Done before step 2.
  2. Independent pre-run critic — an off-panel seat (qwen/qwen3.7-max) that produces no hand and does no coding. Amendments recorded in amendments.md before any other call.
  3. The lead translates 三–八 from the Japanese alone; translator's log frozen with it.
  4. Dispatch P1, P3 on the whole story. Prompt names no channel and no hypothesis.
  5. Overlap checks (F2, F6) by tools/dependence_check.py.
  6. Dispatch coders P2, P5, five hands each.
  7. Compute, verify by an independent recomputation with mutation tests, write the result.

10. Pre-flight cost

Built from the caps actually sent, never from an expected output length — note (abc).

call seat n cap out worst case
pre-run critic qwen/qwen3.7-max 1 12,000 $0.060
whole-story hands openai/gpt-5.6-terra 1 6,000 $0.040
whole-story hands x-ai/grok-4.5 1 6,000 $0.048
coding google/gemini-3.6-flash 5 4,000 $0.195
coding deepseek/deepseek-v4-pro 5 4,000 $0.031
re-dispatch allowance dearest seat 4 4,000 $0.156
declared ceiling $0.60

Today's headroom at freeze: $5.00 − $3.450609849 = $1.549390151. The ceiling fits.

11. What this cannot show

Tier D is not passed. Nothing here is a claim that carrying a mark is better than losing it, or that Shaw's English is good. The question is entirely where the marks land, and the framework section it feeds says so in its first line.