Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260814b-honorific-hands/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260814b-honorific-hands
statusfrozen
created2026-08-14
updated2026-08-14
trackT5
sensesstyle-correspondence, voice, cultural-mediation
internal-judgment-onlytrue
provisionaltrue
linkswiki/arms/ARM-honorific-hands.md, framework/v0.2/README.md, wiki/findings/results/RS-20260813e-slot-typology-ja.md, wiki/findings/results/RS-20260812i-footing-channel.md, workshop/translations/genji-yomogiu/R04-v2/translation.md, workshop/translations/genji-yomogiu/R06-v1/translation.md, config/models.md, wiki/method-notes.md

E-20260814b — does a published human hand carry a Japanese honorific, and by what device?

Frozen 2026‑08‑14, ARM-honorific-hands step 1 (T5). The translation limb (T-genji-yomogiu-R06-v1 at f2cf142, T-genji-yomogiu-R04-v2 at ffa3a5f) was frozen before this file existed, in that order and in separate commits. No published English of this span had been opened when this design was frozen.

1. Question

At sites where 『源氏物語』 grades social footing in grammar English does not have, do published human translators carry the grading — and do they use the rank-noun-in-the-pronoun-slot device that RS-20260813e §4 saw only in 2026 machine hands?

framework/v0.2 §10.3 records the device and refuses to recommend it because "the only hands that do it are the two 2026 language models." §10.5 says the fix is materials: two or more freely readable published human translations. This design is that census.

2. Materials

source 紫式部『源氏物語』ch. 15 「蓬生」 §§2‑3 「叔母、末摘花を誘う」 + 2‑4 「侍従、叔母に従って離京」 — 1,841 characters, 30 paragraphs, stored at wiki/base/anchors/A-yosano-yomogiu/yomogiu-2-3-2-4-murasaki-classical.txt (Shibuya 校訂本文 via the UVA Japanese Text Initiative, same route as the S017 extract)
why this span it holds three speakers at three social levels talking to each other — an aunt married into the provincial governor class speaking up to a royal princess, the princess answering down, and a gentlewoman (Jijū) speaking up to both — plus a narrator who grades the princess and nobody else. The whole deference system is exercised inside 1,872 characters
SUEM Suematsu Kenchō 1882 (rev. 1900), Project Gutenberg #19264, Japanese Literature, ch. XV "Overgrown Mugwort". Published human hand, native Japanese speaker
WALEY Arthur Waley 1926, Project Gutenberg #67111, The Sacred Tree, ch. XV "The Palace in the Tangled Woods". Published human hand, English scholar
P1 openai/gpt-5.6-terra, 2026 machine hand, prompted with no hypothesis and no mention of honorifics
P3 x-ai/grok-4.5, same prompt
LEAD T-genji-yomogiu-R04-v2, the lead's in-session rendering. Hypothesis-aware and therefore excluded from every registered number. Coded and reported for completeness only

Declared priming. During materials search on 2026‑08‑14, before the span was chosen, the lead read SUEM's rendering of §§3‑1–3‑3 — a different, non-overlapping stretch of the same chapter. This is why site selection is mechanical (§3): the lead must not be the one deciding which sites count. It is declared again on the translation artifact.

3. Sites — mechanical, 65, frozen with this file

extract_sites.py (frozen alongside) scans the source for a committed pattern list and emits every match with its paragraph, its speech/narration context and its stratum. The paragraph-level speech/narration state machine is part of the frozen rule. No site was added, removed or re-labelled by hand. Output: sites.tsv.

stratum what the Japanese marks sites speech narration
H1 subject-honorific (尊敬): たまふ, せ/させたまふ, おはす(まし), のたまふ, 思す, 賜ふ — grades the subject's referent above the speaker 34 20 14
H2 object-honorific / humble (謙譲): きこゆ/きこえさす, たてまつる, うけたまはる, まうづ, まかる, 参る — grades the object above the subject 13 10 3
H3 addressee-polite (丁寧): はべり, さぶらふ — grades the addressee, and has never been measured in this project 10 10 0
H4 humble shimo-nidan 給ふ: 思ひたまへ, 見たまふる, おぼえたまふ — the speaker's own perception marked humble, the finest grade the language has 3 3 0
H5 honorific prefix 御 on a noun: 御装束, 御みづから, 御宿世, 御ありさま, 御女, 御心, 御髪 7 5 2
total 67 48 19

3a. Negative controls — mechanical, 10, frozen with this file

The 10 paragraphs of the span that contain zero pattern matches, computed by the same script and stored as controls.json: ¶4, 11, 12, 17, 18, 22, 23, 25, 26, 27. Two of them are the aunt's brusque commands to a servant (「さらば、侍従をだに」, 「いづら。暗うなりぬ」), two are narration that pointedly withholds elevation from the aunt (されど、行く道に心をやりて、いと心地よげなり), four are the two poems, two are speech tags. At a control paragraph the Japanese marks no social relation grammatically, so an English that marks one there has inserted it. Coders are given these ten interleaved with the 67 sites and are not told which are which.

4. Coding

Per site × hand, one of three codes:

NO-RENDERING is reported separately and excluded from the carriage denominator. This is method note (bnj) generalised from S180: a hand that omits a sentence has not lost the marking, it has omitted the sentence, and pooling the two would print an omitting hand as a losing hand. Note (bnj) was written for a locus absent from the translator's own source text; the same reasoning covers a locus the translator's source contains and his English does not render, and the two counts stay apart.

Seats and blinding. Two coders per hand, cross-assigned so that no model codes its own translation (RS-20260813e's rule): P2 google/gemini-3.6-flash on every hand; the second coder is P1c openai/gpt-5.6-terra except on P1's own hand, where it is P3c x-ai/grok-4.5. Coders see the Japanese paragraph, the marked form, a gloss of what it marks, and the hand's English — never the hand's identity, never the stratum labels, never the hypothesis, never any other hand's English. P5 deepseek/deepseek-v4-pro is out as a coding seat on any task shape, note (bne).

Consensus = both coders say CARRIED. Raw agreement is reported per stratum.

5. The device census — coder-free, and the sharper half

A committed regex over each hand's full English of the span, for the rank-noun set:

lordship · ladyship · my lord · my lady · your excellency · your highness · your majesty
your honour/honor · your reverence · your grace · his/her grace · liege · sir · madam · ma'am

Amended (critic BLOCKING 3). "Classified by position" was not a rule, and G3 was defined on the bucket that needed the judgment. Three fully deterministic measures replace it, all regex-only, all computable by device_census.py (frozen alongside), rates per 100 words of that hand's English of the span:

measure rule what it is
DEV-all every hit of the set above the rank-noun rate, no judgment at all — G3 is registered on this
DEV-poss a hit immediately followed by 's / ’s the exact shape of §10.3's exhibit, Your Lordship's hand
DEV-voc a hit with a comma, dash or sentence boundary on both sides vocative address, the shape RS-20260812i was about

DEV-all − DEV-voc is reported as the non-vocative residue. No model is asked anything for this table, and no hand-classification enters any registered number.

6. Registered predictions

Bars set now, before any published English of this span is opened.

# prediction bar
G1 The loss replicates. Pooled H1+H4 consensus carriage over the two published human hands is at or below the S1-a level ≤ 0.15
G2 The polite stratum is lost too. H3 consensus carriage over the two published human hands — a stratum this project has never measured ≤ 0.15
G3 The §10.3 confound is confirmed: the machine hands' DEV-all rate per 100 words exceeds the published human hands' by at least 1.0 ≥ +1.0
G4 No insertion. No hand marks a social relation at more than 2 of the 10 negative-control paragraphs (consensus of both coders) ≤ 2 per hand
G5 The prefix stratum crosses. H5 consensus carriage over the two published human hands, because 御 attaches to a noun English also has ≥ 0.35
G6 Coder agreement floor. Raw two-coder agreement in every stratum with ≥ 5 sites ≥ 0.75

G3 is the prediction the arm exists to test and it is written as the lead's expectation so that its failure is legible. The decision rule is registered with it:

Amended (critic NON-BLOCKING 7). The dissolved branch is reworded so it does not overclaim what two pre-1930 hands can establish: "§10.3's 'only the two 2026 language models do it' is false for this span; period and hand remain confounded (§11), so nothing here says this is what English does." It can falsify the sentence as printed. It cannot establish English practice.

6a. Amendments made after the pre-run critic pass, before any hand or coder was dispatched

P4 moonshotai/kimi-k3 returned NEEDS-AMENDMENT, 8 findings, 4 BLOCKING. All eight are accepted; four in full, one in part.

# finding disposition
B1 the frozen table said 65 sites and the script emitted 67 accepted. The critic is right and the cause is on the record: after freezing at 6297568 the lead found that the H4-over-H1 override compared pattern strings rather than positions, so "たまふ" in "おぼえたまふ" deleted every H1 match of that shape anywhere in a paragraph containing an H4 site. It cost two real sites in ¶13. Fixed positionally, script re-run, §3's table now identical to the output. No hand or coder had been dispatched.
B2 the negative-control list was referenced but never defined accepted — §3a, mechanical, 10 paragraphs, frozen as controls.json
B3 ARG/VOC/REF was not coder-free and G3 rested on it accepted — §5 replaced by three deterministic regex measures; G3 re-registered on DEV-all
B4 G1/G2/G5 undefined if F1 fires on one hand accepted — registered now: if F1 fires on exactly one published human hand, G1/G2/G5 are evaluated on the surviving hand and the result page says so in its headline, not its limits. If it fires on both, F2 withholds them
N5 the both-coder consensus rule is the most conservative estimator and coincides with the lead's low-carriage expectation on G1/G2 accepted — both-coder and either-coder rates are reported side by side for every stratum. G1, G2 and G5 are all decided on the both-coder rate, which makes G5's high bar harder and G1/G2's low bars easier; the asymmetry is stated rather than removed, because the estimator must be the same for all three
N6 the glosses are the lead's unreviewed philology and they are the coders' target accepted in part. A gloss audit is dispatched to P3 — shown the Japanese, the marked form and the gloss, asked whether the gloss states the relation correctly, blind to the hypothesis and to every English hand. Disagreements are logged and those sites' carriage is reported separately. The critic's three named sites (25, 35, and the whole H4 stratum) are flagged in advance
N7 the G3 dissolved-branch overclaims accepted — reworded above
N8 the 0.60–0.74 agreement band is a dead zone accepted — stated explicitly: agreement in [0.60, 0.75) means G6 failed and the stratum is still admissible; below 0.60 F4 fires

7. Failure criteria — when a number may not be read

# criterion consequence
F1 a published human hand returns NO-RENDERING on more than 0.50 of its 65 sites that hand's rates are reported but not pooled; the primary becomes single-hand and says so
F2 both published human hands exceed 0.50 NO-RENDERING the human-hand comparison is withheld and the session reports the materials failure as its result
F3 a coder returns CARRIED on more than 0.20 of the negative-control sites that coder's verdicts are not read and the stratum is single-coded, flagged
F4 raw coder agreement below 0.60 in a stratum of ≥ 5 sites that stratum's rate is reported with the agreement figure and is not used to decide G1/G2/G5

F1 is a live risk and is registered as such: SUEM is an abridging hand — his whole ch. XV runs about 12,850 characters of English against 14,226 characters of Japanese, which for a character-dense source implies heavy compression. If he compresses this span the way the count suggests, F1 fires on him and the primary rests on WALEY alone, which would be one published human hand and not two, and the result page must say so in its headline rather than in its limits.

8. Procedure

  1. Freeze this file and extract_sites.py and sites.tsv. (done — commit recorded on the result page)
  2. Independent pre-run critic pass: P4 moonshotai/kimi-k3, shown this design and the site table, asked for blocking findings. Amend in writing before any other call is dispatched.
  3. Dispatch the two machine hands (P1, P3) on the source span, neutral prompt (§9).
  4. Extract SUEM and WALEY's renderings of the span from the stored Gutenberg texts.
  5. Run the device census (§5) — mechanical, no calls.
  6. Dispatch the coders (§4), one call per (hand × coder), 65 sites per call.
  7. Post-run verification: recompute every reported number from the raw outputs; mutation tests on the tally code.
  8. Write RS-20260814b-honorific-hands.

9. The machine-hand prompt (frozen verbatim)

You are translating a passage of classical Japanese literature into modern English prose.

Below is a passage from 『源氏物語』, chapter 15 「蓬生」, in Shibuya Eiichi's critical text. Translate it into English. Keep the paragraph divisions exactly as given, one English paragraph per Japanese paragraph, and render the two poems as poems on their own lines. Do not add notes, headings, commentary or explanations. Output the translation and nothing else.

[source paragraphs]

Nothing in the prompt mentions honorifics, deference, social rank, register, or this design's question. temperature default, max_tokens 3000.

10. Pre-flight cost estimate

Built from max_tokens, not from expected output — note (abc).

call seat in (tok) max out worst case
critic P4 $3/$15 4,000 3,000 $0.057
machine hand × 2 P1 $1/$6, P3 $2/$6 1,500 ea 3,000 ea $0.020
coder × 10 P2 $1.5/$7.5 (5), P1c $1/$6 (4), P3c $2/$6 (1) 7,000 ea 4,000 ea $0.396
total worst case $0.473

Declared: $0.60. Today's UTC day stands at $0.000000000 of $5.00 before this run.

Declaration revised to $1.60 after the critic pass, and the reason is a seat failure, not a design change. P4 moonshotai/kimi-k3 was dispatched three times on the critic prompt. The first two spent their entire completion budget on hidden reasoning — 4,000 and 14,000 tokens — and returned finish_reason: length with empty content: $0.079782 + $0.233043 = $0.312825 for no output at all. The third, with reasoning: {"effort": "low"} and max_tokens 6,000, returned a full verdict in 91 seconds for $0.083293. The lesson is a rule of future conduct and is written as a method note: on a reasoning seat, max_tokens is not a ceiling on the answer, it is a ceiling on the answer plus the thinking, and the thinking will take all of it. Note (abc) says to build the worst case from the cap the request permits; it does not say that the cap can be consumed without producing anything, and this run is the first time that has happened here.

11. What this design cannot do