Repository path: workshop/experiments/E-20260814b-honorific-hands/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260814b-honorific-hands |
| status | frozen |
| created | 2026-08-14 |
| updated | 2026-08-14 |
| track | T5 |
| senses | style-correspondence, voice, cultural-mediation |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-honorific-hands.md, framework/v0.2/README.md, wiki/findings/results/RS-20260813e-slot-typology-ja.md, wiki/findings/results/RS-20260812i-footing-channel.md, workshop/translations/genji-yomogiu/R04-v2/translation.md, workshop/translations/genji-yomogiu/R06-v1/translation.md, config/models.md, wiki/method-notes.md |
E-20260814b — does a published human hand carry a Japanese honorific, and by what device?
Frozen 2026‑08‑14, ARM-honorific-hands step 1 (T5). The translation limb
(T-genji-yomogiu-R06-v1 at f2cf142, T-genji-yomogiu-R04-v2 at ffa3a5f) was frozen before
this file existed, in that order and in separate commits. No published English of this span had
been opened when this design was frozen.
1. Question
At sites where 『源氏物語』 grades social footing in grammar English does not have, do published human translators carry the grading — and do they use the rank-noun-in-the-pronoun-slot device that
RS-20260813e§4 saw only in 2026 machine hands?
framework/v0.2 §10.3 records the device and refuses to recommend it because "the only hands that
do it are the two 2026 language models." §10.5 says the fix is materials: two or more freely
readable published human translations. This design is that census.
2. Materials
| source | 紫式部『源氏物語』ch. 15 「蓬生」 §§2‑3 「叔母、末摘花を誘う」 + 2‑4 「侍従、叔母に従って離京」 — 1,841 characters, 30 paragraphs, stored at wiki/base/anchors/A-yosano-yomogiu/yomogiu-2-3-2-4-murasaki-classical.txt (Shibuya 校訂本文 via the UVA Japanese Text Initiative, same route as the S017 extract) |
| why this span | it holds three speakers at three social levels talking to each other — an aunt married into the provincial governor class speaking up to a royal princess, the princess answering down, and a gentlewoman (Jijū) speaking up to both — plus a narrator who grades the princess and nobody else. The whole deference system is exercised inside 1,872 characters |
SUEM |
Suematsu Kenchō 1882 (rev. 1900), Project Gutenberg #19264, Japanese Literature, ch. XV "Overgrown Mugwort". Published human hand, native Japanese speaker |
WALEY |
Arthur Waley 1926, Project Gutenberg #67111, The Sacred Tree, ch. XV "The Palace in the Tangled Woods". Published human hand, English scholar |
P1 |
openai/gpt-5.6-terra, 2026 machine hand, prompted with no hypothesis and no mention of honorifics |
P3 |
x-ai/grok-4.5, same prompt |
LEAD |
T-genji-yomogiu-R04-v2, the lead's in-session rendering. Hypothesis-aware and therefore excluded from every registered number. Coded and reported for completeness only |
Declared priming. During materials search on 2026‑08‑14, before the span was chosen, the lead
read SUEM's rendering of §§3‑1–3‑3 — a different, non-overlapping stretch of the same chapter.
This is why site selection is mechanical (§3): the lead must not be the one deciding which sites
count. It is declared again on the translation artifact.
3. Sites — mechanical, 65, frozen with this file
extract_sites.py (frozen alongside) scans the source for a committed pattern list and emits every
match with its paragraph, its speech/narration context and its stratum. The paragraph-level
speech/narration state machine is part of the frozen rule. No site was added, removed or
re-labelled by hand. Output: sites.tsv.
| stratum | what the Japanese marks | sites | speech | narration |
|---|---|---|---|---|
H1 |
subject-honorific (尊敬): たまふ, せ/させたまふ, おはす(まし), のたまふ, 思す, 賜ふ — grades the subject's referent above the speaker |
34 | 20 | 14 |
H2 |
object-honorific / humble (謙譲): きこゆ/きこえさす, たてまつる, うけたまはる, まうづ, まかる, 参る — grades the object above the subject |
13 | 10 | 3 |
H3 |
addressee-polite (丁寧): はべり, さぶらふ — grades the addressee, and has never been measured in this project |
10 | 10 | 0 |
H4 |
humble shimo-nidan 給ふ: 思ひたまへ, 見たまふる, おぼえたまふ — the speaker's own perception marked humble, the finest grade the language has |
3 | 3 | 0 |
H5 |
honorific prefix 御 on a noun: 御装束, 御みづから, 御宿世, 御ありさま, 御女, 御心, 御髪 |
7 | 5 | 2 |
| total | 67 | 48 | 19 |
3a. Negative controls — mechanical, 10, frozen with this file
The 10 paragraphs of the span that contain zero pattern matches, computed by the same script and
stored as controls.json: ¶4, 11, 12, 17, 18, 22, 23, 25, 26, 27. Two of them are the aunt's
brusque commands to a servant (「さらば、侍従をだに」, 「いづら。暗うなりぬ」), two are
narration that pointedly withholds elevation from the aunt
(されど、行く道に心をやりて、いと心地よげなり), four are the two poems, two are speech tags.
At a control paragraph the Japanese marks no social relation grammatically, so an English that
marks one there has inserted it. Coders are given these ten interleaved with the 67 sites and are
not told which are which.
4. Coding
Per site × hand, one of three codes:
CARRIED— the English at the corresponding place marks the same social relation the Japanese marks at that site, by any means: a title or address noun, a deferential periphrasis, a register shift, or an explicit statement of rank.NOT-CARRIED— the content of the clause is rendered and the relation is not marked.NO-RENDERING— the clause the site sits in is not rendered by that hand at all.
NO-RENDERING is reported separately and excluded from the carriage denominator. This is method
note (bnj) generalised from S180: a hand that omits a sentence has not lost the marking, it has
omitted the sentence, and pooling the two would print an omitting hand as a losing hand. Note (bnj)
was written for a locus absent from the translator's own source text; the same reasoning covers a
locus the translator's source contains and his English does not render, and the two counts stay
apart.
Seats and blinding. Two coders per hand, cross-assigned so that no model codes its own
translation (RS-20260813e's rule): P2 google/gemini-3.6-flash on every hand; the second coder
is P1c openai/gpt-5.6-terra except on P1's own hand, where it is P3c x-ai/grok-4.5.
Coders see the Japanese paragraph, the marked form, a gloss of what it marks, and the hand's English
— never the hand's identity, never the stratum labels, never the hypothesis, never any other
hand's English. P5 deepseek/deepseek-v4-pro is out as a coding seat on any task shape, note
(bne).
Consensus = both coders say CARRIED. Raw agreement is reported per stratum.
5. The device census — coder-free, and the sharper half
A committed regex over each hand's full English of the span, for the rank-noun set:
lordship · ladyship · my lord · my lady · your excellency · your highness · your majesty
your honour/honor · your reverence · your grace · his/her grace · liege · sir · madam · ma'am
Amended (critic BLOCKING 3). "Classified by position" was not a rule, and G3 was defined on
the bucket that needed the judgment. Three fully deterministic measures replace it, all regex-only,
all computable by device_census.py (frozen alongside), rates per 100 words of that hand's English
of the span:
| measure | rule | what it is |
|---|---|---|
DEV-all |
every hit of the set above | the rank-noun rate, no judgment at all — G3 is registered on this |
DEV-poss |
a hit immediately followed by 's / ’s |
the exact shape of §10.3's exhibit, Your Lordship's hand |
DEV-voc |
a hit with a comma, dash or sentence boundary on both sides | vocative address, the shape RS-20260812i was about |
DEV-all − DEV-voc is reported as the non-vocative residue. No model is asked anything for this
table, and no hand-classification enters any registered number.
6. Registered predictions
Bars set now, before any published English of this span is opened.
| # | prediction | bar |
|---|---|---|
G1 |
The loss replicates. Pooled H1+H4 consensus carriage over the two published human hands is at or below the S1-a level |
≤ 0.15 |
G2 |
The polite stratum is lost too. H3 consensus carriage over the two published human hands — a stratum this project has never measured |
≤ 0.15 |
G3 |
The §10.3 confound is confirmed: the machine hands' DEV-all rate per 100 words exceeds the published human hands' by at least 1.0 |
≥ +1.0 |
G4 |
No insertion. No hand marks a social relation at more than 2 of the 10 negative-control paragraphs (consensus of both coders) | ≤ 2 per hand |
G5 |
The prefix stratum crosses. H5 consensus carriage over the two published human hands, because 御 attaches to a noun English also has |
≥ 0.35 |
G6 |
Coder agreement floor. Raw two-coder agreement in every stratum with ≥ 5 sites | ≥ 0.75 |
G3 is the prediction the arm exists to test and it is written as the lead's expectation so that
its failure is legible. The decision rule is registered with it:
G3met (machines exceed humans by ≥ 1.0) → the confound is confirmed;framework/v0.2§10.3's sentence stands and gains the measurement, and the device may not be recommended.G3fails and the two rates are within 1.0 → the confound is dissolved; §10.3 may be restated as a description of English practice, still short of a recommendation because Tier D is not passed.G3fails in the other direction (humans exceed machines by ≥ 1.0) → §10.3's stated basis is false as printed and the sentence is corrected, not merely qualified.
Amended (critic NON-BLOCKING 7). The dissolved branch is reworded so it does not overclaim what two pre-1930 hands can establish: "§10.3's 'only the two 2026 language models do it' is false for this span; period and hand remain confounded (§11), so nothing here says this is what English does." It can falsify the sentence as printed. It cannot establish English practice.
6a. Amendments made after the pre-run critic pass, before any hand or coder was dispatched
P4 moonshotai/kimi-k3 returned NEEDS-AMENDMENT, 8 findings, 4 BLOCKING. All eight are
accepted; four in full, one in part.
| # | finding | disposition |
|---|---|---|
| B1 | the frozen table said 65 sites and the script emitted 67 | accepted. The critic is right and the cause is on the record: after freezing at 6297568 the lead found that the H4-over-H1 override compared pattern strings rather than positions, so "たまふ" in "おぼえたまふ" deleted every H1 match of that shape anywhere in a paragraph containing an H4 site. It cost two real sites in ¶13. Fixed positionally, script re-run, §3's table now identical to the output. No hand or coder had been dispatched. |
| B2 | the negative-control list was referenced but never defined | accepted — §3a, mechanical, 10 paragraphs, frozen as controls.json |
| B3 | ARG/VOC/REF was not coder-free and G3 rested on it |
accepted — §5 replaced by three deterministic regex measures; G3 re-registered on DEV-all |
| B4 | G1/G2/G5 undefined if F1 fires on one hand |
accepted — registered now: if F1 fires on exactly one published human hand, G1/G2/G5 are evaluated on the surviving hand and the result page says so in its headline, not its limits. If it fires on both, F2 withholds them |
| N5 | the both-coder consensus rule is the most conservative estimator and coincides with the lead's low-carriage expectation on G1/G2 |
accepted — both-coder and either-coder rates are reported side by side for every stratum. G1, G2 and G5 are all decided on the both-coder rate, which makes G5's high bar harder and G1/G2's low bars easier; the asymmetry is stated rather than removed, because the estimator must be the same for all three |
| N6 | the glosses are the lead's unreviewed philology and they are the coders' target | accepted in part. A gloss audit is dispatched to P3 — shown the Japanese, the marked form and the gloss, asked whether the gloss states the relation correctly, blind to the hypothesis and to every English hand. Disagreements are logged and those sites' carriage is reported separately. The critic's three named sites (25, 35, and the whole H4 stratum) are flagged in advance |
| N7 | the G3 dissolved-branch overclaims |
accepted — reworded above |
| N8 | the 0.60–0.74 agreement band is a dead zone | accepted — stated explicitly: agreement in [0.60, 0.75) means G6 failed and the stratum is still admissible; below 0.60 F4 fires |
7. Failure criteria — when a number may not be read
| # | criterion | consequence |
|---|---|---|
F1 |
a published human hand returns NO-RENDERING on more than 0.50 of its 65 sites |
that hand's rates are reported but not pooled; the primary becomes single-hand and says so |
F2 |
both published human hands exceed 0.50 NO-RENDERING |
the human-hand comparison is withheld and the session reports the materials failure as its result |
F3 |
a coder returns CARRIED on more than 0.20 of the negative-control sites |
that coder's verdicts are not read and the stratum is single-coded, flagged |
F4 |
raw coder agreement below 0.60 in a stratum of ≥ 5 sites | that stratum's rate is reported with the agreement figure and is not used to decide G1/G2/G5 |
F1 is a live risk and is registered as such: SUEM is an abridging hand — his whole ch. XV runs
about 12,850 characters of English against 14,226 characters of Japanese, which for a
character-dense source implies heavy compression. If he compresses this span the way the count
suggests, F1 fires on him and the primary rests on WALEY alone, which would be one published
human hand and not two, and the result page must say so in its headline rather than in its limits.
8. Procedure
- Freeze this file and
extract_sites.pyandsites.tsv. (done — commit recorded on the result page) - Independent pre-run critic pass:
P4moonshotai/kimi-k3, shown this design and the site table, asked for blocking findings. Amend in writing before any other call is dispatched. - Dispatch the two machine hands (
P1,P3) on the source span, neutral prompt (§9). - Extract
SUEMandWALEY's renderings of the span from the stored Gutenberg texts. - Run the device census (§5) — mechanical, no calls.
- Dispatch the coders (§4), one call per (hand × coder), 65 sites per call.
- Post-run verification: recompute every reported number from the raw outputs; mutation tests on the tally code.
- Write
RS-20260814b-honorific-hands.
9. The machine-hand prompt (frozen verbatim)
You are translating a passage of classical Japanese literature into modern English prose.
Below is a passage from 『源氏物語』, chapter 15 「蓬生」, in Shibuya Eiichi's critical text. Translate it into English. Keep the paragraph divisions exactly as given, one English paragraph per Japanese paragraph, and render the two poems as poems on their own lines. Do not add notes, headings, commentary or explanations. Output the translation and nothing else.
[source paragraphs]
Nothing in the prompt mentions honorifics, deference, social rank, register, or this design's
question. temperature default, max_tokens 3000.
10. Pre-flight cost estimate
Built from max_tokens, not from expected output — note (abc).
| call | seat | in (tok) | max out | worst case |
|---|---|---|---|---|
| critic | P4 $3/$15 |
4,000 | 3,000 | $0.057 |
| machine hand × 2 | P1 $1/$6, P3 $2/$6 |
1,500 ea | 3,000 ea | $0.020 |
| coder × 10 | P2 $1.5/$7.5 (5), P1c $1/$6 (4), P3c $2/$6 (1) |
7,000 ea | 4,000 ea | $0.396 |
| total worst case | $0.473 |
Declared: $0.60. Today's UTC day stands at $0.000000000 of $5.00 before this run.
Declaration revised to $1.60 after the critic pass, and the reason is a seat failure, not a design
change. P4 moonshotai/kimi-k3 was dispatched three times on the critic prompt. The first two
spent their entire completion budget on hidden reasoning — 4,000 and 14,000 tokens — and
returned finish_reason: length with empty content: $0.079782 + $0.233043 = $0.312825 for no
output at all. The third, with reasoning: {"effort": "low"} and max_tokens 6,000, returned a
full verdict in 91 seconds for $0.083293. The lesson is a rule of future conduct and is written
as a method note: on a reasoning seat, max_tokens is not a ceiling on the answer, it is a
ceiling on the answer plus the thinking, and the thinking will take all of it. Note (abc) says to
build the worst case from the cap the request permits; it does not say that the cap can be consumed
without producing anything, and this run is the first time that has happened here.
11. What this design cannot do
- One work, one chapter, one language pair. Two published human hands is the minimum that can separate a hand from a period, and it is still two.
- The two human hands are 44 years apart (1882, 1926) and both predate 1930, so a period norm of
English translating practice is not separated from anything else.
RS-20260813gnames the same limitation on its own four hands and it is not solved here. SUEMtranslated into a second language andWALEYinto his first. If they differ, this design cannot say which of the two differences explains it.- Nothing about quality. Tier D is NOT PASSED. No verdict here says any hand's English is better than any other's.
- The lead's hand is not evidence. It is hypothesis-aware, written after the account existed, by the agent that designed the census.