Repository path: workshop/experiments/E-20260816d-lexical-channel/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260816d-lexical-channel |
| status | frozen |
| created | 2026-08-16 |
| updated | 2026-08-16 |
| links | wiki/arms/ARM-supplied-footing.md, wiki/findings/results/RS-20260815d-supplied-footing.md, framework/v0.2/README.md, wiki/base/anchors/A-morita-christmas-carol/A-morita-christmas-carol.md, workshop/translations/christmas-carol-ja/R06-v1/translation.md, config/models.md, config/budget.md |
| senses | — |
| internal-judgment-only | true |
E-20260816d — the lexical channel, and the sites where English says nothing
ARM-supplied-footing step 2 (T4). This is v2, rebuilt after the pre-run critic returned
NEEDS-REDESIGN with 15 BLOCKING findings — thirteen implemented, four overruled in writing
(critic-response.md). v1 was never dispatched. Frozen 2026-08-16 before any run call. The lead's Japanese
and its log were frozen first, at 89669e16.
1. The question
framework/v0.2 §10 owes a written subsection on two things RS-20260815d named and could not
measure, because it had six sites, one scene, one work, and two blind hands that turned out to be
measurably dependent on each other:
- the lexical channel — where English marks social footing at all, it marks it in the utterance's own lexis and illocution, and that crosses into a marking language without help;
- the silent sites — where the English words say nothing about who is above whom, and a translator into a language that marks footing compulsorily is inventing.
Both are claims about translating literature, and the practitioner-relevant half is the second: if it is true, the places a translator into Japanese should watch are not the loud ones.
What this unit teaches about translating literature (subject rule, wiki/tracks.md): where a
translator into a language that marks social footing compulsorily gets the marking from when the
English source marks none of it, and which sites are the ones being invented at.
2. What is new against step 1
step 1 (E-20260815d) |
here | |
|---|---|---|
| sites | 6 | 70 |
| scenes | 1 | 4, in two works-within-a-work |
| dyads | 1 | 16 |
| author / date | Doyle 1892 | Dickens 1843 |
| published Japanese hands | 三上 1930 + its own revision | 森田草平, independent |
| blind hands | 2, and dependent (33-char shared run) | 3, dependence measured |
3. Materials
Charles Dickens, A Christmas Carol (1843), Project Gutenberg #46, public domain. 森田草平 訳『クリスマス・カロル』, Aozora Bunko card 4328, public domain — the whole book read, both scenes read whole.
- Scene A — Stave One, 1,595 words, the nephew's visit and the two charity gentlemen.
47 sites, five dyads. Built by
materials/build_sites.pyfrom the English alone. - Scene B — three self-contained spans, 682 words: the counting-house closing (Stave One), the
Cratchit house (Stave Three) and Bob's toast (Stave Three). 23 sites, eleven dyads, including
the one this arm most wants: a husband and wife disagreeing. Built by
materials/build_sites_b.py.
Why scene B exists and why the lead translated it and not scene A. The lead read 森田's Japanese
of scene A while locating the material, and is therefore primed on it. Scene B's Japanese was
not opened until the lead's own rendering and its log were frozen and committed at 89669e16.
The lead is excluded from every primary either way (charter §5, and RS-20260815d's precedent).
4. The outcome variable
analysis/code_politeness.py, copied verbatim from E-20260815d — a frozen, deterministic
coder of the addressee-directed (対者敬語) axis:
2 尊敬語 / 謙譲語 / 丁重語 token present
1 丁寧体 present and no such token
0 neither, and no contemptuous second person
-1 no polite form and a contemptuous second person
No model and no human judgment enters it. Reusing it unchanged is deliberate: it makes this run commensurable with step 1, and it removes the freedom to tune an instrument to a hypothesis.
Two failure modes of this coder, named before the run because both have already been seen in this material:
- Referent honorifics are scored as addressee politeness. 森田's children say
「阿父さんが帰っていらっしゃるところだ」 — an honorific about their father, addressed to
the room. The coder scores
2. §7'sS1sensitivity analysis exists for this. - Deference carried by an address term is invisible. 森田's clerk answers "If quite
convenient, sir" with 「ご都合が宜しければ、貴方。」 — no polite predicate, so the coder
scores
0, while the lead's 「お差支えございませんようでしたら」 scores2. The scale is one axis of footing, not footing (RS-20260815dlimit 6). Reported, not repaired.
5. Conditions, and who does what
Every site is put to every seat one site per call, never batched — batching would leak adjacency,
which is the variable ISO exists to remove.
| stage | what the seat sees | seats |
|---|---|---|
RATE |
the English line alone: no speaker, no addressee, no scene | P1, QR |
ISO |
translate this line into Japanese — the line and nothing else | P1, P2, P3 |
ISOL |
the line plus one clause naming speaker and addressee | P1, P2, P3 |
SWAP |
as ISOL, on the 7 husband↔wife sites, with the two labels exchanged |
P1, P2, P3 |
GATE |
a paragraph and one stretch of speech in it: who is addressed? | P4, QR |
Hands are P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5.
QR is qwen/qwen3.7-max, the first reserve named in config/models.md §reserves, brought in
precisely to widen the voice count; verified live at list price from GET /api/v1/models.
P5 is out on any task shape, note (bne).
AMENDMENT 5, and it is a partial retreat from the critic's F5 fix. v2 as frozen put both
raters (P4 moonshotai/kimi-k3 and QR) outside the hands, so that nothing classified a site and
then translated it. P4 had to be dropped: it honours neither reasoning.max_tokens nor
reasoning.effort, spending whatever cap it is given on reasoning and returning an empty string
(8 of its first 18 bodies at cap 300, still failing at 557), and a cap large enough for it costs
~$1.07 at its $15/M output over 70 sites — which fits under no stop-loss this day's headroom allows.
With five usable models, three hands and two disjoint raters is exactly all of them, so dropping one
forces an overlap. P1 is now both a rater and a hand, and the registered answer is S4
(§8). Every amendment was made on an instrument fact — what a provider does with a token budget
— and never on what any body said. P4's 19 bought bodies are ledgered and discarded; its 8 GATE
bodies were bought before the drop and stand.
RATE asks for a position on a 1–7 scale — 1 = speaks as a clear social inferior, 4 = as an equal,
7 = as a clear social superior — with X = cannot be determined from the words alone offered in
the same sentence as the scale, and named as a real and expected answer.
Non-primary hands, at $0: M = 森田 (published, independent, both scenes) and L = the lead
(scene B only). Neither enters any Q.
6. Site classification, from the English, by seats that never see Japanese
SILENT— both raters answerX.DETERMINATE— both raters answer a number.MIXED— they split. Excluded fromQ1–Q3, and the count is reported.
PREV — a prevalence floor, and it is not a validity check (critic F10). If SILENT or
DETERMINATE holds fewer than 7 of 70 sites, the comparison has no cell to make and Q1–Q3
are withheld. Non-empty cells are all this guarantees; it validates nothing, and the design no
longer calls it a manipulation check.
CTL1 — the validity check, and it gates Q1, Q2 and Q3 (critic F10). Three sites carry an
explicit English deference marker: V02 "If quite convenient, sir", U34 "Both very busy,
sir", U31 "…comforts, sir". Registered: all three are DETERMINATE and their mean RATE
sits below 4. If the raters cannot see sir, the classification underlying all three predictions
is not trustworthy and all three are withheld.
GATE — the eight contestable addressee labels, ratified off the lead. Six sites are paragraphs
split between two hearers (U28/U21, V08/V09, V13/V14) and three are assigned to the
family at large (V13, V15, V18). Every other site carries an explicit narrative tag naming
speaker and addressee in the same or the adjacent sentence, which is why only these eight are
bought. Withhold everything if ≥ 4 of the 8 are contradicted by both raters.
3P flag. Each site is flagged from the English, before any Japanese is read, for whether it
names or refers to a specific person who is neither speaker nor addressee — mechanically, a personal
proper name or a third-person singular human pronoun. 16 of 70 are flagged. Frozen in
materials/three_p.json.
7. Registered predictions
Renamed Q1–Q5 because v1 used P1–P5 for both seats and predictions (critic F11).
Every threshold is joined with an exact permutation test, stratified by scene, and holds only if
both clear (critic F6/F7/F8). Where the permutation space exceeds 20,000 it is sampled at 20,000
with the seed frozen at 20260816.
Q1 — THE PRIMARY. Where the English says nothing, independent hands do not converge.
Mean pairwise disagreement in coded level across the three hands, ISOL condition, is higher at
SILENT than at DETERMINATE sites by ≥ 0.15, at permutation P ≤ 0.05. Disagreement at a site =
the fraction of the three seat-pairs whose codes differ (0, ⅓, ⅔, 1).
What Q1 does and does not say (critic F2). Convergence of independent hands is a necessary
condition for the source having decided the choice, not a sufficient one, and the coder measures one
axis — addressee-directed politeness — not footing. No claim on the result page will call this a
measurement of "invention".
Q2 — the lexical channel. Same-seat ISO↔ISOL agreement is higher at DETERMINATE than
at SILENT sites by ≥ 0.15, at permutation P ≤ 0.05. If the words carry it, naming the parties
adds nothing; if they do not, naming the parties is the only thing that decides.
Q3 — direction, on the line alone. Over DETERMINATE sites, the Spearman correlation between
mean RATE and mean ISO coded level is negative, ρ ≤ −0.35, at permutation P ≤ 0.05.
v1 used ISOL here and that was the wrong variable: ISOL supplies the status information whose
contribution is the thing in question (critic F4/F8).
Q4 — the lead's frozen list, scene B. DESCRIPTIVE (critic F16 — 4 sites against 19 cannot
carry a confirmatory claim). The lead's log, frozen at 89669e16, calls 4 of the 23 scene-B sites
FIX — {V02, V05, V12, V14} — and the rest SUP. Reported: the DETERMINATE rate on each
set, with an exact Fisher test and no threshold. It asks whether a practising translator's own
account of where the source decided for him matches two readers of the same English who never saw
his Japanese.
Q5 — the husband and the wife. At the seven husband↔wife sites of scene B — V16, V17,
V19, V20, V21, V22, V23 — 森田 (1930s) and the lead (2026) already agree at 7 of 7:
the wife carries the polite forms and the husband speaks plain. On the coder's scale that
means the Japanese marks the wife as the deferential party, which is the opposite of what
Dickens's scene does — she wins the argument, and he answers her twice with the same four words
because he has nothing else. (That measurement was made before this design was written and is
descriptive. v1's prose called this "marking the wife above the husband", which inverts the
scale — critic F1, correct, and the number was always the other way.)
Registered: the three blind hands reproduce it in ISOL — wife mean > husband mean — in
≥ 2 of 3 hands.
SWAP, the control that makes Q5 worth having (critic F9). The same seven English lines are
re-dispatched with the speaker and addressee labels exchanged. If the politeness is coming from
the words, the levels stay where they were; if it is coming from the labels — a stereotype about
husbands and wives, supplied by the prompt and not by Dickens — the levels move with the labels.
Registered: the mean level assigned to the utterances now labelled "spoken by the wife to the
husband" exceeds the mean assigned to those labelled the other way, in ≥ 2 of 3 hands. A SWAP
that reproduces the asymmetry on the swapped labels is the strong result, and it is a result about
the model and not about Dickens; the result page must say so.
8. Sensitivity and secondary analyses, all registered
S1— the third-party honorific (critic F3).Q1,Q2,Q3are recomputed on the 54 sites with3P == false. A verdict that does not survive both computations is reported as unstable, not as a result.S2— length (critic F14).Q1recomputed on sites of 3–30 words, dropping the one-word fragments and the 158-word speech. Mean site length per class is reported whatever happens.S4— the rater/hand overlap (amendment 5).Q1is recomputed on theP2–P3pair alone, which shares no model with either rater. That statistic is binary per site (the two hands agree or they do not) and is therefore coarser, but it is fully independent of the classification. A verdict that differs betweenQ1andS4is reported as unstable.S3— model contamination on canonical Dickens (critic F15).tools/dependence_check_cjk.pyon each hand's concatenated Japanese against 森田's, and hand against hand, with the reference cellsRS-20260815dused. $0. Reported whatever it shows: these are famous lines and a memorised published rendering is a live alternative explanation for any convergence.
9. Execution, frozen (critic F12)
Temperature 0, one sample per cell, no seed available from the API and none claimed. Serial
dispatch in frozen census order, scene A then scene B. max_tokens = 300 for RATE/GATE,
300 + 4 × the site's English word count for ISO/ISOL/SWAP. Reasoning budget 150 tokens
explicitly on every call — note (bph): step 1 lost 21 bodies to a reasoning budget set at or
above the content cap. Parsers: POS:\s*([1-7]|X) and WHO:\s*([A-Za-z_]+) anchored per line; a
Japanese body is the outermost 「…」 or the whole reply, requiring ≥ 4 CJK characters.
Missing data. A body with finish_reason: length, or one whose parser returns nothing, is
not kept and is re-bought once, selected mechanically on finish_reason and on the parse
failing — never on what the body says — at double the cap. Both attempts are ledgered, note
(boe). A body that fails twice is dropped, and its site is dropped from that statistic's
denominator only, with the count reported. One rendered request per template is written to
materials/rendered_prompts.txt before dispatch.
10. Spend
Worst case built from the max_tokens caps the requests permit, not from expected length
(note (abc)). 597 bodies: RATE 140, ISO 210, ISOL 210, SWAP 21, GATE 16.
Caps are per seat, because they must be: P2 is the verbose hand (614 output tokens on the longest
site where P1 took 231 and P3 196), so it alone gets 400 + 5w; P1 and P3 get 300 + 3w;
the raters get 700 + 3w.
| stage | seats | bodies | worst case |
|---|---|---|---|
RATE |
P1, QR |
140 | $0.61 |
ISO + ISOL + SWAP |
P1, P2, P3 |
441 | $1.09 |
GATE + the discarded P4 bodies |
— | 35 | $0.27 already spent |
| total, this session | 616 | $1.97 |
Worst case $1.97 < stop-loss $2.05 < declared ceiling $2.20 (critic F13). v1 put the stop-loss
at $1.25, below its own worst case, which licensed a truncated run — the mirror image of the
mistake note (bpq) was written for. The stop-loss must sit above the worst case so that it fires
only when something is actually wrong, and below the ceiling so that it is still a margin. The
stop-loss is cumulative over the whole run file, so the discarded P4 bodies count against it.
Day arithmetic: the UTC day opened for this session at $2.561078 of $5.00. Worst case here takes the day to $4.76.
Headroom at design time: $2.438922, less $0.045802 already spent on the critic. Lead translation, both collations, 森田's coding and every contamination measurement are $0 and are never ledgered (charter §3, A4).
11. Verification
verify.py recomputes every number reported on the result page from run.jsonl, asserts no kept
body was truncated, asserts analysis/code_politeness.py is byte-identical to E-20260815d's, and
runs mutation tests that must be caught. Byte-identity is a reproducibility check and not a
validity argument (critic F17): it proves the same code ran, not that the code measures what the
design wants.
12. What this design cannot do, stated before it runs
- It cannot separate addressee-directed politeness from formal style, because the coder is a
token scale on one axis and the project has no human annotators (critic F2/F3/F7, overruled with
reasons in
critic-response.md). - Three model seats are not readers and not three observers (charter §4). The word reader appears in no claim on the result page. Nothing here is about quality; no sense is scored; Tier D is NOT PASSED.