Repository path: workshop/experiments/E-20260815d-supplied-footing/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260815d-supplied-footing |
| status | frozen |
| created | 2026-08-15 |
| updated | 2026-08-15 |
| links | wiki/arms/ARM-supplied-footing.md, wiki/base/anchors/A-mikami-silver-blaze/A-mikami-silver-blaze.md, workshop/translations/silver-blaze-ja/R06-v1/translation.md, framework/v0.2/README.md, wiki/findings/results/RS-20260815c-footing-direction.md, config/models.md, config/budget.md |
| senses | — |
| internal-judgment-only | true |
E-20260815d — when the source marks footing nowhere: does the relation cross anyway?
Frozen before dispatch. ARM-supplied-footing step 1 (T4). The lead's own rendering of the
same scene, T-silver-blaze-ja-R06-v1, was frozen at commit fb58476b before this design
existed, and is excluded from every registered primary below.
1. The question
framework/v0.2 §10 has English as the target in every cell it is built on. Everything it can
therefore say is about loss: a source's grammatical footing mark does not reach English, and
translators do not put it anywhere else (S1-a, five pairs). Those numbers cannot distinguish
- the mark was lost — English has no slot, so the inflection dies but the reader still knows who is above whom, from the scene; from
- the relation was lost — the reader of the English does not know.
RS-20260815c measured, on constructed control passages, that narrated behaviour dominates
wording on the direction of a perceived social grading, and §10.6 now carries the limit that
nothing in §10 has ever controlled for what a passage narrates. An EN→JA cell is the control
§10 cannot otherwise buy: an English source has no grammatical footing marks at all, so anything
a Japanese hand marks was read off something that is not a mark.
The question, in one sentence: when the source marks footing nowhere and the target must mark it at every predicate, do independent hands supply the same marking, and do they get it from the words of the utterance or from the scene around it?
2. The wire to the translation limb
The translation limb is T-silver-blaze-ja-R06-v1 — the lead rendering this scene into Japanese,
log frozen first. The translating generated the study limb's question in the ordinary way: the
lead could not write a single utterance of it without choosing a politeness level the English does
not contain, and the log's twelve decisions are twelve records of what the level was read off
(D5: 貴様 taken from what the devil and the hunting-crop; D9: 仰せ/いたす taken from a
narrative sentence forty words earlier). The study limb asks whether other hands read it off the
same things. The lead's log also registers three predictions about those hands (P-L1–P-L3),
which is the only role the lead's rendering plays in the numbers.
3. Materials
| source | Doyle, "Silver Blaze" (1892), the Mapleton stables scene, 501 words, PG #834. materials/mapleton-scene-en.txt |
| sites | materials/sites.json — 21 quoted spans → 17 utterances, each with speaker, addressee, dyad and phase. Built from the English before any Japanese was read. Span 5 ("What's this, Dawson!") is addressed to the groom and is excluded from every dyad |
M |
三上於菟吉 訳「白銀の失踪」, 平凡社《世界探偵小説全集》第三卷, 1930. Public domain in Japan. Stored whole |
O |
大久保ゆう 改訳「シルヴァブレイズ」, 2021, CC BY 4.0. A revision of M, not an independent hand — descriptive only, never counted in a primary |
L |
T-silver-blaze-ja-R06-v1, the lead. Hypothesis-aware. Excluded from every primary |
H1 H2 |
two blind non-lead hands: P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash |
Dyads and phases, fixed in sites.json:
| dyad | pre | post | role |
|---|---|---|---|
BH Brown → Holmes |
3 (U06, U07, U09) | 3 (U12, U14, U16) | the primary |
HB Holmes → Brown |
3 (U05, U08, U11) | 3 (U13, U15, U17) | story-position control |
GH groom → Holmes |
1 (U01) | 1 (U03) | descriptive (the half-crown reversal) |
HW Holmes → Watson |
— | 1 (U10) | register anchor |
Amendment A1, made before any call was dispatched and after sites.json was committed
(1cd69018). U04 is span 6, "No gossiping! Go about your business! And you, what the devil do
you want here?" — its first two sentences are addressed to Dawson and only its last to Holmes.
A site whose text mixes two addressees cannot be coded on one dyad, so U04 is excluded from
BH and from ISO/RATE, leaving BH at 3 pre and 3 post. sites.json is not rewritten;
the exclusion is carried in materials/excluded.json and asserted by the verifier. The U04-included
figure is reported as a robustness check and is not the primary. Perfect rank separation on 3-vs-3
has exact null probability 1/C(6,3) = 0.05 per hand. (The cross-hand product first written here
is withdrawn on the critic's F4, §12: the three hands' independence cannot be established.)
The phase boundary is Doyle's own: the twenty minutes in the parlour, and the narrative sentence "His bullying, overbearing manner was all gone too, and he cringed along at my companion's side like a dog with its master." Nothing in either English utterance set carries a grammatical footing mark, because English has none to carry.
4. Conditions
FULL— the hand is given the whole 501-word scene, narration included, and renders it into Japanese. One call per hand. The prompt never mentions politeness, honorifics, register, rank, deference or this experiment's question; it asks for a translation of a passage of English fiction and nothing else.M,OandLareFULLby construction.ISO— the hand is given one utterance, alone, in its own call, with no narration, no speaker name, no adjacent line, and no other utterance in the context. 14 dyad utterances (BH,HB,GH) × 2 hands = 28 calls. One utterance per call is the whole point: a shuffled list in a single call would let the hand reconstruct the scene.ISOL— added before dispatch on the critic'sF5. The same isolated line, in its own call, with the speaker and the addressee named ("spoken by Brown to Holmes") and nothing else. 14 × 2 = 28 calls.ISOremoves four things at once — the narration, the identities, the sequence, and the knowledge that pre and post share a speaker — andISOLputs the identities back so that the narration can be separated from them.RATE— three seats (P1,P2,P3) rate each of the same 14 utterances out of context, one utterance per call, on a 1–7 scale: how far does the speaker place himself BELOW the person he is speaking to? (1 = far above, 4 = level, 7 = far below), withX= cannot be determined from the words alone, announced as a real and expected answer (criticF6). 14 × 3 = 42 calls.GOLD— added on the critic'sF9. Two non-lead seats (P3,P4) label the coder's ten gold lines on the coder's own four-level scale, seeing neither the coder nor this design. 10 × 2 = 20 calls. This is gateG1b.
P1 and P2 appear both as ISO/FULL hands and as RATE seats. The objects differ — they
produce Japanese in one role and rate English in the other, in separate calls with no shared
context, and no seat ever sees any Japanese. Declared here rather than treated as clean; the panel
is four models wide (P5 is out, note (bne)) and three rating seats cannot be disjoint from two
translating hands.
5. Coding — mechanical, and frozen with this design
analysis/code_politeness.py, committed with this file before any Japanese comparator text was
extracted. Addressee-directed politeness (対者敬語) only:
2 honorific/humble token present (尊敬語・謙譲語・丁重語)
1 polite copula/suffix, no such token (丁寧体)
0 neither, and no contemptuous second person (常体)
-1 no polite form, and a contemptuous second person
Precedence is strictly H > P > C > plain. The critic's F2 asked for C > P > H and that half
of F2 is OVERRULED, with the arithmetic written out (§12): the gap is mean(post) − mean(pre), so
raising a pre utterance shrinks the gap. Under H > P > C a rough-but-polite pre utterance
(「貴様、黙ってください」) scores 1 rather than −1, which raises pre and therefore shrinks the
gap — the conservative direction, not the anti-conservative one the critic read it as. What was
genuinely wrong was the content of the H list, and that half of F2 is accepted in full: the
list no longer treats a request marker as exaltation. The token lists are general Japanese grammar,
written from the language and not from any of these texts. No human or model judgment enters the
outcome variable; the coder is deterministic and its calibration is now checked off the lead by
G1b.
6. Registered predictions — REWRITTEN BEFORE DISPATCH ON THE PRE-RUN CRITIC (§12)
The critic's F1 is accepted and it changes the shape of the run: P1 alone cannot license the
headline, because the post-reversal utterances are pleas, offers and compliance formulas and the
pre-reversal ones are imperatives and insults, and Japanese has conventional renderings for those
sentence types that owe nothing to any scene. P2 therefore becomes the primary and P1 becomes
a manipulation check. P2 holds the sentence types constant by construction — the FULL and
ISO conditions code the same six English utterances — so what FULL adds over ISO is
everything outside the utterance and nothing inside it.
P2 — THE PRIMARY. For each blind hand (H1 = P1, H2 = P2), on the six BH sites:
gap(FULL) − gap(ISO) > 0, with an exact paired permutation null: condition labels are
swapped within each site independently, 2⁶ = 64 relabelings, one-sided.
Criterion: 2 of 2 hands positive, and at least one hand at exact P ≤ 0.05.
P2b — THE PRIMARY'S OWN CONTROL, added on F5. ISO+LABEL gives the same isolated line with
the speaker and addressee named and nothing else, separating the narration is missing from the
identities are missing. Registered: for each hand,
|gap(ISOL) − gap(ISO)| < |gap(ISOL) − gap(FULL)|, 2 of 2. If ISOL sits with FULL
instead, what FULL supplies is who is speaking to whom, not the narrated behaviour, and P2's
reading changes accordingly — a fourth outcome the first version of this design did not have.
P1 — MANIPULATION CHECK, no longer a primary. In FULL, for each independent hand
(M, H1, H2): mean LVL(BH-post) − mean LVL(BH-pre) ≥ 1.0, 3 of 3. A failure here
makes P2 uninterpretable (there is no reversal to attribute), so P1 failing in all three
hands withholds P2.
P1r — rank separation, registered as its own statistic on F4. Perfect rank separation of the
three post over the three pre, per hand, exact null 1/C(6,3) = 0.05. The cross-hand product is
withdrawn: M is a 1930 text that both models may have seen, and dependence_check_cjk.py
measures output overlap, not memorisation, so the three hands' independence cannot be established
and 0.05³ is not claimed.
P3 — DESCRIPTIVE, demoted on F3. The critic is right that the HB dyad cannot fire: Holmes
is silky before and curt after, so any competent hand renders HB flat or falling and a "control"
that cannot fail certifies nothing. HB is reported as description. P2's within-site design is
what replaces it as the control that can actually fail.
P4 — the English side, with an escape. On F6, RATE gains an explicit X = cannot be
determined from the words alone, and X is announced as a real and expected answer. Registered:
across the six BH sites, in at least 2 of 3 seats, either some pre site is rated at or above some
post site, or at least one site returns X. Per-seat mean(post) − mean(pre) over the
determinate sites is reported either way.
P5 — the lead's frozen predictions (P-L1–P-L3), descriptive, the lead being
hypothesis-aware.
7. What each outcome means — the full cross, rebuilt on F1 and F7
P1 |
P2 |
P4 |
reading |
|---|---|---|---|
| holds | holds | holds (words leave it open) | the relation crosses a source that marks it in no grammatical slot, and it crosses through the SCENE. §10's loss is a loss of marking, not of relation |
| holds | holds | fails (words separate) | the relation crosses, and both channels carry it — the utterance's own words and the scene. §10 owes a sentence on the lexical channel, and the scene's share is what P2 measures |
| holds | fails | either | the reversal is produced by the utterance types alone: a hand that never saw the scene marks it just as sharply. Reading (a) is refused, and the honest finding is that English carries this lexically |
| holds | holds | either, and P2b puts ISOL with FULL |
what FULL supplies is who is speaking to whom, not the narrated behaviour. The relation crosses, but through the cast list |
| fails in all 3 | — | — | independent hands do not supply the same footing from an unmarked source; P2 is withheld, and §10's loss claim extends to the relation |
| fails in 1–2 | reported | reported | a hand-type split. If M passes and the models fail, or the reverse, the difference is period-or-training and not English, and no reading above is licensed |
No cell is a null to be explained away. The page reports whichever occurs.
8. Failure criteria and the withholding rule
G1— the coder's mutation test, ten constructed utterances including the critic's hostile cases (くださいon a plea,なさいas a downward imperative,ましな,でしかない,申し訳,たまえ). 10 of 10 or the run stops. (Passed on the amended coder before dispatch.)G1b— the gold labels ratified off the lead, added onF9. Two non-lead seats (P3,P4) label the same ten lines on the same four-level scale, seeing neither the coder nor this design. If the coder disagrees with the majority of ratifiers on 3 or more of the 10, every registered primary is WITHHELD. This is the gate on the instrument, and it can fire.G2alignment. EachFULLhand must yield an identifiable Japanese rendering for all 14 dyad utterances in order. A hand failing ≥ 2 is re-dispatched once; failing again it is dropped, the denominator changes, and the drop is declared in the headline.G3isolation. The verifier reads the stored request bodies and asserts that everyISO,ISOLandRATEcall contains exactly one utterance and no narration.- If
G1,G1b,G2(both blind hands) orG3fires, every registered primary is WITHHELD and reported as withheld, not as a null — with the measured numbers printed, carrying the withholding, so that no successor buys them again.
9. Spend
Pre-flight, rebuilt after the critic's amendments, from the max_tokens cap each request
permits, not from expected length (note (abc)):
| stage | calls | max_tokens |
worst case |
|---|---|---|---|
pre-run critic (P4) — SPENT, $0.098946 |
1 | 9,000 | actual |
GOLD gate G1b (P3, P4) |
20 | 600 | $0.126 |
FULL (P1, P2) |
2 | 4,000 | $0.039 |
ISO |
28 | 400 | $0.056 |
ISOL |
28 | 400 | $0.056 |
RATE (P1, P2, P3) |
42 | 500 | $0.116 |
| inputs, all stages | 120 | — | ~$0.040 |
| re-dispatch allowance | — | — | $0.100 |
| declared ceiling | 121 bodies | $1.00 |
Worst case $0.632 including the critic already spent. Today's UTC ledger stood at $3.080393 of $5.00 across five sessions when this session opened, so headroom is $1.919607 and the ceiling fits inside it with $0.92 to spare.
The key-usage delta is not a cross-check and is not reported as one — note (bof), S192:
the key is billed continuously by something outside these sessions (~$0.81/hour while idle).
Per-request "usage": {"include": true} is the ledger, and a re-dispatched body's discarded
attempt keeps its own usage row (note (boe)).
10. Procedure, in order
T-silver-blaze-ja-R06-v1frozen and committed. Done,fb58476b.- This design,
sites.jsonandcode_politeness.pyfrozen and committed. Before any Japanese comparator text is extracted. G1mutation test.- Independent pre-run critic (
P4), amendments accepted or overruled in writing. - Dispatch
FULL,ISO,RATE. Raw request/response preserved inrun/. - Extract
MandOrenderings of the 15 sites — only now, after 2–5. - Contamination:
tools/dependence_check_cjk.pyonL,H1,H2againstMandO. analysis/verify.pyrecomputes every reported number from the raw outputs, independently of the analysis path, with mutation tests.- Result page; framework consequence at step 2 of the arm.
11. Limits, registered in advance
- Three model seats and two model hands are not readers and not translators of record. Charter §4. The word reader may appear in no claim this page licenses.
- One work, one scene, one language pair, one published human hand.
Mis 1930; period is confounded with hand, exactly as §10.6 says of everything else in the section. Ois a revision ofMand can never be a second independent observation. Its only use is descriptive: what a 2021 reviser does to a 1930 hand's footing marking, which is a period contrast with the wording held largely constant.- The scene is unusually explicit. Doyle states the reversal in narration. That makes it the
easiest possible case for the scene-carries-it reading, and a
P1pass therefore licenses nothing about scenes that do not state it. - Nothing here is about quality. No sense is scored. Tier D is NOT PASSED.
LVLis one axis of footing, not footing (criticF8). Address terms, sentence-final particles (ぞ, ぜ, わ, さ), and command forms carry addressee footing too, and the coder scores most of them 0. A hand that marks the whole reversal in forms the coder cannot see would read as a hand that did not mark it, and that would be a coder artifact reported as a finding about a translator.たまえ/給え/やがれwere added toCon this finding; the particles were not, because a bare〜ぞis as often self-directed as addressee-directed and a wrong rule is worse than a declared blind spot.- The premise is not that English marks footing nowhere (critic
F8). English marks it lexically and indexically — my good sir, Mr. Brown, gadabout, what the devil, the imperative mood. What English has no slot for is a grammatical footing mark, which is what §10 is about.P4is the measurement of the lexical channel and is registered as such. - Three model seats are not three observers (critic
F10). They may share training data, and fourteen calls from one seat are one observer's fourteen answers. "2 of 3 seats" is a majority of three correlated instruments.P1andP2serve as both hands andRATEseats; declared in §4.
12. The pre-run critic, and what was done with it
P4 moonshotai/kimi-k3, one call, $0.098946, verdict NEEDS-REDESIGN, ten findings, five
BLOCKING. Raw response at critic.json. Every finding and its disposition:
| finding | disposition | |
|---|---|---|
| F1 | BLOCKING. P1 can pass on sentence type alone — pre are imperatives and insults, post are pleas, offers and compliance formulas, and Japanese has conventional renderings for those types that owe nothing to any scene |
ACCEPTED, and it restructures the run. P2 becomes the primary and P1 a manipulation check; §7 rebuilt as a P1 × P2 × P4 cross |
| F2a | BLOCKING. The H list is fitted: ください puts every plea at the ceiling; なさい is a downward imperative; お願い/頂戴 mark requests; まし/でし/申し/拝 over-match on ましな, でしかない, 申し訳, 拝啓 |
ACCEPTED IN FULL. Coder amended before dispatch, every change moving it against the predicted rise; G1 re-run at 10 of 10 with the critic's own hostile cases in it |
| F2b | BLOCKING. Precedence should be C > P > H; H > P > C "helps P1 at both ends" |
OVERRULED, with the arithmetic (§5). The gap is mean(post) − mean(pre); H > P > C raises a rough-but-polite pre utterance, which shrinks the gap. The direction was read backwards. The content half of F2 was the real defect and is accepted |
| F3 | BLOCKING. P3 cannot fire — Holmes is silky before and curt after, so HB falls by construction and the "control" certifies nothing |
ACCEPTED. P3 demoted to descriptive. P2's within-site design replaces it as the control that can fail |
| F4 | BLOCKING. The exact nulls attach to a statistic §6 never registered, and 0.05³ assumes an independence between a 1930 text and two models that cannot be established | ACCEPTED. Rank separation registered as P1r with its own exact null; the cross-hand product withdrawn; P2 gets an exact paired permutation over 2⁶ = 64 relabelings |
| F5 | BLOCKING. ISO removes the narration and the identities, the sequence, and the shared speaker; either outcome is uninterpretable, and §7 had no branch for P2 at all |
ACCEPTED. ISOL added — the same line with speaker and addressee named — as P2b, and P2 given registered readings in §7 |
| F6 | NON-BLOCKING. RATE forces an answer on lines carrying no addressee signal; P4's overlap criterion is near-unfalsifiable |
ACCEPTED. X = indeterminate added and announced as expected; P4 restated to count X |
| F7 | NON-BLOCKING. §7 does not exhaust the outcomes; a hand-type split would be reported as a finding about English | ACCEPTED. §7 is now six rows including the hand-type split and the ISOL-with-FULL branch |
| F8 | NON-BLOCKING. "marks footing nowhere" is false — English marks lexically; and たまえ/給え, the likeliest downward imperative in a 1930 hand, was invisible to the coder |
ACCEPTED IN PART. たまえ/給え/やがれ added to C; the premise restated in §11; sentence-final particles not added, with the reason written |
| F9 | NON-BLOCKING. G1's gold labels are the lead's own intentions, so G1 can catch typos but not miscalibration |
ACCEPTED. G1b added: two non-lead seats label the ten lines blind to the coder, and 3 disagreements of 10 withhold every primary |
| F10 | NON-BLOCKING. Seat reuse; three correlated model seats are not three observers | ACCEPTED as a limit (§11) |
Nine of ten accepted, one overruled in writing. The amendments were all made before a single main-run body was dispatched, and the coder amendments were made before any comparator text had been read.
13. G1b result, and the robustness check it forces — written before any main-run body
G1b PASSES: the coder disagrees with the ratifiers' majority on 1 of 10, against a withholding
threshold of 3. P3 x-ai/grok-4.5 and P4 moonshotai/kimi-k3, 20 bodies, $0.063886, blind to
the coder and to this design.
| the one disagreement | coder | P3 |
P4 |
|---|---|---|---|
| 「黙りたまえ」 | −1 | 0 | 0 |
Both ratifiers say たまえ is not contemptuous, and they are right about the form: 〜たまえ is a
condescending familiar imperative — what a superior says easily to a subordinate — not an insult,
and it was added to C on the pre-run critic's F8 without a ratifier having seen it. Two further
near-misses in the same direction are recorded rather than adjudicated: P3 alone read
「そこらをうろつかれちゃ困るんだがね」 as 2 and 「どうか信じてください」 as 2, and on both the
other ratifier agrees with the coder.
The consequence is registered here, before the main run:
R1— pre-registered robustness check. Every registered quantity is recomputed a second time withたまえand給えremoved fromC(scoring 0, the ratifiers' unanimous reading). This matters in the anti-conservative direction: aBH-pre utterance rendered with 〜たまえ scores −1 under the coder and 0 under the ratifiers, and the lower pre value inflates the gapP1andP2are measured on. If any registered criterion changes verdict between the two codings, the primary is reported as UNSTABLE and the ratifiers' coding governs.