Repository path: workshop/experiments/E-20260807g-indirect-deference/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260807g-indirect-deference |
| status | frozen |
| created | 2026-08-07 |
| updated | 2026-08-07 |
| senses | style-correspondence, voice, naturalness |
| provisional | true |
| links | wiki/arms/ARM-ja-register.md, wiki/base/anchors/A-morri-botchan/A-morri-botchan.md, wiki/base/anchors/A-shaw-spider-thread/A-shaw-spider-thread.md, wiki/findings/results/RS-20260807b-register-collapse.md, workshop/translations/botchan/R04-v1/translation.md, workshop/translations/botchan/R04-v2/translation.md, workshop/translations/saigo-no-ikku/R04-v2/translation.md, config/models.md, config/budget.md |
E-20260807g — does deference survive the move out of quotation marks?
Frozen 2026-08-07 (S131) before any call. ARM-ja-register step 2, target (b).
1. The question, and why it is not about the apparatus
RS-20260807b §5 found that Yasotarō Morri lost Botchan's high register pole — the deference that
constitutes Kiyo and, through her, the narrator's social position — and A-morri-botchan §6 named
the mechanism from inspection: he converts her direct speech to indirect discourse, which deletes
the Japanese morphology outright. That is an inference from four sites in one translator's hand, and
it is confounded with everything else Morri did at those sites.
The question here is a craft question about English, not about Morri: when a deferential utterance
is moved out of quotation marks and reported instead, holding the content fixed, how much of the
deference does English lose, and can a translator get it back inside indirect discourse? No regime
in workshop/regimes/ and no recommendation in framework/v0.1 mentions the direct/indirect choice
at all; if it is a register carrier, that is an omission with consequences for both.
Subject-rule sentence (continue-prompt.md §4.5): this unit teaches whether the direct/indirect
discourse choice is a register decision in JA→EN literary translation, and what a translator who must
make it can and cannot recover.
2. Materials
Seventeen deferential utterances, each taken with the minimal narrative frame that identifies who speaks to whom, in three cells:
| cell | source | rendering, frozen | n | hypothesis-aware? |
|---|---|---|---|---|
| A1 | Sōseki 「坊っちゃん」 ch. 1 (1906) | T-botchan-R04-v1, S126 |
4 | no |
| A2 | Sōseki 「坊っちゃん」 ch. 1 (1906) | T-botchan-R04-v2, S131, this session |
8 | yes |
| B | Ōgai 「最後の一句」 (1915) | T-saigo-no-ikku-R04-v2, S128 |
5 | no |
A1 + B = 9 utterances whose English was written before this design existed. The primary is read on those nine. A2's eight were written today by a translator who knew the question, and are reported as a separate, flagged arm — see §7 limit 1.
Arms. Each utterance appears in four forms. DIR is the filed rendering verbatim; the other
three are written for this design and are the only new prose in the materials:
DIR— the lead's filed rendering. Direct discourse.IND— the same content in indirect discourse under a mechanical rule: quotation removed, reporting verb supplied, deixis and tense backshifted. No compensating material added.INDC— indirect discourse with licit compensation: reporting verbs, adverbs, honorific third-person description, self-lowering predicates. No fact added.DIRF— direct discourse with the deference devices removed. The lexical-flattening control: it holds the form and takes the marking, which is the mirror of whatINDdoes.
MORRI — Yasotarō Morri's published 1918 English, on the twelve Botchan utterances. Not a fifth
manipulated arm: the published comparator, entering as itself.
Materials: materials/build_loci.py → materials/loci.json. 80 English items (17 × 4 + 12).
Contamination, measured before this design was written (materials/contamination.json):
T-botchan-R04-v2 × Morri chapter 1 — 0 shared 7-grams, 0 twelve-grams, longest run 6 tokens,
clean, cleaner than S126's 4 / 0 / 9. Note (bhb) applied as a gate: the project had not
rendered these five paragraphs before (R04-v2 × R04-v1: 0 / 0, run 5).
3. Procedure
Stage 0 — pre-run critic. One non-panel adversarial pass over this frozen design. Findings applied as numbered amendments before any other call.
Stage 1 — source-side gate. Three seats rate 24 Japanese items — the 17 marked utterances and 7 unmarked comparators (the father's 貴様, the classmate's taunt, the magistrate's 帰れ帰れ, and four more) — on the deference scale, reading Japanese only, no English anywhere in the prompt.
Stage 2 — the rating stage. The same three seats rate the 80 English items, shuffled with opaque ids, in three blocks each. Nothing in any prompt names an arm, a translator, a work, a date or this hypothesis. Blocks are hash-fixed per seat so the shuffle is reproducible.
Stage 3 — parity gate. One non-panel model, not a seat, is shown each DIR beside its IND and
INDC and asked whether the second reports the same content as the first. Four planted content
errors (W1–W4) are mixed in, in IND form, as the positive control: a parity call that cannot
catch a planted error cannot certify the real ones. 38 judgments.
Stage 4 — device census. One non-panel model, not a seat, is shown the 17 DIR renderings and
asked (i) to list the words and constructions that mark the speaker's deference and (ii) to say for
each whether it could still be present if the utterance were reported rather than quoted. It is
not shown the lead's own table (T-botchan-R04-v2 §1), the hypothesis, or any other arm.
The scale (identical in stages 1 and 2), an integer 0–6:
How much deference — social subordination to the person being addressed — does the speaker's language convey? 0 = none, or the speaker is talking down; 3 = neutral, between equals; 6 = deep deference, the speaker placing themselves far below the person addressed.
Judgment is never parallelised. All raw bodies preserved in runs/ before anything is computed.
4. Gates, fixed now
G1— instrument sensitivity. PooledDIR − DIRF≥ 1.00, and positive at 3 of 3 seats. If the scale cannot see deference removed lexically while the form is held, it cannot be trusted to see deference removed formally while the lexis is held.G2— source-side marking. Pooledmarked − unmarkedon the Japanese ≥ 1.50, and positive at 3 of 3 seats. Establishes that the 17 are marked in the source and that the 7 comparators are not.G3— parity. ≥ 80% of the 34 realIND/INDCitems judged content-equivalent to theirDIR, and ≥ 3 of the 4 planted errors caught. Both halves required.G4— no leak. The verifier reconstructs every dispatched prompt from the pre-call maps and confirms no arm name, translator, work, date or hypothesis token appears in any.
5. Predictions, registered
P1 (primary), read on cells A1 + B only (9 utterances).
DIR − IND ≥ 0.50 × (DIR − DIRF), positive at 3 of 3 seats, and positive at ≥ 7 of 9
utterance-medians. The syntactic move alone costs at least half of what stripping the deference
lexically costs.
P2 — compensation is partial. INDC < DIR on A1 + B, at 3 of 3 seats. The recovery fraction
(INDC − IND) / (DIR − IND) is reported whatever P2 does.
P3 — Morri's loss is accounted for by the conversion. On the 7 Botchan utterances Morri renders
indirectly, MORRI ≤ IND + 0.5; on the 5 he keeps direct, MORRI > IND. If both halves hold,
the published translator's register collapse is the same effect measured on a hand that was not
trying to demonstrate it.
P4 — the structural explanation. In stage 4's independent census, ≥ 60% of the deference
devices found in the 17 DIR renderings are classified as unable to survive reporting.
P5 — the discourse-form census, no jury, computed from the texts. Morri converts ≥ half of
the 12 Botchan deferential utterances to indirect discourse. (The lead's count before dispatch is
7 of 12; this is registered as a claim to be verified, not discovered.)
6. Failure criteria, fixed now
F1— ifG1,G2orG3fails,P1,P2andP3are withheld and the run is descriptive. Not weakened after firing.F2— if more than 20% of the realINDitems are judged content-inequivalent,P1is withheld: the conversion changed more than form.F3— if any seat's ratings have a standard deviation below 0.5 across the 80 items, that seat is reported and excluded from pooled figures (a seat not using the scale is not evidence).F4— a body that omits items or returns empty content is re-dispatched once to the same seat, and the dead body is preserved inruns/discarded/and ledgered as waste (notes (bhd), (bhf), (bhq)). A second failure drops the seat and is reported.F5— if cells A1+B and A2 disagree in direction onP1, the A2 result is reported and the primary stands on A1+B alone, with the disagreement named as evidence of the §7 limit 1 bias.
7. Known limits, written before the run
- The A2 renderings are hypothesis-aware. The lead wrote them today knowing this question would
be asked. This is why the primary excludes them; it is not a fix, because the same lead wrote the
IND,INDCandDIRFarms for all cells. Nothing here puts the manipulations outside the lead.G3bounds content drift; it does not bound the lead's choice of how much deference to put intoDIRin the first place. INDCis deliberately effortful English and may read as awkward. The scale asks about deference, not quality, but a seat that dislikes the prose may mark it down anyway.INDitems are structurally shorter than theirDIRcounterparts (the vocative goes). Length is a confound the design does not control and the verifier will report it.- Three seats, Tier D NOT PASSED. Nothing here is a jury verdict. The seats rate a described social relation; no arm is scored for goodness and no translation is ranked.
- Two works, one language pair, one direction. Whatever holds here is JA→EN.
- The 17 utterances are the lead's selection, though
G2tests that they are marked and the comparators are not.
8. Budget
Pre-flight worst case from max_tokens, not from expected length (note (abc)):
| stage | calls | cap | worst case |
|---|---|---|---|
| 0 critic | 1 | 16,000 | $0.06 |
| 1 source | 3 | 6,000 | $0.09 |
| 2 rating | 9 | 8,000 | $0.36 |
| 3 parity | 1 | 10,000 | $0.05 |
| 4 census | 1 | 14,000 | $0.07 |
re-dispatch reserve (F4) |
≤ 3 | 8,000 | $0.12 |
Declared ceiling $0.75. Headroom at freeze: $1.823046555 of the UTC-day $5.00 cap, six prior sessions today. If a stage would breach the declaration it is cut, not silently overrun.
9. Amendments, all made before any rating call was dispatched
Stage 0 pre-run critic (nvidia/nemotron-3-ultra-550b-a55b, non-panel): NEEDS-REDESIGN, 5
BLOCKING, 6 NON-BLOCKING. Its BLOCKING 2, 3 and 4 would have voided the primary. Ten findings
accepted, two remedies overruled.
A1(critic BLOCKING 3, 4, 5) — the three manipulated arms are no longer the lead's.IND,INDCandDIRFare written by an independent model that is not a seat, from theDIRpassages alone, under instructions that state what to change and not why. TheINDinstruction never mentions deference, register or the hypothesis. This is the amendment that matters: the critic was right that excluding cell A2 protected only theDIRside, while the lead still wrote every comparison arm in every cell, and thatG1was circular because its control was the lead's own idea of what flattening looks like. The lead's versions are preserved unrated asLEAD_IND/LEAD_INDC/LEAD_DIRFinmaterials/arms.jsonand enter no figure.A2(BLOCKING 2) — the rating task is reworded. "The person being addressed" is ambiguous in reported speech, where the grammatical addressee becomes the reader. The scale now says: one person's words are given, quoted or reported; rate how much deference those words show toward the person they were spoken to; when the words are reported, judge the speaker who said them, not the narrator.A3(BLOCKING 1) — speech-act preservation is required and tested. TheINDinstruction requires that a request stay a request and a question stay a question, andG3tests it. The critic was right that the lead's ownINDhad recast "be so good as to hand them over" as a reported obligation, which is a change of illocutionary force and not a backshift.A4(NON-BLOCKING 6, 11) — mean item length per arm is computed and reported; the A1+B vs A2 comparison reports effect sizes, not only direction.A5(NON-BLOCKING 10) —P4is DEMOTED from a registered prediction to a reported descriptive figure. The critic is right that one model classifying deference devices is that model's linguistic intuition, not evidence.A6(NON-BLOCKING 8) — aG2failure is declared in advance to be ambiguous between "the items are not marked" and "the seats cannot read Japanese deference", and will be reported as such.A7— form-compliance screen on the generated arms: mechanical check only (is the speech quoted or reported?). A non-compliant item is reported, never edited into compliance.A8— the declared ceiling rises from $0.75 to $0.90 to pay for stage 0b. Headroom at the moment of the raise: $1.823046555.A9— the first generating model returned zero content on theINDarm at caps 14,000 and 30,000 (note (bhf); note (bhq)'s cap-raise remedy failed, as it did at S128) while writingINDCandDIRFcleanly. Rather than let two hands write different arms — which would confound arm with author — all three arms were regenerated by a single second hand. Both discarded bodies are preserved inruns/discarded/and ledgered as waste.A10— an arithmetic correction to §2 and §5, made before the rating stage. A1 + B is 9 utterances, not 11;P1's utterance-median criterion is ≥ 7 of 9, not ≥ 8 of 11. The design as first frozen misstated its own n.
Overruled, with reasons. (i) The critic's remedy "restrict the primary to DIR vs DIRF, treat
IND/INDC as exploratory" is refused: that is the arm's question, and abandoning it would leave
A-morri-botchan §6's mechanism claim exactly as unmeasured as it was. A1 and A2 address the
same worry without giving up the comparison. (ii) "Add a naturalness rating for all items" is
refused on cost — it doubles stage 2 against a $0.90 declaration — and carried instead as §7 limit 2.