Repository path: workshop/experiments/E-20260820c-mimetic-subtraction/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260820c-mimetic-subtraction |
| status | frozen |
| created | 2026-08-20 |
| updated | 2026-08-20 |
| version | v3-post-critic |
| senses | perceived-source-carriage, style-correspondence, accuracy |
| provisional | true |
| internal-judgment-only | true |
| links | wiki/arms/ARM-mimetic-subtraction.md, wiki/findings/results/RS-20260817e-mimetic-reading.md, workshop/experiments/E-20260817e-mimetic-reading/design.md, workshop/translations/botchan-ch2/R06-v1/translation.md, workshop/translations/botchan-ch3/R06-v1/translation.md, wiki/method-notes.md, framework/v0.2/README.md |
E-20260820c — is §7.19's subtractive reading of a Japanese mimetic driven by the property or by the mutilation
One-sentence design. On the same nine sites RS-20260817e §3 admitted a clean deletion for,
put the same two English strings — the enacting rendering and the stating rendering, byte-identical
— to the same three model seats twice: once against the original Japanese, and once against the
Japanese with the mimetic replaced by a plain, grammatical Japanese phrasing of the same event
that carries no phonaesthetic property. Two independent hands write those substitutes and their
outputs are compared under a rubric frozen below before the reader run is dispatched (note bqm).
1. Question and predictions
Question. §7.19 tells a translator: cover the mimetic in the source and ask whether your
English word still has anything to answer to. The evidence for it (RS-20260817e, S204) came from
a design in which the source lost the mimetic by deletion — a mutilation the result page named
as its live threat and did not address. This experiment asks: when the mimetic is removed without
mutilation — replaced by the plainest Japanese phrasing of the same event — does the enacting
English still read as an addition?
Registered predictions (v3, per critic-response.md).
P1(primary — directional). On the SEVEN primary sites (M01 M03 M04 M05 M07 N06 N07, i.e. the nine deletable sites minus M08 and M12), the enacting arm's majorityADDSflag flips from no (mimetic present) to yes (plain substitute) atflips_no_to_yes ≥ 5of the 7 evaluable primary sites ANDflips_yes_to_no ≤ 1. This is a directional criterion — a run with 5 forward and 3 reverse flips FAILS. Reasoning: counting only forward flips is not a valid sign-test criterion (critic BLOCKING 6); an explicit net-direction rule with a cap on reverse flips answers what the earlier bar meant to. 5 of 7 tracks S204's 6 of 9 as a ratio (0.71 vs 0.67) at a slightly higher bar, which the smaller n demands.- The exact two-sided sign test p-value on discordant pairs is reported as a descriptive statistic, computed by enumerating all 2ⁿ sign sequences. Its 0.5 null is stated as contestable — substituting wording changes semantic specificity and demand cues, and exchangeability is not guaranteed — so it does not decide the primary.
P2(secondary). The stating arm'sOMITSrate is reported without a directional prediction; the S204 mirror did not move (RS-20260817e§3) and no reason to expect substitution to fix that.P3(continuity — aggregate). In the mimetic-present cell of this run, the enacting arm reads "adds nothing" at ≥ 6 of the 7 primary sites, matching S204's per-site outcomes at these sites.P3′(continuity — per-item, new v3). At ≥ 6 of the 7 primary sites, the majority of this run's mimetic-present cell equals the majority of S204's mimetic-present cell at the same uid. A miss at 2 or more of 7 voidsP1as a comparison against S204.- M08 and M12 are dispatched and reported in §7 (side data) but do not enter the primary. Reasons and remedies in critic-response.md, BLOCKING 1 and 2.
Failure criteria (v3).
F1. GateG3(substitute-parity, §4) fails: the two hands' substitutes disagree at ≥ 3 of 9 sites. DISCHARGED — G3 measured 8/9 at stage B, passes; the disagreeing site (M08) is excluded from the primary under BLOCKING 1's remedy. M08's data still enters §7.F2. GateG0fails: a substitute reintroduces a mimetic word or changes the event. That site drops from the primary; if ≥ 3 of 7 primary sites drop, the primary isWITHHELD.F3. GateG1fails: within the actually-dispatched primary arms, the enacting flag isnoeverywhere oryeseverywhere on the mimetic-present cell.F4.P3orP3′(continuity, above) misses. Named as an explicit primary void, not a soft warning.F5— complete-case rule (v3, per critic MAJOR 9). A pair is evaluable if both cells (mk × presentandmk × sub_lead) parse three seats each. A pair with any dead cell drops from the primary and its uid is named. If fewer than 6 of the 7 primary sites are evaluable, the primary isWITHHELDand step 2 of the arm rebuilds the design.F6— parse rate. Overall parse rate on the dispatched reading run below 90% voids the run; per-condition parse rates < 80% void that condition. Both include re-dispatch under (bqk).
What no outcome licenses on this design.
- Not a claim about Japanese. Three named model seats, on nine curated items from one novel. Charter §4 forbids treating panel agreement as validation. Everything in the frozen result page will carry that qualifier. Tier D NOT PASSED at any grain that would license a jury reading.
- Not a per-site verdict. With nine sites, only the pooled sign test survives.
- Not a decision about §7.19 as a whole. §7.19 also carries a mirror instruction about the stating arm (§7.19 second-half warning); that half is not tested here.
2. Materials
Frozen in materials/items.json, built by materials/build_materials.py and identical, at every
row that appears there, to the corresponding row in E-20260817e's items.json. Nine sites, one
per row: M01 M03 M04 M05 M07 M08 M12 N06 N07. For each site:
ja_sentence— the original Botchan sentence, verbatim.mimetic_span— the substring to be replaced.plain_span_lead— the lead's plain substitute.ja_sub_lead— the lead's substituted sentence.en_frame— the English sentence, with a{X}slot.mk— the enacting English fill.pl— the stating English fill.lead_gloss,lead_note— for the log, not read by any seat.
Each row also carries this experiment's class label (sound/manner/contested) inherited from
E-20260817e §5 G0b, and is not re-classified here.
The independent hand's substitutes are collected by run_stage_a.py and written to
materials/hand-subs.json; they are not part of items.json because comparing them to the lead's
is a step of this design.
3. Procedure
- Stage A — the independent hand writes plain-Japanese substitutes.
run_stage_a.pysends one call toopenai/gpt-5.6-terra(P1, non-Anthropic, non-reader-panel role) with the nine sentences and one instruction: replace the marked span with the plainest Japanese phrasing of the same event, use no mimetic word, leave the rest of the sentence unchanged. The prompt names no English, no arm, no phonaestheme, no target, and does not mention that a second version of the task exists. Output is written verbatim tomaterials/hand-subs.json. - Stage B — freeze both substitute sets and compare under §4's
G3rubric. The comparison is done byanalyse.py --stage-bbefore any reading body is dispatched. Both sets are printed side by side intostage_b.json. IfG3fires, the reading run is dispatched with both substitute sets as parallel arms. - Stage C — the reading run. For each site, each English arm (
mk,pl) and each source condition (present,sub_leadand, ifG3fires,sub_hand), three seats read the pair(ja, en)and answer with two flags:ADDSandOMITS, one per seat, each ayes/no, with a one-line reason. Nine sites × two English arms × two source conditions × three seats = 108 reading bodies for one substitute arm. With both arms it is 162. - Verifier.
verify.pyrecomputes every reported number fromrun.jsonlby a route that re-parses raw bodies with different regexes, re-derives which arm each body scored from the frozen slots map, and recomputes the exact sign-test p-value by enumerating all 2ⁿ sign sequences. It does not importanalyse.py. Mutation tests: flip one reading flag; swap one seat's arm assignment; drop one substitute site; and one negative control that changes fields the verifier ignores and must produce no change.
4. Gates and rubrics
G0— substitutes contain no mimetic word. For each of the two substitute sets, the replacement span is checked mechanically against a small mimetic-form pattern list (reduplicated bimoraic form; a bimoraic form + り/と; the canonical common mimetics listed inrun_stage_a.py's prompt) and against the lead's frozen expansion. A hit at any site drops that site from that arm.G0PASSES if ≥ 7 of 9 sites are clean in the substitute arm actually dispatched.G1— variance floor. The enacting flag is neither allnonor allyesacross the reading run.G2— symmetry (identity items). Not built here: this design has no identity items because the primary is already a within-item pairing (same English, different source), and the S204 identity items were built for a different arm structure. Symmetry is instead measured on the mimetic-present cell againstRS-20260817e(P3continuity, above).G3— substitute-parity rubric. For each site, the two hands' substitutes agree if all of the following hold. (a) Both replacements contain no mimetic word (perG0). (b) Both replacements preserve the same event (same verb-and-object semantics), judged by the lead in writing on this page under a table filled in againststage_b.json. (c) Both replacements are of the same grammatical class (adverb-of-manner vs. verb-phrase vs. adjective) or, if not, the difference does not change what property the source-of-that-arm asserts. If ≥ 7 of 9 agree,G3PASSES and the lead's substitutes are the primary arm; if 6 or fewer agree,G3FIRES and both substitute sets are dispatched as parallel arms.
The rubric is what makes (b) and (c) tractable when the two hands' wordings differ: the question is not are the two substitutes identical — they will not be — but does either substitute reintroduce a phonaesthetic property, or change the event.
5. Stopping rule against critic regress (S206, note (bqp))
A stopping rule was fixed before the third critic pass at S206, and it is imported here verbatim:
one round of critique per version; findings addressed on the page; next round is only bought if
the last one killed a numbered primary, and never bought to widen the primary. Two rounds this
session on this design is the ceiling; if a second round returns BLOCKING findings after they have
been addressed, the design is NEEDS REDESIGN and step 2 of the arm rebuilds it. This design will
not be re-critiqued to death.
6. Cost pre-flight
Ceiling for this experiment: $1.20 against the day's remaining $4.11.
- Stage A (P1, one call at 3,000 cap, doubled once under bqk): worst case $0.030 — actual on S206's identically shaped stage was $0.010494 at 3,000 and would double to $0.021 on the second call; the prompt here is shorter.
- Stage-B analysis: $0 (no dispatch).
- Independent pre-run critic (
P1, one call at 3,000+3,000 cap, doubled once): worst case $0.15. S206 spent $0.063–$0.085 per pass at similar caps. - Stage C reading run: at 108 bodies × three seats × mean ~$0.005 per body from
RS-20260817eand under (bqk)'s doubled cap for every call at the dearest seat's rate, worst case $0.87. Actual on S204's reading was $0.480 for 108 usable bodies. - Verifier + writing: $0.
Worst case ≈ $1.05 under (bqk); stop-loss set at $1.35, ceiling $1.50 wall-clock and $1.20
declared in NEXT.md's spend section.
7. What this run cannot answer
Kept here rather than in the result page's limits section because these bind before any number exists.
- The stating arm's failure to reproduce §7.19's mirror is not this design's question.
RS-20260817e§3 already foundP2(stating-armOMITS) did not move under deletion; there is no reason a substitute would fix that, and the design does not predict it. - Six of nine mimetics were excluded from S204's deletion primary by the same grammatical
constraint (
RS-20260817e§3deletion_refused); this arm inherits that exclusion and does not test substitutability where deletion was impossible. - A substitute is still a lead-written text, in the primary arm.
G3's bar of 7 of 9 is where the arm licenses the lead's set as primary; below that, both sets run and the primary is a joint-pass finding.G3cannot rescue a case where both hands are systematically wrong in the same direction (e.g. both replace with a still-mimetic form). - One work, one author, one language pair, one translator. Both chapters are Botchan;
T-botchan-ch2-R06-v1andT-botchan-ch3-R06-v1are the only translations that supply items. - Substitutes preserve heterogeneous residues (critic MAJOR 3). They keep sound occurrence (汽笛を鳴らして, 音を立てて), force (勢いよく), mouth size or eating intensity (大口で, 盛んに), thinness (薄手の), speed (早く/すぐに), and duration/aspect (ずっと) — different properties at different sites. A flip cannot be attributed specifically to loss of mimetic form rather than to a changed specificity or aspect at that site. A single primary direction across the seven sites is still informative — it says the enacting arm's flip does not depend on which residue happens to be preserved — but no per-item verdict is licensed.
- The mechanical mimetic-form list is naive (critic MAJOR 5). Every lead classification is a lead judgement; expert Japanese raters are not available in-session. ずっと's status is the one contested case named, and the result page carries that flag.
- Reader-seat reuse from S204 (critic BLOCKING 7). The API is stateless, so no seat retrieves its prior answer; the plausible variant is that each seat's stable response habit on identical items reproduces itself. That stability is what P3′ measures; it does not confound the within-item present-vs-substitute comparison because that comparison is inside the same seat. Reported as a limit, not a stop.
8. Stage-B outcome (frozen before stage C is dispatched)
analyse.py --stage-b ran on materials/items.json and materials/hand-subs.json and wrote
stage_b.json. Agreement 8 of 9 under the G3 rubric; G3 PASSES; the primary arm is
sub_lead. Per-site verdicts:
| uid | lead's substitute | hand's substitute | verdict |
|---|---|---|---|
| M01 | 汽笛を鳴らして | 汽笛を鳴らして | AGREE — identical text |
| M03 | 遅くあるき出した | ゆるやかに歩き出した | AGREE — same event, same manner-axis, both plain adverbs |
| M04 | 音を立てて | 音を立てて | AGREE — identical text |
| M05 | 勢いよく飛び込んで | 勢いよく飛び込んで | AGREE — identical text |
| M07 | 大口で食っている | 盛んに食っている | AGREE — same event; both drop the phonaesthetic property |
| M08 | ずっと笑ってる | 薄笑いを浮かべてる | DISAGREE — one preserves temporal aspect, the other preserves the specific "wry smile" quality of にやにや |
| M12 | 薄手の | 薄手の | AGREE — identical text |
| N06 | 何か音を立てて食ってた | 何かを音を立てて食ってた | AGREE — near-identical text |
| N07 | 早く講義を済まして | すぐに講義を済まして | AGREE — same manner-axis, both plain adverbs |
G0 lead PASSES at 9 of 9 (no mimetic in any lead substitute against the frozen pattern list).
G0 hand PASSES at 9 of 9. All nine hand-substituted sentences keep the non-marked context
byte-for-byte, so context_ok is True at every site.
What the M08 disagreement means for the primary. The primary arm (sub_lead) uses ずっと.
This is a stronger subtraction than the hand's — it drops both the temporal-continuous
grounding and the wry-smile-kind aspect that にやにや carries. If the primary passes here, that is
what it passes on. If the hand's substitute had been the primary at M08, the comparison would be
against a substitute that preserves the wry-smile aspect, and the enacting English grinning and
grinning to herself might read as less of an addition. The M08 comparison is therefore reported
in the result page's limits section as a known-anisotropic site, and if the primary lands close
to its 6-of-9 bar, the M08 direction is flagged as a case where the answer depends on which hand
wrote the substitute.
No amendment is made to run BOTH substitute sets. G3 passed, and the frozen design says the
primary arm is sub_lead alone. Deviating post-hoc would let the design float; the disagreement
belongs in the limits, not in the primary.
9. Contamination declaration
The Japanese sentences here are the same as in E-20260817e; contamination measurements against
Morri 1918 apply verbatim (RS-20260817e §2). Nothing in this experiment reads or references
Morri's English; contamination-status: none for this run because no comparator is present.