Repository path: workshop/experiments/E-20260817e-mimetic-reading/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260817e-mimetic-reading |
| status | frozen |
| created | 2026-08-17 |
| updated | 2026-08-17 |
| senses | accuracy, perceived-source-carriage, style-correspondence |
| provisional | true |
| internal-judgment-only | true |
| links | wiki/arms/ARM-mimetic-carriage.md, wiki/findings/results/RS-20260816e-mimetic-carriage.md, workshop/experiments/E-20260816e-mimetic-carriage/design.md, workshop/translations/botchan-ch3/R06-v1/translation.md, workshop/translations/botchan-ch2/R06-v1/translation.md, config/models.md, wiki/goodness-senses.md, framework/v0.2/README.md |
E-20260817e — does the Japanese mimetic license the English phonaestheme? ARM-mimetic-carriage step 2
v3. Two independent adversarial pre-run passes, both NEEDS-REDESIGN, both answered in writing in
critic-response.md; v1 and v2 were never dispatched. The design that runs is the third, and its
central control was invented by the second critic's objection, not by the designer.
1. The question, and two designs that died before it
RS-20260816e §3: at a Japanese mimetic site an English rendering that enacts the sound or
manner and one that states it are not paraphrases of each other. Eight pairs of eleven failed
a propositional-parity gate. Every design that would price the marking against a constant was
killed by that, and the arm registered the successor question before this chapter was translated:
If a depiction and a statement are different claims, then a translator at a mimetic site is choosing between two readings of the source. Which one do readers who can see the Japanese say the Japanese makes?
v1 asked it as a forced choice — which rendering asserts what the Japanese asserts, no more and
no less — and the critic's BLOCKING 1 showed that this demands exact cross-linguistic identity,
returns NEITHER almost everywhere, and makes the run vacuous.
v2 replaced the forced choice with two independent binary judgments per arm (ADDS, OMITS;
SAYS = neither) and read the enacting arm's SAYS rate at real mimetic sites against its rate at
eight chapter-3 sentences containing no mimetic word. The second critic killed that too, on two
counts that are both right:
- BLOCKING 1 — the threshold sat inside the noise. At n = 15 and n = 8 the standard error of a difference of proportions is √(.25/15 + .25/8) ≈ 0.219, larger than the registered 0.20 bar. About 98 items per condition would be needed for 80% power. The bar could be cleared by sampling variation alone.
- BLOCKING 2 — "contains no mimetic word" is not "states no sound or manner." Seven of the eight
null sentences state a manner outright —
なるべく大きな声をして,急いで,威勢よく,遠慮なく,泳ぎ巡って,追っ払う,引き下がって. They were not a baseline; they were a second, uncontrolled set of items.
2. v3: the control the critic's objection implies
The comparison moves inside the item. The same two English arms, the same prompt, the same seats, are put to the Japanese sentence with its mimetic and to the same sentence with the mimetic deleted.
| Japanese | A / B | |
|---|---|---|
paired condition M03 |
おれは…革鞄を二つ引きたくって、のそのそあるき出した。 | shambled off / walked off slowly and heavily |
paired condition X-M03 |
おれは…革鞄を二つ引きたくって、あるき出した。 | the same two strings, unchanged |
Three things follow, and they are the whole reason for the redesign.
- Every item-level defect the critic named is now a constant. plunged in noisily is awkward, rattled and clattered is inflated, flimsy, flappy doubles the descriptive load — and each of them is identical in both conditions. They cannot produce a difference between the conditions. The critic's BLOCKING 8, its longest finding, is answered by construction rather than by argument.
- Any demand characteristic in the wording is a constant too. BLOCKING 3 and 10 said the prompt invites the seat to find additions in the marked arm. It does — equally in both conditions. The leading example list was removed anyway, the English target span is now marked, and a reason is required per arm; but the design no longer depends on the prompt being neutral.
- The statistic changes from an unpaired difference of proportions to a paired sign test, on which BLOCKING 1's power arithmetic does not bite in the same way: nine paired items flipping in one direction is p = 2⁻⁹ ≈ 0.002 exact, and the design registers the sign test in advance.
The entry rule for the paired set is grammaticality of the deletion, applied before any call:
delete the mimetic and nothing else, and keep the item only if the Japanese survives as Japanese.
Nine of fifteen qualify. The six that do not are named with their reasons in items.json and
here, because a silently shrunken denominator is the failure this project keeps writing notes about:
| refused | why the deletion is not a deletion |
|---|---|
M09 |
the rest of the sentence states the noise (無暗に仰山な音がする), so removing がらがら does not remove the property |
M10, M13 |
ぱちつかせて is the predicate; deleting it leaves no verb |
M11 |
「急にがやがやする」→「急にする」 is ungrammatical |
N01 |
「足の裏がむずむずする」→「足の裏がする」 is ungrammatical |
N05 |
「砂でざらざらしている」→「砂でしている」 is ungrammatical |
One declared edit inside a deletion: at M08 the degree adverb やに is removed with にやにや,
because 「やに笑ってる」 leaves an intensifier stranded on a bare verb.
What it teaches about translating literature (the subject rule, continue-prompt.md §4.5): it
tells a translator whether the English phonaestheme at a mimetic site is licensed by the mimetic —
by testing whether readers' judgement of that same English changes when the mimetic is taken out of
the Japanese. The site set is a grammatical class; the readers are not the translator and are not
one of the project's instruments being audited.
3. Materials
Built by materials/build_materials.py into materials/items.json; every chapter-2 row is copied
verbatim from E-20260816e/materials/sites.json and every chapter-3 row is asserted against the
frozen files it was transcribed from. 27 reading items:
| kind | n | what it is |
|---|---|---|
| real | 15 | mimetic sites with two buildable arms: 11 from Botchan ch. 2 (M01 M03 M04 M05 M07 M08 M09 M10 M11 M12 M13) and 4 from ch. 3 (N01 N05 N06 N07) |
| deleted | 9 | X-M01 X-M03 X-M04 X-M05 X-M07 X-M08 X-M12 X-N06 X-N07 — the same English, the mimetic gone from the Japanese |
| identity | 3 | D-M06 D-M15 D-M02: chapter-2 sites where the frozen file carries the same string in both arms, because no enacting rendering could be built |
The translation limb. T-botchan-ch3-R06-v1 — 「坊っちゃん」chapter 3 whole, 5,821 Japanese
characters → 3,410 English words under R06, frozen at 673f0bed with a nine-site mimetic
census and, at every site, the two renderings the translator was choosing between, written before
this design existed. Its four buildable sites are the N items; its five unbuildable ones join
chapter 2's four in the buildability stage.
Contamination, measured after the freeze and before this design was written. Against Morri 1918
chapter 3: 10 shared 7-grams, 1 twelve-gram, 0 fifteen-grams, longest common run 12 tokens,
tools/dependence_check.py verdict DEPENDENT?. The run is "i was taken for a locksmith when i
went to see the", rendering 「錠前直しと間違えられた事がある。ケットを被って、鎌倉の大仏を見物した
時は」. No figure in this design depends on independence from Morri: he is not a comparator here,
appears in no item, and is read by no seat. Against the lead's own chapter 2 (note (bhb)):
1 shared 7-gram, longest run 7, clean.
The A/B slot for every (item, seat) pair was fixed by random.Random(20260817) inside
build_materials.py before any call was made and is stored in items.json. The analysis reads
the slot back rather than assuming it. The two conditions of a pair get independent slots, so a
seat cannot recognise a repeat by position.
4. The seats
P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5, QR qwen/qwen3.7-max for the reading task;
P2 and P3 for the census, the class and the buildability second hand; P1
openai/gpt-5.6-terra for the two critic passes and nothing else. P4 and P5 are out on any
task shape (notes (bps), (bne)); GL is out on any long prompt, which is why the chapter-3
census does not use it.
These are three specified systems, not a reader population (v2 critic, NON-BLOCKING 14, taken). Every claim below is a claim about what these three seats said on a curated Botchan item set. The phrase "the source licenses" is not used anywhere in the result without that qualifier.
5. Predictions, registered before dispatch
- The enacting arm's
ADDSrate rises when the mimetic is deleted. This is the experiment. The lead predicts it rises at 6 or more of the 9 pairs. - The lead predicts the enacting arm is already
ADDSat a majority of the 15 real sites — against the lead's own instinct while translating, which was that the phonaestheme was the faithful choice. If prediction 1 holds and 2 holds, the reading is that the mimetic licenses the phonaestheme partially: it makes the addition smaller, not zero. - The stating arm's
OMITSrate falls when the mimetic is deleted — for the same reason, mirrored. A design in which only the enacting arm moves is a design measuring one arm. - The identity items produce identical flags within a seat.
(Prediction 4 of v2 — a sound/manner class effect — has been struck on the critic's NON-BLOCKING 16: the design says it cannot test it, so it is not a prediction of this run. The by-class counts are reported as counts and refused as evidence.)
6. Gates, and what each one withholds
| gate | what is bought | bar | what fails if it fails |
|---|---|---|---|
| G0 | independent census of ch. 3, P2 and P3, the Japanese chapter only |
each of the four N sites listed by at least one seat |
a site not listed is dropped, with it its deleted twin, and the drop is named in the headline |
| G0b | independent SOUND/MANNER class of the four N sites |
the two seats agree | disagreement → the site is contested and is excluded from any by-class count |
| G1 | floor/ceiling check on the primary flag | the enacting arm's ADDS majority must not be YES at all 24 real+deleted items nor NO at all of them |
the primary is WITHHELD. A flag with no variance measures nothing, which is the v2 critic's BLOCKING 5 in the form that actually applies to a paired design |
| G2 | the 3 identity items | on an identity item the arms are one string, so a seat's A-flags must equal its B-flags. A seat differing on ≥ 2 of 3 fails | one seat failing → its rows are flagged and the primary is reported twice, with and without it. Two or more → primary WITHHELD. Reported as a symmetry check, not an attention measure (v2 critic NON-BLOCKING 13) |
| G3 | complete data on a pair | a pair missing any of its 6 cells is dropped from the primary, not patched | the count of dropped pairs is reported in the headline |
Missing-data and tie rules, registered (v2 critic BLOCKING 6). Three seats per cell; the majority is 2 of 3 and cannot tie. If a body is dead after one mechanical re-dispatch, that cell has two seats; if they split 1–1 the cell is unresolved, the pair is dropped under G3, and both the unresolved cell and the dropped pair are named. Unresolved cells never enter a denominator. G1 and G2 are read before the primary is computed (note (boa)).
7. Procedure
Serial, temperature 0, one item per call, max_tokens 600 against a reasoning cap of 120 (note
(bpv)), 3,000 against 400 on the two census calls. run.py enforces the stage order
census → classify → reading → build and refuses a stage whose dependencies are not complete; it
writes a dead row after one mechanical re-dispatch so a permanently unparsable job is never
re-bought; it checks the stop-loss before each call and the hard ceiling after each; and the reading
parser requires all six lines, rejects a duplicated flag, and requires a reason for each arm.
(All six are the v2 critic's BLOCKING 12, taken.)
Judgment is not parallelised. Every call is one seat on one item.
8. The primary
Primary — the paired flip. For each of the 9 pairs, the majority ADDS flag on the enacting
arm in the mimetic-present condition and in the mimetic-deleted condition. Count:
- flips to ADDS (present NO → deleted YES) — the mimetic was licensing the phonaestheme;
- flips away (present YES → deleted NO) — the opposite;
- concordant pairs, in either direction.
Registered test: an exact two-sided sign test on the discordant pairs. With 9 pairs, 9–0 gives p = 0.004, 8–1 gives p = 0.039, 7–2 gives p = 0.180. The bar is not a rate threshold and there is no 0.20 anywhere in this design.
The same statistic is computed for the stating arm's OMITS flag, which is prediction 3.
Reported alongside, descriptively and with no test: the 4 × 2 table of ADDS and OMITS
majorities for both arms at all 15 real sites; the per-seat marginals for every flag (the v2 critic's
BLOCKING 5 — a raw agreement figure without marginals is uninterpretable); flag-specific pairwise
agreement; and every WHY clause verbatim at the sites that flip, because on this project's record
the reasons have been worth more than the counts.
What this design still cannot do, written before the numbers exist.
- 9 pairs is 9 pairs. A 6–3 split is p = 0.51 and means nothing. The design can return a clear answer or no answer, and the second is reported as no answer.
- It cannot separate "the Japanese says this" from "these three systems read Japanese this way." Charter §4 forbids treating panel agreement as validation.
- It cannot test the sound/manner conjecture. 6 clean sound sites against 2 clean manner sites is not a contrast.
- A deletion is not a natural sentence. 「汽船がとまると」 is grammatical Japanese but it is Sōseki with a word taken out, and a seat may be responding to the mutilation rather than to the absence of the property. Nothing in this design excludes that, and the result page says so.
M12andM13share a Japanese sentence.M13has no clean deletion and so is absent from the primary; the clustering is resolved by that accident and not by design.- One work, one author, one language pair, one translator. Both chapters are Botchan.
Secondary, rebuilt blind (v2 critic BLOCKING 11). The nine sites where the translator could build
no enacting arm go to P2 and P3, which are shown the Japanese, the span and the stating English
with its slot marked, and are asked for up to three alternative renderings at least as accurate.
The prompt names no lexical class, no theory, and no example. Every candidate is printed
verbatim. Registered: a NONE answer has no evidentiary force whatever — a seat declining to
better a phrase is not evidence that English lacks a resource — and any classification of a returned
candidate as phonaesthemic is a lead judgment, marked internal-judgment-only where it appears.
9. Failure criteria, stated as failures
- G1 or G2 withholds → the primary is not printed, and the arm closes on that.
- Sign test p > 0.05 on both arms → no answer, reported as no answer, and the arm closes
retiredwith the registered statement that the choice at a mimetic site is not shown to be recoverable from the source by this design at this n — not the stronger claim that it is a style choice, which the v2 critic's BLOCKING 7 correctly refused. - Fewer than 7 pairs survive G3 → the sign test is not run and the counts are reported bare.
- Any stage crossing the stop-loss → the run halts and the result page reports what was not bought.
10. Budget
Pre-flight from the max_tokens cap each request permits (note (abc)), including the one
mechanical re-dispatch at double cap.
| stage | calls | seats | worst case |
|---|---|---|---|
| census | 2 | P2 P3 |
$0.043 |
| classify | 8 | P2 P3 |
$0.028 |
| reading | 81 | P2 P3 QR |
$0.281 |
| build | 18 | P2 P3 |
$0.064 |
| re-dispatch contingency (×1.5 on output) | — | — | $0.208 |
| pre-run critics, both already spent | 2 | P1 |
$0.115 |
| worst case total | 111 | $0.74 |
Worst case $0.74 < stop-loss $0.82 < declared ceiling $0.95 (note (bpq), two-sided). The UTC day 2026-08-17 opens at $0.00 of $5.00 and this is its first session. The two critic calls cost $0.115 and killed two designs, which is what the money is for.
11. Decision log — investigator choices, not mechanical ones
The v2 critic's NON-BLOCKING 15 asked for this list, and it is right that the word frozen prevents
later change without making a choice neutral. Free choices made by the lead: which spans count as
mimetic (the census rule); which of the fifteen deletions are grammatical; removing やに with
にやにや at M08; the three identity items; G2's two-of-three cutoff; the seed 20260817 and the
slot schedule; which nine sites go to the buildability stage; keeping chapter-2 items the critic
called defective rather than editing a frozen artifact. Inherited constraints, not chosen here:
the eleven chapter-2 pairs and their exact strings; the chapter-2 classes; the panel composition; the
budget cap.
12. Verification
analyse.py --json produces every number in the result page. verify.py recomputes them from
run.jsonl by a route that re-parses the raw bodies rather than trusting the stored parses, does
not import analyse.py, re-derives the slot decoding from items.json independently, and recomputes
the sign-test p-value by exhaustive enumeration rather than from a table. Mutation tests: flip one
reading flag, break one identity item's symmetry, delete one census word, flip one class, and one
negative control that corrupts an unused field and must change nothing.