Repository path: workshop/experiments/E-20260815e-fluent-carriage-2/amendments.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260815e-amendments |
| status | frozen |
| created | 2026-08-15 |
| updated | 2026-08-15 |
| senses | perceived-source-carriage, naturalness |
| internal-judgment-only | true |
| provisional | true |
| links | workshop/experiments/E-20260815e-fluent-carriage-2/design.md, wiki/arms/ARM-fluent-carriage.md, wiki/findings/results/RS-20260815b-fluent-carriage.md |
E-20260815e — amendments, in the order they were made
Every amendment below was made before any judging call was dispatched. The dispatch log
(runs/bodies.jsonl) is the check on that claim: no row with a cond in AB AC AD QB QC QD AS
precedes any row recorded here.
A-1 — build 1 of the arms passed the gate at 5 of 7, and the two flags are my errors, not the auditor's
Dispatched: run_stage.py gate, build 1 arms (frozen fffce5f7). Cost $0.086253000.
Positive control: 5 of 5 planted errors caught, each named exactly — place, referent, negation ×2, quantity. The seat is competent and its parity verdicts carry the weight the design registered.
Result: SAME on 5 of 7 segments. Step 1's control, on the seven-class flattening, returned SAME on 0 of 7. That difference is itself a measurement and it is reported as the run's first finding: dropping the suspension, the reduplication and the held keyword removed five sevenths of the parity problem.
The two remaining flags, verbatim:
S1 — "P explicitly says C 'called to B'; Q and R only have C calling 'Hey!' with his head out of water, omitting the addressee. Q/R also explicitly say the striped girl 'cried' her words, while P only quotes them."
S5 — "P says 'and then again as if I were setting,' implying a subsequent, repeated moment, whereas Q and R say 'at the same time,' implying simultaneity — a tense/timing difference."
Both are right, and both are defects in my flattening rather than properties of the device class. They are the same kind of error as step 1's §6(iii) pair — content changed while a form was being removed — and they were invisible to 507 mechanical checks for the same reason: no check tests meaning.
- S1, A6.02. Folding a standalone attribution into a dialogue tag made me choose a word order,
and the order I chose dropped to B. The addressee is in the source (
Bを呼んだ) and in theFSarm; it is not inFF2. A dropped argument is a content change. - S1, A6.01. The
FSarm attributes "Me too, me too!" by juxtaposition alone — "—and with that the other one … caught hold of young B" — with no speech verb anywhere. Folding it into a tag made me supply one, and I supplied cried, which asserts a manner of utterance the source does not. A supplied speech verb is a content change. - S5, A5.06. Breaking the piled period turned "and then again as if I were setting" into "I
felt at the same time as though". The source's
またis sequential; at the same time is simultaneous. A changed temporal relation is a content change.
The finding this deposits, before any repair: A6 — utterance first, attribution after — is not
unconditionally form-only either. At a site where the source attributes by juxtaposition rather
than by a verb, the flattening operator must invent an attribution, exactly as flattening a
suspension must invent a completion. That is a fourth device class joining RS-20260815b §6(i)'s
three, and it is the sharpest thing this gate has to give. It is recorded here whatever the repair
does, because the repair does not make the difficulty go away — it works around it at two sites.
A-2 — the arms are repaired rather than the run shrunk, and the reason is written
What the frozen design says (F3): 1–3 segments flagged → those segments are excluded and every
statistic recomputed on the remainder.
What is done instead: the three defects are repaired at the three sites named above, the gate is re-dispatched on the repaired arms, and — if it passes — all seven segments are judged on the registered bars.
Why this is the more faithful reading of the frozen design, not a relaxation of it.
- The registered bars are stated on seven segments.
P1,P2andP3fire at "all three seats ≥ 6 of 7". Excluding two segments would force me to re-derive every bar after seeing a measurement, which is the thing pre-registration exists to prevent. Repairing the arms keeps the frozen inferential structure exactly as it was written. - The gate exists to let a design learn for one call's price — note (boa), the standing remedy this run is the first application of. Its remedy is "dispatch the parity control first, as a gate; a design that cannot pass it has no primary". The point of learning early is to fix the arms while fixing is still free. Shrinking the run instead would be paying the gate's price and declining its benefit.
- The repairs restore the
FSarm's content; they do not move anything toward a hypothesis. Each is a one-way restoration — put to B back, remove cried, put and then again back — and each is checkable againstFS, which is unchanged. None touches theODDarm, none touches a device site's presence or absence, and none changes the 33-site census. - No judging call has been dispatched. Nothing about carriage, quality or direction has been measured. The repair is blind to every statistic the run exists to produce.
What is given up, and it is stated rather than buried: the build-1 gate result is a measurement on arms that no longer exist, and the build-2 result is a measurement on arms whose two known defects were repaired because the same instrument found them. A parity control cannot certify arms it has already been used to correct at the same sites. The result page reports both builds, and reports build 2's parity as repaired-and-rechecked, not as an independent pass.
The repairs, exactly:
| site | seg | build 1 | build 2 |
|---|---|---|---|
| A6.02 | S1 | "Hey!" called C, the young man who had got there ahead of A, with his head out of the water. |
C, the young man who had got there ahead of A, called "Hey!" to B with his head out of the water. |
| A6.01 | S1 | "Me too, me too!" cried the other one, the girl in the broad red-and-black stripes, catching hold of young B. |
"Me too, me too!" And with that the other one, the girl in the broad red-and-black stripes, caught hold of young B. |
| A5.06 | S5 | … I felt at the same time as though I were setting beside it a figure … |
… And then again I felt as though I were setting beside it a figure … |
A6.01's build-2 operator is weaker than the class's declared one and that is declared here. The
class is flattened by folding the attribution into a dialogue tag; at this site there is no verb to
make a tag from, so what build 2 does instead is remove the paragraph break and the dash-led
continuation, leaving the utterance and its attribution in one paragraph. The delay is removed;
the tag is not created. The site is kept in the census at reduced strength, and the reduced
strength is a reason the FF2 arm is weaker than declared, not stronger.
A-3 — one flag survives the repair, and it is NOT repaired: S4 is excluded and the bar is recomputed before any judging call
Dispatched: run_stage.py gate, build 2 arms. Cost $0.072254 + $0.007371 re-dispatch.
Positive control: 5 of 5 again. Result: SAME on 6 of 7 segments. S1 and S5 — the two build-1 flags — now pass, so the three repairs did what A-2 said they would.
One new flag, verbatim:
S4 — "In P the schoolboys run off 'at a scatter' (all at once, dispersing), while Q and R say they ran off 'in ones and twos' (in small groups) — different descriptions of the same event."
The auditor is right again, and this one is a Class A2 site. パラ/\と駆け出して行く is rendered off at a scatter in the carrying arm; the flattening operator for A2 is generic verb plus a stated manner adverb, and the manner I stated — in ones and twos — is a different manner. Dispersal at once is not departure in small groups.
It is not repaired, and the reason is the one A-2 gave for repairing the first three. A-2 was defensible because it kept the frozen bars intact and because no judging call existed. Repairing a second time, at a site the same instrument has just named, would make that instrument a co-author of the arms it is certifying — and a parity control that has written the text it passes is not a control. One amendment on evidence is a repair; a habit of them is fitting. The frozen design already says what to do with 1–3 flags, and this run does that.
F3 applied, and the recomputation registered HERE, before the first judging call:
- S4 is excluded. The primary analysis runs on six segments — S1, S2, S3, S5, S6, S7.
- The bar becomes: all three seats ≥ 5 of 6. One-sided exact P per seat = 7/64 = 0.109375; jointly, on the design's independence assumption, 1.31 × 10⁻³. This is looser per seat than the frozen ≥ 6 of 7 (P = 0.0625) and it is stated as looser rather than presented as equivalent.
- All seven segments are still dispatched and the seven-segment figures are reported alongside,
as
F3requires. Where the two disagree, the six-segment figure governs and the disagreement is the headline.
What A2's failure deposits, independently of this run's statistics. Three device classes failed the form-only test in step 1; A6 and A2 have now failed it too, at one site each. That is five of the seven declared classes shown, by an independent reader, to carry something propositional. The surviving unchallenged classes on this story are A5, the piled period, and A7, the withheld subject — both of which are pure constituent order, which is exactly what one would predict and is now measured rather than assumed.
A dispatch note. The build-2 PAR call on S2 returned finish_reason: "length" with zero
visible characters — note (bny) firing again, at reasoning.max_tokens 180 inside a content
cap of 900. Re-dispatched at 2,000 total with reasoning 400: stop, "SAME.", $0.007371. The hole
was repaired, not reported as a flag; the scorer takes the last body per key so the empty one cannot
be double-counted.
A-4 — the pre-run critic returned NEEDS-REDESIGN, 13 findings, 4 BLOCKING; all 13 accepted
P4 moonshotai/kimi-k3, one call, stop, 18,550 characters, $0.180201. Full text:
runs/critic.jsonl. Note (bny) applied and worked: explicit reasoning: {"max_tokens": 4000}
inside 14,000 total, no effort pin, no truncation.
Two BLOCKING findings changed the arms and a third changed what the run may conclude.
F-2 [BLOCKING] — the 33-site census was the translator marking his own homework, and it flattened
a site a blind auditor had already refused. C-8d — a site with no carriage is never flattened —
was wired to the translator's FAILED verdict only. The auditor's refusals were not wired in, and
step 1's auditor had refused A2.4 specifically. Accepted in the strong form: the blind carriage
audit was re-run on all 33 surviving sites, three batches, and the auditor's NO is now treated
exactly as the translator's FAILED — the site is not flattened.
The audit's result, which is a finding in its own right: 25 of 33 carried, 8 refused.
| class | flattened | refused | the auditor's words |
|---|---|---|---|
| A2 mimetic | 1 of 6 | A2.1–A2.5 | "onomatopoeia dropped, plain verb only", "mimetic reduplication flattened into plain verb", "onomatopoeic adverb lost" |
| A5 piled period | 9 of 10 | A5.05 | "English front-loads noun, loses end-suspension" |
| A6 standalone attribution | 12 of 12 | — | — |
| A7 withheld subject | 3 of 5 | A7.2, A7.3 | "English inverts, action first instead of verb-final" |
The mimetic class collapses. Five of six sites refused: a single English sound-symbolic verb
does not, on an independent reader's judgment, do the work of a reduplicated mimetic. That is a
direct hit on R29's most confident craft claim (R29-v1 log, §A2: "English has a large stock of
these and they cost nothing"), it replicates step 1's A2.4 refusal on five more sites, and it is
the sharpest craft finding this session has. A6 is the only class the auditor passes entire.
Consequence for the manipulation: 25 sites, and the one surviving A2 site (A2.6) is the only
lexical one. The remaining 24 are pure constituent order — which makes the design's "purely
syntactic" claim true by measurement rather than by assertion, which is what F-3 [BLOCKING] said
it was not.
F-1 [BLOCKING] — the O4 operator was Japanese-transfer English. Comma-bracketed medial
adverbials — "he was, back there, the whole time, giving her sidelong looks" — are the surface
signature of an under-digested translation from a verb-final language, which is direction
information; an ODD arm built with it re-confounds H-CARRIAGE with H-LITERALITY inside the primary
cell. Accepted: O4 is retired (joining O6 and O7) and all eleven O4 edits were rewritten
under nominalisation, stilted collocation, circumlocution and a new O10 cleft. The critic's
own proposed audit was added and is amendment A-5.
The other eleven findings, accepted and where each is discharged: F-3 (strike "purely syntactic")
→ the audit made it true; the phrase is still struck and the claim is stated as measured, not
asserted. F-4 (G2 is an acceptance region, not a parity test) → "parity" is removed from every
downstream licence; G2 passing licenses only "no quality difference detected at n = 21, power
roughly 0.5–0.6 against a true rate of 0.65", and that sentence is what the result page will say.
F-5 (the A-2/A-3 line is convenient) → accepted in full, in the critic's own words: the repair
rule is registered prospectively as at most one repair round, and build 3's parity is reported as
co-authored by the auditor, not as an independent pass. F-6 (excluding S4 both weakens the
manipulation and loosens the bar) → moot at build 3, see A-5, and both figures are reported anyway.
F-7 (the null row is stated at ablation strength) → rewritten. F-8 (recognition) → the probe was run;
see A-5. F-9 (author degrees of freedom) → a standing limitation, partly closed by the F-1 and F-2
audits, and named as unclosed where it is unclosed. F-10 (gate denominators) → registered in A-5.
F-11 (A6.01 counted at full strength after its operator was weakened) → the census is reported as
25 sites, one at reduced strength. F-12 (register shift inside dialogue) → the PAR prompt now
asks about voice and register; see A-5. F-13 (a critique that cannot stop the run is decoration) →
registered: every BLOCKING finding either has its amendment frozen before dispatch, as F-1 and F-2
do, or stands as a written caveat on the primary's licence, as F-4 does.
A-5 — build 3 passes parity 7 of 7, and the language audit fails against my own repair
Three pre-judging measurements, all dispatched before the first judging cell.
(1) PAR, build 3: 5 of 5 plants caught, SAME on 7 of 7 segments. No exclusion; the frozen
seven-segment bars (≥ 6 of 7 per seat) apply unamended, and A-3's recomputation is vacated.
But S4 — flagged DIFFERENT at build 2 and SAME at build 3 — moved on a phrasing that did not
change. A2.6's "off at a scatter" against "in ones and twos" is present in both builds and
the same seat judged it a content difference once and a stylistic one once. The gate is not a
deterministic instrument, and one call is one draw. Two consequences, both registered here: the
S4 exclusion is also reported as a sensitivity analysis at the A-3 bar (≥ 5 of 6), and the
seven-of-seven pass is reported with this non-replication attached to it rather than as a clean bill.
(2) RC, the recognition probe (F-8): all three judging seats answered Japanese; none
recognised the story. "Unfamiliar, though geisha and yukata indicate Japanese origin clearly."
The recall confound F-8 feared most — a seat answering from memory of the actual original — is
closed. The language leak is total and content-borne, replicating step 1's G5 at 9 of 9.
(3) LS, the language-signature audit (F-1's own proposed check): it FAILS, and it fails against
the repair that was made to satisfy it. 24 of 35 oddity edits were attributed to Japanese. The
reasons name the surviving operators, not the retired one: "nominalization" eleven times,
"cleft mirrors Japanese emphasis structure", "'in the event that' translates Japanese
conditional", "'for what reason' literal interrogative".
This is not fixed by a fourth rebuild and the reason is worth stating plainly. Retiring O4
moved the attribution rate not at all; the operators that replaced it draw the same verdict. The
honest reading is that on a source whose content names its language, no English-side awkwardness
can be built that an independent reader will not attribute to that language — and step 1 measured
the same thing from the other side, its FS arm identified as Japanese from "verb-final inverted
syntax kept". Both manipulations leak, symmetrically. That is a property of the material, not a
defect a fourth build removes, and iterating a fourth time against an instrument I have already
edited to twice is precisely what F-5 forbids.
What it costs, registered before any judging number exists:
P1(AB,FSvsODD) loses its strong reading. A firingP1onFSmay not be reported as "carriage beats non-fluency" without the live alternative in critic F-1's words — "competent-Japanese-sounding beats incompetent-Japanese-sounding", i.e. H-CARRIAGE winning via theODDarm.P1is demoted from primary to reported-with-caveat.P2(AC,FSvsFF2) is promoted to the run's principal cell, and the promotion is argued, not convenient: neither arm inACis the awkward one, so F-1's confound cannot reach it. Content parity is 7 of 7, quality is gated byG2, and the 24 sites separating the arms are now — by the F-2 audit rather than by my say-so — pure constituent order.P3(AD) gains rather than loses. If the seats callODDthe source-follower while an independent seat callsODDJapanese-transfer English, that conjunction is H-LITERALITY, stated more sharply than the four figures on the record state it.- A further confound is now measured rather than suspected (F-12). The
PARseat reported a register shift in all seven segments: "R's dialogue is more formal/stilted … shifting the speakers' voices away from the colloquial tone of P and Q." TheODDarm varies fluency AND character register together and this run cannot separate them. Standing caveat onP1andP3.
Registered gate denominators (F-10). All seven segments are dispatched and every gate is stated
on 21 cells: G1 ≥ 19, G2 non-rejection 8–13, G3 ≥ 19, G4 ≥ 17. The six-segment sensitivity
recomputes primaries only, at ≥ 5 of 6 per seat.