Repository path: workshop/experiments/E-20260815e-fluent-carriage-2/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260815e-fluent-carriage-2 |
| status | frozen |
| created | 2026-08-15 |
| updated | 2026-08-15 |
| senses | perceived-source-carriage, naturalness |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-fluent-carriage.md, wiki/findings/results/RS-20260815b-fluent-carriage.md, wiki/findings/results/RS-20260814f-carriage-elevation-2.md, wiki/findings/results/RS-20260808d-carriage-decoupled.md, workshop/translations/atsui-suna/R29-v1/translation.md, workshop/translations/atsui-suna/R14-v2/translation.md, workshop/translations/atsui-suna/device-sites.md, workshop/regimes/R29-fluent-carriage.md, wiki/goodness-senses.md, config/models.md |
E-20260815e — carriage against non-fluency, on a manipulation that survives a parity control
ARM-fluent-carriage step 2. Frozen before any call is dispatched. The three arms and the
rebuilt translator's record were frozen at fffce5f7, before this file existed.
This is a rebuild of the manipulation, not a replication of the result (ARM-fluent-carriage
§Steps, registered before step 1's numbers existed). Step 1 built the three arms, passed 342
mechanical checks, and then lost its whole carriage half to a content-parity control dispatched
last. Note (boa) is the standing consequence and this design is its first application.
1. The question, unchanged
When carriage and non-fluency point at opposite renderings, which one does the judgment follow?
RS-20260814f §7 named the missing arm — fluent and source-ward — and step 1 built it. What
step 1 could not do is read the answer, because the texts it contrasted were not propositionally
identical.
2. What changed, and why each change is forced
(i) Three of the seven device classes are dropped. RS-20260815b §6(i): an independent
auditor, having caught 3 of 3 planted factual errors, held that flattening a suspension invents
its completion, that a reduplication asserts more than a single token, and that varying a
held keyword changes the degree asserted. A device is form-only if removing it changes nothing
the text says, and the test is the auditor's to apply, not the translator's. A1, A3 and A4 are
therefore not manipulated; they stay carried in all three arms.
(ii) What is left is a purely syntactic manipulation, and that is a sharpening. The dropped classes are lexical and punctuational; the four that remain are all word order and sentence shape — mimetic manner (6 sites), the piled period (10), the attribution standing alone (12), the withheld subject (5). 33 sites. Whatever this run measures, it measures about syntax, with every word of the vocabulary and every mark of the punctuation held identical.
(iii) A7.1 is excluded. R29-v1 records it FAILED — English cannot leave a finite main
clause subjectless — so the carrying arm does not carry it. Step 1 flattened it anyway, which put a
difference between the arms that was not a carriage difference. Registered here: a FAILED site is
never flattened, asserted by checks.py C-8d.
(iv) The oddity arm retires the two operators that produced step 1's defects. O6 dangling
modifier — it reversed agency at S2 and garbled referents at S5, and it does so by construction,
because a modifier that dangles necessarily re-attaches to a different subject. O7 hedge insertion
— basically, sort of, needless to say are semantically empty and pragmatically not, and the
parity auditor was right to flag them (§6(ii)). What remains — nominalisation, stilted collocation,
adverbial displacement, heavy circumlocution — adds no proposition and moves no agency. A committed
filler blacklist stops the retired words returning through another operator.
(v) The parity control is dispatched FIRST, as a gate, and nothing is judged until it passes.
3. The three arms
| arm | what it is | carriage | fluency | words |
|---|---|---|---|---|
FS |
T-atsui-suna-R29-v1, unchanged from step 1 |
high | high | 1,485 |
FF2 |
T-atsui-suna-R14-v2 — the same text with exactly 33 syntactic device sites flattened, every other string byte-identical |
reduced | high | 1,525 |
ODD |
FF2 plus 35 propositional-neutral markedness edits, every one at a paragraph where FS and FF2 are byte-identical |
reduced | low | 1,564 |
FF2 is NOT a zero-carriage arm and the design does not treat it as one. The 29 dash
suspensions, the 8 reduplications and the 11 tokens of the held keyword are carried in FF2 exactly
as in FS. The FS/FF2 contrast is a syntactic dose contrast, not an ablation — 33 sites
against step 1's 84 — and every reading below is stated at that strength. Registered consequence: a
null on this manipulation is weaker evidence than a null on step 1's would have been, and is
reported as such rather than as an absence of the effect.
Orthogonality is mechanical, not claimed: FS/FF2 differ only at flattened spans, FF2/ODD
only at edit sites, the two sets disjoint. 507 checks, 0 failures, 9 of 9 mutation tests caught.
Eighteen dropped-class sites sit inside a rewritten span and are not exempted: each is asserted
to survive in kind — dash present, doubling verbatim, held-keyword count unchanged over the same
stretch of story (C-3b).
Length ratios FF2/FS 1.0269, ODD/FS 1.0532, ODD/FF2 1.0256; declared
tolerance ±6% whole-text, ±10% per segment, asserted by checks.py (C-7). Worst segment S3 at 1.090.
4. Segments and seats
Seven segments, the same cuts as step 1 (analysis/segments.py, STARTS unchanged), so the
two steps are directly comparable. Alignment is computed by difflib and asserted.
| seg | ¶ (FS) | FS w |
FF2 w |
ODD w |
device sites | oddity sites |
|---|---|---|---|---|---|---|
| S1 | 1–15 | 264 | 273 | 277 | 5 | 3 |
| S2 | 16–27 | 137 | 140 | 143 | 3 | 4 |
| S3 | 28–49 | 189 | 199 | 206 | 7 | 7 |
| S4 | 50–67 | 217 | 216 | 225 | 8 | 5 |
| S5 | 68–75 | 178 | 187 | 189 | 2 | 5 |
| S6 | 76–94 | 282 | 293 | 300 | 5 | 6 |
| S7 | 95–105 | 217 | 216 | 223 | 3 | 5 |
| all | 1,484 | 1,524 | 1,563 | 33 | 35 |
Seats: P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5 — the
same three as ARM-carriage-direction and as step 1, so the instrument is imported unchanged. P5
excluded on note (bne). P4 moonshotai/kimi-k3 runs the parity control and the pre-run
critique and takes no judging cell.
Conditions. 7 × 7 × 3 = 147 cells, one call each.
| condition | pair | source shown | question | role |
|---|---|---|---|---|
AB |
FS vs ODD |
no | carriage | PRIMARY P1 — the decisive cell |
AC |
FS vs FF2 |
no | carriage | CO-PRIMARY P2 — does the sense respond to SYNTAX alone? |
AD |
ODD vs FF2 |
no | carriage | P3 — the false positive, in forced-choice form |
QB |
FS vs ODD |
no | quality | GATE G1 |
QC |
FS vs FF2 |
no | quality | GATE G2 — the fluency-parity gate P2 depends on |
QD |
ODD vs FF2 |
no | quality | GATE G3 — the manipulation check on ODD, which step 1 left at 5 of 21 |
AS |
FS vs ODD |
yes | carriage | GATE G4 — does showing the original change the answer? Not dispatched at all in step 1 |
Prompts are step 1's verbatim, so nothing in the instrument moved between the steps.
Presentation order fixed per cell by sha256("E-20260815e|" + cond + "|" + seg + "|" + seat) mod 2.
The word Japanese appears in no dispatched prompt in any condition, nor the author's name, the
title, any regime id, or any of carriage, flatten, device, fluent, odd — asserted by the
runner and by the verifier; a violation voids the run.
LG is not re-run, and the omission is declared here rather than discovered later. Step 1
measured the language leak at 9 of 9 cells — total, and arm-invariant, the stated cue being
content that is byte-identical across the arms (RS-20260815b §7). The dropped classes make the
arms more alike lexically than step 1's, so the leak can only be more arm-invariant, not less.
Nine calls saved; the step 1 figure is imported and cited, not re-measured.
5. The gate that runs first
PAR — content parity, dispatched before any judging call. P4, which takes no judging cell,
is given the three arms per segment and asked whether they state the same things. Five planted
factual errors — place, referent, negation ×2, quantity — are the positive control, one call each,
against the undamaged FS text.
- If
P4does not catch 5 of 5 plants, its parity verdicts carry no weight and the run proceeds with parity unestablished, reported exactly as step 1 reported it. - If
P4catches the plants and reports a real content difference on a segment, that segment is excluded and every statistic is recomputed on the remainder. - If it reports a real content difference on 4 or more of 7 segments, the manipulation is not
content-neutral and NOTHING IS JUDGED. The arm closes
retiredand the finding is that this project cannot build a content-neutral carriage manipulation on this material — whichRS-20260815b§6 already half-says and which is worth saying plainly for one call's price rather than a run's.
This is the whole procedural point of step 2 and it is registered before the first call.
6. Hypotheses and what each predicts
- H-CARRIAGE — the seats track which way the translation is marked. Predicts
FSonABandFSonAC; makes no prediction onAD, where neither arm's syntax carries. - H-LITERALITY — the seats choose whichever text reads more like a translation, without reading
for direction. Predicts
ODDonABandODDonAD. OnACit predictsFS, and this is stated plainly because it is the design's honest limit: source-ward syntax and translation-sounding syntax are the same surface.ACtherefore does not separate the two hypotheses;ABdoes, andABis the primary for that reason. - H-LENGTH — the seats choose the longer text. Predicts
ODDonABandADandFF2onAC. Registered because both manipulations add words. - H-QUALITY — the seats read "followed the original" as general merit. Predicts
FSonAB, and is indifferent onACiffG2shows quality parity.
What AC is for, given that it does not separate H-CARRIAGE from H-LITERALITY. Step 1 measured
AC at 21 of 21 with all seven classes varying. Here only syntax varies. If AC still fires, the
sense responds to sentence shape alone; if it does not, step 1's 21 of 21 was carried by the lexis
and punctuation that this version holds fixed — and that is a substantive statement about what the
sense reads, in either direction. It is registered as a co-primary on that ground, not on separation.
The diagnostic table, written before the numbers exist:
P2 (AC) |
P3 (AD) |
P1 (AB) |
what it means |
|---|---|---|---|
FS |
ODD |
FS |
carriage wins when opposed at 40% of step 1's dose — H-LITERALITY survives as a contributor, refuted as the account |
FS |
ODD |
ODD |
non-fluency beats carriage. The perceived-source-carriage entry is rewritten as a non-fluency detector and framework/v0.2 §8 Q-e loses its reader-side leg |
FS |
ODD |
indeterminate | step 1's indeterminacy replicates at reduced dose; the entry gains the width, and the arm closes saying the decisive cell is not resolvable on this material |
| null | ODD |
any | the sense does not respond to syntax alone. Step 1's AC was carried by lexis and punctuation, and the entry gains that — a real narrowing of what the sense tracks |
| any | null | — | the four false-positive figures on the record do not reproduce in forced-choice form, and P1 is uninterpretable until that is explained |
7. Registered statistics, bars and failure criteria
The unit is the seat, not the cell (E-20260813h's critic, imported): seven cells from one model
are not seven draws. Each seat contributes a count out of 7; a primary is stated on all three jointly.
Bars are step 1's, unchanged — a bar lowered after seeing a weaker manipulation is bar-shopping,
and the two steps must be comparable.
| id | statistic | fires when | one-sided exact P |
|---|---|---|---|
P1 |
FS chosen on AB, per seat, out of 7 |
all three seats ≥ 6 of 7 | 0.0625³ = 2.4 × 10⁻⁴ |
P1′ |
ODD chosen on AB, per seat |
all three seats ≥ 6 of 7 | as above, opposite direction |
P2 |
FS chosen on AC, per seat |
all three seats ≥ 6 of 7 | as above |
P3 |
ODD chosen on AD, per seat |
all three seats ≥ 6 of 7 | as above |
G1 |
FS chosen on QB |
≥ 19 of 21 cells | manipulation check on ODD |
G2 |
FS chosen on QC |
8–13 of 21 cells | parity gate |
G3 |
FF2 chosen on QD |
≥ 19 of 21 cells | manipulation check on ODD, completed |
G4 |
agreement between AB and AS, same seat and segment |
≥ 17 of 21 | source-visibility gate |
D1 |
Spearman ρ between a segment's oddity-site count and seats choosing ODD on AB; and ρ against device density |
descriptive, reported whatever it is | dose check |
P1 and P1′ cannot both fire. If neither fires the result is indeterminate, reported as
indeterminate with the width written down.
What voids the run.
F0—PARfails on 4 or more of 7 segments. Nothing is judged; the arm closesretired. (This is the gate that runs first and it is the reason step 2 exists.)F1—G2fails. If the seats tellFSfromFF2on quality at better than chance,FSis not the fluent-and-source-ward arm the design needs,P2is confounded with quality, andP1is reported but may not be read as separating carriage from non-fluency.F2—G3fails. IfODDis not read as the worse-written arm, the oddity operator did not operate andP1/P3are void.F3— aPARflag on 1–3 segments. Those segments are excluded and every statistic recomputed; both the full and the excluded figures are reported.F4— a banned string reaches a dispatched prompt. Voids the run.F5— position lock. Any seat answering the same slot letter in ≥ 20 of its 21 carriage cells is reported as position-locked and excluded fromP1,P2,P3, which are recomputed.F6—checks.pynon-zero at run time. Voids the primary. (It exits 0 at freeze.)F7— truncation. Note (bph):finish_reason != "length"is a kept-ness condition in the runner and an assertion in the verifier. A truncated body is re-dispatched, never counted.
Registered in advance about what a firing P1 does NOT license. It does not establish that the
FS arm carries anything — the census is a count, not a perception (RS-20260813g §10), and a blind
auditor already cut step 1's self-assessment from 0.976 to 0.80. It does not establish quality;
Tier D is NOT PASSED and every evaluative word here is internal-judgment-only and provisional.
And three models sharing a training distribution are not three readers: note (bpi) measured
two of these seats agreeing at a 33-character contiguous run on an unrelated task, so unanimity is
what a shared projection looks like and this run cannot tell that apart from perception.
8. Money
Pre-flight, built from the cap each request permits and not from an assumed length (note (abc)).
| stage | calls | cap each | worst case |
|---|---|---|---|
PAR + PLANT gate |
12 | 900 | $0.20 |
| pre-run critic | 1 | 14,000 total, reasoning 4,000 | $0.20 |
| judging | 147 | 900 | $0.85 |
| re-dispatch headroom | ~10 | 900 | $0.10 |
| $1.35 declared ceiling |
UTC-day headroom at freeze: $1.528904 (2026-08-15 stands at $3.471096 of $5.00 after six
sessions). The ceiling fits inside it with $0.18 to spare; if the gate stage overruns, the AS
condition is cut first (21 cells, ~$0.14) and the cut is reported.
Note (bny): every reasoning-capable seat gets an explicit reasoning: {"max_tokens": N} and
never an effort pin; the two keys cannot both be sent. Note (bph): the reasoning budget is held
at about a fifth of the content cap, never at or above it. Note (bnx): every body is appended to
runs/bodies.jsonl as it returns and a re-invocation resumes. Note (bof): no key-usage delta is
reported as a cross-check; per-request billed cost is the ledger.
Lead translation is free and is not ledgered (charter §3, A4): the rebuilt arm, the 507-check verifier, the alignment and the analysis cost $0.