Repository path: workshop/experiments/E-20260815b-fluent-carriage/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260815b-fluent-carriage |
| status | frozen |
| created | 2026-08-15 |
| updated | 2026-08-15 |
| senses | perceived-source-carriage, naturalness |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-fluent-carriage.md, wiki/findings/results/RS-20260814f-carriage-elevation-2.md, wiki/findings/results/RS-20260813h-carriage-or-elevation.md, wiki/findings/results/RS-20260808d-carriage-decoupled.md, workshop/translations/atsui-suna/R29-v1/translation.md, workshop/translations/atsui-suna/R14-v1/translation.md, workshop/translations/atsui-suna/device-sites.md, workshop/translations/atsui-suna/provenance.md, workshop/regimes/R29-fluent-carriage.md, wiki/goodness-senses.md, config/models.md |
E-20260815b — carriage against non-fluency, with the two predicting opposite arms
AMENDED 2026-08-15 on the pre-run critic pass, before any judging call was dispatched.
P4moonshotai/kimi-k3returnedNEEDS-REDESIGN, ten findings, 2 BLOCKING, and all ten were accepted.amendments.mdis the authoritative record of what this design now says; where it and the text below differ, the amendment governs. In particular: the ODD arm was rebuilt (F-1), the licensed conclusion became a conjunction ofP1andP2(F-2),G2's acceptance region was corrected (F-3), and three new conditions —LG,AUand the parity control — were added.
ARM-fluent-carriage step 1. Frozen before any call is dispatched. The three arms and both
translator's records were frozen at 2af0d84, before this file existed.
1. The question
ARM-carriage-direction established that perceived-source-carriage is oriented — it survives
reversing which arm is the more elevated one, across two language families — and then, at S184,
installed the rival that explains every one of its results equally well:
H-LITERALITY — the seats choose whichever text reads more like a translation, without reading for direction at all.
RS-20260814f §7: "Separating H-LITERALITY needs a fourth arm that is fluent and source-ward …
and that arm does not exist in this project." This run builds it.
The question: when carriage and non-fluency point at opposite renderings, which one does the judgment follow?
2. Why the arm can exist now — the three-way construction
One hand, one story, three texts, and two orthogonal manipulations:
| arm | what it is | carriage | fluency |
|---|---|---|---|
FS |
T-atsui-suna-R29-v1, the whole story translated from the Japanese under a rule set that carries the source's declared devices using English resources and forbids importing the exponent |
high | high |
FF |
T-atsui-suna-R14-v1, the same text with exactly the 84 declared device sites flattened and every other string byte-identical |
zero | high |
ODD |
FF with 31 propositional-neutral markedness edits — nominalisation, agentless passive, stilted collocation, adverb interpolation, redundant relative — every one of them at a paragraph where FS and FF are byte-identical |
zero | low |
The orthogonality is not a claim, it is a mechanical property, asserted by analysis/checks.py:
FS and FF differ only at device spans; FF and ODD differ only at edit sites; the edit
sites are disjoint from the device spans. 218 checks, 0 failures, 6 of 6 mutation tests caught,
run before this file was written.
Materials. 牧野信一 «熱い砂の上» (1935), Aozora Bunko, the whole story — 3,198 characters, 106
paragraphs. Japanese→English is the pair where carrying the source is most reliably translationese,
so it is the hardest case for the FS arm and therefore the informative one. Copy-text, the
single-witness declaration and the contamination gate: ../../translations/atsui-suna/provenance.md.
No published English rendering of this story is reachable, so the overlap statistic has no second
operand and was not run; that is stated on the artifacts rather than glossed.
The FS arm's own result, independent of any call. 84 sites, 82 CARRIED, 2 FAILED, and both
failures are grammatical — English cannot leave a finite main clause subjectless (A7.1) and has no
dash-then-stop (A1.21) — not stylistic. On this story, on this device set, the carriage/fluency
trade did not bite. Whether the result reads as unmarked English is exactly what G1 measures and
is not the translator's to say.
3. Segments
Seven, cut at the story's own joints (the four section breaks and three scene turns), chosen and
committed in analysis/segments.py before any judging. analysis/segments.json records the
paragraph range in each arm; the alignment is computed by difflib and asserted — every boundary
falls on a paragraph the arms agree on.
| seg | ¶ (FS) | FS w |
FF w |
ODD w |
device sites | oddity sites |
|---|---|---|---|---|---|---|
| S1 | 1–15 | 264 | 281 | 281 | 12 | 1 |
| S2 | 16–27 | 137 | 138 | 138 | 8 | 3 |
| S3 | 28–49 | 189 | 199 | 207 | 11 | 8 |
| S4 | 50–67 | 217 | 214 | 218 | 10 | 4 |
| S5 | 68–75 | 178 | 188 | 193 | 4 | 3 |
| S6 | 76–94 | 282 | 280 | 289 | 10 | 7 |
| S7 | 95–105 | 217 | 216 | 216 | 5 | 5 |
| all | 1,484 | 1,516 | 1,542 | 84 | 31 |
Whole-text ratios FF/FS 1.0216, ODD/FS 1.0391, ODD/FF 1.0171; declared
tolerance ±6% whole-text and ±10% per segment, asserted by checks.py (C-7). The worst segment is
S3 at 1.095.
S1 has one oddity site and twelve device sites. That is a property of the story — S1 is
device-dense and has almost no paragraph outside a device span — and it is declared here rather than
discovered afterwards. It licenses a registered dose check (D1, §6).
4. Conditions
| condition | pair | source shown | question | role |
|---|---|---|---|---|
AB |
FS vs ODD |
no | carriage | PRIMARY P1 — the decisive cell |
AC |
FS vs FF |
no | carriage | CO-PRIMARY P2 — does the sense respond to carriage with fluency held? |
AD |
ODD vs FF |
no | carriage | P3 — the false-positive, in forced-choice form |
QB |
FS vs ODD |
no | quality | GATE G1 |
QC |
FS vs FF |
no | quality | GATE G2 — the fluency-parity gate P2 depends on |
QD |
ODD vs FF |
no | quality | GATE G3 — the manipulation check on ODD |
AS |
FS vs ODD |
yes | carriage | GATE G4 — does showing the Japanese change the answer? |
7 conditions × 7 segments × 3 seats = 147 cells, one call each.
Seats: P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5 — the
same three as both steps of ARM-carriage-direction, so the instrument is imported unchanged.
P5 excluded on note (bne). P4 moonshotai/kimi-k3 is the pre-run critic and takes no cell.
Presentation order fixed per cell by sha256("E-20260815b|" + condition + "|" + segment + "|" +
seat) mod 2.
The carriage question is E-20260814f's as dispatched, verbatim, with one word changed — short
story stays, nineteenth-century and another European language would both be false here:
Two English translations of the same passage from a twentieth-century short story are printed below. They translate the same original, which is not in English. Neither is the original.
One of these two translators set out to follow the original's own way of putting things as closely as English allows. The other set out to do something else.
Which of the two followed the original's own way of putting things?
Answer on ONE line, in exactly this form, and print nothing else:
AorB, then a semicolon, then at most fifteen words naming the one feature you used.
The quality question, built parallel to it in form:
Two English translations of the same passage from a twentieth-century short story are printed below. They translate the same original, which is not in English. Neither is the original.
One of these two is written in better, more natural English than the other.
Which of the two is written in the better, more natural English?
Answer on ONE line, in exactly this form, and print nothing else:
AorB, then a semicolon, then at most fifteen words naming the one feature you used.
In AS the Japanese is prepended under the heading ORIGINAL. The word Japanese appears in no
dispatched prompt in any condition, and neither does the author's name, the title, any regime id,
any rule name, or any of the strings carriage, flatten, device, fluent, odd — asserted by
the verifier, and a violation voids the run.
5. The hypotheses, and what each predicts
- H-CARRIAGE — the seats track which way the translation is marked, i.e. whether its form
follows the source's. Predicts
FSonAB,FSonAC, and makes no prediction onAD, where neither arm carries anything. - H-LITERALITY — the seats choose whichever text reads more like a translation, without reading
for direction. Predicts
ODDonABandODDonAD. OnACit predictsFFif flattening makes the English read more translated, and is otherwise indifferent;G2decides which, and the prediction is therefore conditional and stated as such rather than free. - H-LENGTH — the seats choose the longer text. Predicts
ODDonABandADandFFonAC. It is registered becauseFFis the longer arm in 4 of 7 segments andODDin 5 of 7: flattening and oddity both add words, so this rival travels with both manipulations. - H-QUALITY — the seats read "followed the original" as a general merit judgment and choose the
better-written arm. Predicts
FSonABand is indifferent onACiffG2shows the two arms are of equal quality.
AB is the primary because H-CARRIAGE and H-LITERALITY predict opposite arms there, which is
what no design in the predecessor arm could construct. AC is the co-primary because it is the
only cell in which H-CARRIAGE and H-LENGTH predict opposite arms, and because it is the cell
H-QUALITY cannot explain if G2 passes.
Read together, the three carriage cells are diagnostic and the pattern is stated before the numbers exist:
P2 (AC) |
P3 (AD) |
P1 (AB) |
what it means |
|---|---|---|---|
FS |
ODD |
FS |
the sense responds to carriage, responds to bare non-fluency too, and carriage wins when they conflict — H-LITERALITY survives as a contributor and is refuted as the account |
FS |
ODD |
ODD |
non-fluency beats carriage. The entry is rewritten as a non-fluency detector and framework/v0.2 §8 Q-e loses its reader-side leg |
| null | ODD |
ODD |
the sense does not respond to carriage at all on this material; H-LITERALITY takes the entry outright |
FS |
null | — | the four false-positive figures on the record do not reproduce in forced-choice form, and P1 is uninterpretable until that is explained |
6. Registered statistics, bars and failure criteria
The unit is the seat, not the cell — E-20260813h's critic BLOCKING 2, imported: seven cells
from one model are not seven draws. Each seat contributes a count out of 7 segments; the primary is
stated on all three seats jointly.
| id | statistic | fires when | one-sided exact P |
|---|---|---|---|
P1 |
FS chosen on AB, per seat, out of 7 |
all three seats ≥ 6 of 7 | 0.0625³ = 2.4 × 10⁻⁴ |
P1′ |
ODD chosen on AB, per seat, out of 7 |
all three seats ≥ 6 of 7 | as above, opposite direction |
P2 |
FS chosen on AC, per seat |
all three seats ≥ 6 of 7 | as above |
P3 |
ODD chosen on AD, per seat |
all three seats ≥ 6 of 7 | as above |
G1 |
FS chosen on QB |
≥ 19 of 21 cells | manipulation check |
G2 |
|FS − FF| on QC |
8–13 of 21 cells for FS (i.e. no seat-level majority beyond chance) |
parity gate |
G3 |
FF chosen on QD |
≥ 19 of 21 cells | manipulation check on ODD |
G4 |
agreement between AB and AS, same seat and segment |
≥ 17 of 21 | source-visibility gate |
D1 |
Spearman ρ between a segment's oddity-site count and the number of seats choosing ODD on AB |
descriptive; reported whatever it is | the dose check S1 licenses |
P1 and P1′ cannot both fire. If neither fires the result is indeterminate and is reported
as indeterminate, with the width of the indeterminacy written down; the arm's step 2 then becomes a
rebuild rather than a replication (ARM-fluent-carriage §Steps, registered before this run).
What voids the run.
F1—G2fails. If the seats can tellFSfromFFon quality at better than chance, theFSarm is not in fact the fluent-and-source-ward arm the design needs,P2is confounded with quality, andP1is reported but may not be read as separating carriage from non-fluency. This is the single most likely way this run fails, and it is named first deliberately.F2—G3fails. IfODDis not read as the worse-written arm, the oddity operator did not operate andP1/P3are void.F3— the content-parity control fails. An independent seat (P4, which takes no judging cell) is asked per segment whether the three arms state the same things, and is given three planted factual errors as a positive control. If it does not catch 3 of 3, its verdicts on parity carry no weight and every primary is reported with parity unestablished. If it reports a real content difference, the affected segment is excluded and every statistic recomputed.F4— a banned string reaches a dispatched prompt. Voids the run.F5— position lock. Any seat answering the same slot letter in ≥ 20 of its 21 carriage cells is reported as position-locked and excluded fromP1,P2,P3, which are recomputed on the remaining seats.F6—checks.pynon-zero at run time. Voids the primary. (It exits 0 at freeze.)
Registered in advance about what a firing P1 does NOT license. It does not establish that the
FS arm carries anything — the census is a count, not a perception (RS-20260813g §10). It does not
establish quality; Tier D is NOT PASSED and every evaluative word here is
internal-judgment-only and provisional. And three models sharing a training distribution are
not three readers: unanimity is what a shared projection looks like and this run cannot tell that
apart from perception.
7. Money
Pre-flight, built from the cap each request permits and not from an assumed length (note (abc)). 147 judging calls at ≈ 700 prompt tokens and ≤ 120 completion tokens, plus 7 parity calls, plus one critic call. Declared ceiling $1.50, against a UTC-day headroom of $4.678 (2026-08-15 stands at $0.321811 of $5.00 after S188).
Every reasoning-capable seat is sent an explicit reasoning: {"max_tokens": N} and never an
effort pin — note (bny), which fired yesterday for $0.100636 and zero characters. The two keys
cannot both be sent (HTTP 400). The runner appends each body to bodies.jsonl as it returns and a
re-invocation resumes — note (bnx).
Lead translation is free and is not ledgered (charter §3, A4): all three arms, the site table, the alignment and the 218-check verifier cost $0.