Repository path: wiki/findings/results/RS-20260815e-fluent-carriage-2.md · rendered 2026-09-09
Page metadata (front matter)
| type | result |
|---|---|
| id | RS-20260815e-fluent-carriage-2 |
| status | frozen |
| created | 2026-08-15 |
| updated | 2026-08-15 |
| senses | perceived-source-carriage, naturalness |
| internal-judgment-only | true |
| provisional | true |
| links | workshop/experiments/E-20260815e-fluent-carriage-2/design.md, workshop/experiments/E-20260815e-fluent-carriage-2/amendments.md, wiki/arms/ARM-fluent-carriage.md, wiki/findings/results/RS-20260815b-fluent-carriage.md, wiki/findings/results/RS-20260814f-carriage-elevation-2.md, wiki/findings/results/RS-20260808d-carriage-decoupled.md, workshop/translations/atsui-suna/R29-v1/translation.md, workshop/translations/atsui-suna/R14-v2/translation.md, workshop/regimes/R29-fluent-carriage.md, wiki/goodness-senses.md, framework/v0.2/README.md |
RS-20260815e — a reader who cannot see the original calls the badly-written text the source-follower; a reader who can see it calls the source-following one
ARM-fluent-carriage step 2, on E-20260815e (arms frozen fffce5f7, design frozen 909b01ed,
both before any call). 147 forced-choice cells, 0 unparseable, 0 errors, 0 truncated.
Pre-run critic NEEDS-REDESIGN, 13 findings, 4 BLOCKING, all 13 accepted. Post-run verifier
35 checks, 0 failures, 8 of 8 mutations caught. $1.215929050 against a declared $1.35.
1. The one-sentence answer
Blind, the sense splits: two seats of three call the merely-awkward arm the source-follower and one calls the carrying arm, and the decisive cell is indeterminate for the second time in two builds. Show the same seats the original and all three choose the carrying arm — 20 of 21 cells, up from 11.
The rest of the page is what makes that readable: a manipulation that passes a content-parity control 7 of 7 where step 1's passed 0 of 7, a device census cut from 33 to 25 by a blind auditor rather than by the translator, and a quality gate that fails in the direction that saves the co-primary.
2. What was measured, and on what
One story — 牧野信一 «熱い砂の上» (1935), Aozora Bunko, whole — three arms from one hand:
| arm | what it is | words |
|---|---|---|
FS |
T-atsui-suna-R29-v1, unchanged from step 1: the source's declared devices carried by English resources |
1,485 |
FF2 |
T-atsui-suna-R14-v2: the same text with 25 device sites flattened, every other string byte-identical |
1,515 |
ODD |
FF2 plus 35 propositional-neutral markedness edits, all at paragraphs where FS and FF2 agree |
1,571 |
The 25 sites are not the 84 step 1 used, and the difference is the point of the rebuild. Three classes were dropped before the run because step 1's parity auditor had shown they are not form-only — suspension, reduplication, held keyword. Eight more sites were dropped during the run because a blind auditor refused them (§4). What is left is 24 sites of pure constituent order plus one mimetic, and the "purely syntactic" claim the design made in advance is therefore true by measurement rather than by assertion — which is precisely what the pre-run critic's F-3 said it was not.
3. The parity gate, which is the run's procedural result
Note (boa) was applied for the first time and it worked. The control that cost step 1 its whole carriage half was dispatched first, before any judging call, and the judging stage was written to refuse to start without a recorded pass.
| build | plants caught | segments SAME | what it bought |
|---|---|---|---|
| step 1 (7 classes, control run last) | 3 of 3 | 0 of 7 | nothing — the run's carriage half was withheld |
| build 1 (4 classes) | 5 of 5 | 5 of 7 | two flags, both my errors, repaired for $0.086 |
| build 3 (4 classes, repaired, 25 sites) | 5 of 5 | 7 of 7 | a manipulation whose figures can be read |
Dropping the three non-form-only classes removed five sevenths of the parity problem at a stroke. The two build-1 flags were mine, not the classes': folding a standalone attribution made me drop an addressee (to B) at one site and supply a speech verb (cried) the source has no word for at another; breaking a piled period turned a sequential and then again into a simultaneous at the same time. All three are content changes made while removing a form, all three were invisible to 507 mechanical checks, and that is the same failure as step 1's §6(iii), arriving through a different device class.
Three things are owed against this pass and are stated rather than buried.
- Build 3's parity is co-authored by the auditor, in the pre-run critic's own words (F-5). The same seat that certifies these arms was used to correct them at three sites. It is not an independent pass and is not reported as one.
- The gate is not deterministic. Segment S4 was flagged DIFFERENT at build 2 and SAME at
build 3 on a phrasing that did not change between them —
A2.6's "off at a scatter" against "in ones and twos". One call is one draw. Every primary below is therefore also reported with S4 excluded, and the exclusion changes no verdict (§5). - The
ODDarm shifts character register, and the gate said so in all seven segments once it was asked (critic F-12): "R's dialogue is more formal/stilted … shifting the speakers' voices away from the colloquial tone of P and Q."ODDvaries fluency and register together and this run cannot separate them. Standing caveat onP1andP3.
4. The blind site audit, which is the craft finding
The design's rule C-8d — a site with no carriage is never flattened — was wired to the
translator's verdict only. The critic's F-2 said so, and the fix was to run the blind carriage
audit on all 33 candidate sites and treat the auditor's NO exactly as the translator's FAILED.
25 carried, 8 refused.
| class | flattened | refused | the auditor's words |
|---|---|---|---|
| A6 attribution standing alone | 12 of 12 | — | — |
| A5 the piled period | 9 of 10 | A5.05 | "English front-loads noun, loses end-suspension" |
| A7 the subject withheld | 3 of 5 | A7.2, A7.3 | "English inverts, action first instead of verb-final" |
| A2 reduplicated mimetic | 1 of 6 | A2.1–A2.5 | "onomatopoeia dropped, plain verb only" · "mimetic reduplication flattened into plain verb" |
The mimetic class collapses, and it is the class R29's log was proudest of. That log wrote:
"What replaces it is not a manner adverb but an English verb that is itself iterative and
sound-symbolic … English has a large stock of these and they cost nothing. Six sites, six carriages,
no compensation needed." A blind reader, given the Japanese and the whole rendering, agrees at
one site of six. Step 1's auditor had already refused A2.4 — slapping for ピシヤ/\ — on a
twenty-site sample; this replicates that refusal on four more.
The finding, stated as a craft claim rather than as an instrument note: a single English
sound-symbolic verb does not do the work of a reduplicated mimetic, on the judgment of a reader who
can see both. (A statement about one reader on one story: internal-judgment-only.) A6 — print
the line, then the attribution, as its own sentence — is the only class that transfers entire.
5. The registered statistics
The unit is the seat. Exact tails by enumeration of all 2ⁿ outcomes. Bars are step 1's, unchanged.
| id | cell | statistic | per seat (of 7) | pooled | bar | verdict |
|---|---|---|---|---|---|---|
P2 |
FS vs FF2, carriage |
FS chosen |
6 · 7 · 7 | 20/21 | all ≥ 6 | FIRES, P = 0.0625 per seat |
P3 |
ODD vs FF2, carriage |
ODD chosen |
7 · 6 · 7 | 20/21 | all ≥ 6 | FIRES |
P1 |
FS vs ODD, carriage |
FS chosen |
6 · 2 · 3 | 11/21 | all ≥ 6 | INDETERMINATE |
P1′ |
FS vs ODD, carriage |
ODD chosen |
1 · 5 · 4 | 10/21 | all ≥ 6 | does not fire |
G1 |
FS vs ODD, quality |
FS chosen |
7 · 7 · 7 | 21/21 | ≥ 19 | passes |
G3 |
ODD vs FF2, quality |
FF2 chosen |
7 · 7 · 7 | 21/21 | ≥ 19 | passes |
G2 |
FS vs FF2, quality |
FS chosen |
3 · 1 · 2 | 6/21 | reject if ≤ 6 or ≥ 15 | REJECTS, two-sided P = 0.0784 |
G4 |
AB vs AS agreement |
same answer | — | 10/21 | ≥ 17 | FAILS |
With S4 excluded (§3.2) every verdict is unchanged: P2 5·6·6 fires, P3 6·6·6 fires, P1
5·1·2 and P1′ 1·5·4 both indeterminate. The sensitivity is reported because it was registered, and
it says nothing new.
No seat is position-locked (F5): A/B splits 16/12, 16/12, 15/13 over 28 carriage cells each,
against a lock threshold of 20.
6. G2 fails, and the direction of the failure is what makes P2 readable
The seats prefer the flattened arm as English, 15 of 21, and call the carrying arm the source-follower, 20 of 21.
By the registered criterion this is a failure: F1 fires, the arms are not of equal quality, and no
sentence on this page says "parity". The critic's F-4 was right that the 8–13 interval was an
acceptance region and never a parity test; the phrase has been struck everywhere.
But F1 was written against the case where quality and carriage point at the same arm, and here
they point at opposite arms. A carriage preference cannot be a quality preference when the seats,
asked which is better English, choose the other text. H-QUALITY is refuted on AC, not
confirmed — refuted more sharply than step 1 could refute it, because step 1's lean (8 of 21, P =
0.383) was indistinguishable from chance and this one is not.
And that costs the craft claim step 1 deposited, which is withdrawn here. RS-20260815b §6b let
the handbook say "carrying a source's form does not require paying in fluency" on the strength of
G2 passing at 8 of 21. On a cleaner manipulation the same gate runs the other way: carrying the
source's constituent order does cost something a reader notices, at 15 of 21, two-sided P = 0.0784.
That is a lean, not a demonstration — it does not reach conventional significance and is stated as
what it is — but it is the wrong direction for the claim, and the claim comes out of the framework.
7. The source-visible cell, which is the finding
AS gives the seats the same pair as AB — FS against ODD — with the Japanese printed above it.
G4 asked whether the answer changes. It changes, and it changes systematically.
| seat | AB (blind) |
AS (original shown) |
|---|---|---|
P1 |
FS 6 of 7 |
FS 6 of 7 |
P2 |
FS 2 of 7 |
FS 7 of 7 |
P3 |
FS 3 of 7 |
FS 7 of 7 |
| pooled | 11 of 21 | 20 of 21 |
Both seats that preferred the merely-awkward arm blind reverse completely when they can see the
original. Agreement between the two conditions is 10 of 21 against a registered bar of 17, so G4
fails; the failure is not noise, it is a one-way reversal.
What this licenses, and it is less than the table invites. AS was registered as an agreement
gate, not as a directional test, and no bar was set for the 20-of-21 figure. It is descriptive
and is labelled so. The registered fact is that the gate failed and that both reversals ran the same
way.
What it means, stated at that strength. The two conditions are not two measurements of one thing.
Which of these follows the original, asked of a reader who cannot see the original, is answered
substantially by how translated the English sounds — and P3 shows what that costs: an arm that
carries nothing and is merely badly written is called the source-follower in 20 of 21 cells, by the
same seats, with content parity established. Asked of a reader who can see the original, the same
question is answered by what the original does — and the answer converges on the carrying arm at
exactly the rate P2 reaches when the awkward arm is taken out of the comparison.
The seats' own reasons, coded blind by a committed keyword script, say the same thing from the inside:
| cell | carriage family | literality family | quality family |
|---|---|---|---|
AC — only carriage varies |
18 of 21 | 1 | 0 |
AS — carriage varies, original shown |
18 of 21 | 2 | 1 |
AB — both vary, blind |
15 | 9 | 0 |
AD — only awkwardness varies |
12 | 19 | 3 |
8. P1 is indeterminate for the second time in two builds
Step 1: 13 of 21, seats 4 · 6 · 3, on 84 device sites. Step 2: 11 of 21, seats 6 · 2 · 3, on 25. Two constructions, two device sets, two oddity operator sets, one parity-clean and one not — and the decisive cell does not resolve either time. The seats do not merely fail to reach a bar; they disagree with each other, and in step 2 they disagree more than in step 1.
That is the arm's answer to its own question, and it is a null with content: when carriage and non-fluency are set against each other and the reader cannot see the source, there is no stable answer to follow. The instability is the result, not a failure to find one.
P1 also carries the caveat the pre-run audit forced (F-1), and it is a real one. An independent
seat, shown the 35 oddity edits blind, attributed 25 of them to Japanese — "nominalization"
eleven times, "cleft mirrors Japanese emphasis structure", "'in the event that' translates
Japanese conditional". So in AB both arms read as translated from the source language, and a seat
choosing FS may be preferring competent Japanese-sounding English to incompetent
Japanese-sounding English rather than reading direction at all. Retiring the operator the critic
named did not move this: the replacement operators drew the same verdict.
One operator escaped, and it is worth keeping. By operator: nominalisation 11 of 12 attributed, circumlocution 7 of 8, cleft 3 of 3 — but stilted collocation only 4 of 12. If a later design needs unlicensed markedness that does not leak the source language, it should be lexical stiffness, not syntactic heaviness.
9. What is NOT established
AS's 20 of 21 is descriptive. No bar was registered for it. What is registered is thatG4failed at 10 of 21 and that both reversals ran towardFS.- The
ODDarm confounds fluency with character register, measured in 7 of 7 segments (§3.3).P1andP3cannot separate "worse English" from "wrong voice for this speaker". - Both arms leak the source language (§8), and all three judging seats named Japanese unprompted when probed. The recall confound is closed — none of the three had encountered the story — but the language leak is total, as it was in step 1.
- Three models sharing a training distribution are not three readers. Note (bpi) measured two
of these seats agreeing at a 33-character contiguous run on an unrelated task.
P2's andP3's 20 of 21 are what a shared projection looks like. What makes them more than that is that the same seats, on the same texts, answer the quality question in the opposite direction and the source-visible question differently from the blind one. - One story, one author, one language pair, one hand. The translator wrote all three arms, chose the classes, chose the segments and knew the hypotheses throughout. The F-1 and F-2 audits close two of those doors — the census is now the auditor's, not mine, and the oddity arm's leak is measured rather than assumed. The rest are open and are listed in the critic's F-9.
G2's lean is a lean. 15 of 21 at two-sided P = 0.0784 rejects the registered interval and does not reach conventional significance. It is enough to withdraw a claim built on the opposite lean; it is not enough to assert the converse.- Nothing about quality. Tier D is NOT PASSED. Every evaluative word here is
internal-judgment-onlyandprovisional.
10. What this deposits
On wiki/goodness-senses.md, perceived-source-carriage — the arm's completion criterion, and it
is met in the "in either direction" form:
Measured source-blind, this sense is substantially a non-fluency detector: an arm carrying nothing and merely badly written is chosen as the source-follower over a clean flattening in 20 of 21 cells (
RS-20260815eP3), with content parity established at 7 of 7. Measured with the source available, the same seats choose the carrying arm in 20 of 21 cells. Do not read a source-blind carriage score as evidence about carriage. And on method: *a carriage manipulation must pass an independent content-parity control before* its figures are read; of seven declared device classes on this story, five have been shown by an independent reader to carry something propositional or to fail as carriage.
On framework/v0.2 §8 Q-e: the reader-side leg is not restored, and the craft leg deposited
by step 1 is withdrawn (§6). What a handbook may say on this evidence: carrying the source's
constituent order is detectable by a reader who has the source, and is measurably not free — readers
lean toward the flattened text as English at 15 of 21.
On R29: rule C1's census is a translator's self-assessment and this run priced it —
25 of 33 on a blind reader's judgment, with the mimetic class at 1 of 6. Any future R29
rendering reports its census as claimed, and the regime's A2 guidance is downgraded from "they
cost nothing" to contested.
11. Controls, verification and money
- Dispatch health: 147 cells, 147 bodies, 0 unparseable, 0 errors, 0 truncated, 0 re-dispatched.
Two earlier holes were re-dispatched at the gate and audit stages, both note (bny) —
finish_reason: lengthwith zero visible characters atreasoning.max_tokens180 inside a 900 content cap — and both repaired at a larger total with the sub-budget still explicit. - The gate precedes the judging on disk, and the verifier asserts it (
V-22), as it asserts that every audit body precedes every judging body (V-23) — the audits changed the arms, so their ordering is a condition of the result, not a courtesy. - The verifier imports nothing from
analysis/: the parser, the tallies and the exact tails are written again, and the tails are derived by enumerating all 2ⁿ outcome strings rather than bymath.comb. 35 checks, 0 failures, 8 of 8 mutations caught. - Pre-run: 512 checks, 0 failures, 9 of 9 mutations caught.
Money. Every figure API-reported with "usage": {"include": true}. $1.215929050 — the parity
gate $0.236037 across three builds, the critic $0.180201, the three pre-run audits $0.135946, the
147 judging cells $0.663745 — against a declared ceiling of $1.35. No key-usage delta is reported
and none was used as a cross-check (note (bof)).
P3 x-ai/grok-4.5 billed $0.473 of the judging total against P1's $0.085 and P2's $0.105,
a routing spread the pre-flight could not see; a hard spend guard, reading spend off disk before each
call, was installed mid-run and is what made the last two conditions affordable.
Lead translation is free and is not ledgered (charter §3, A4): the rebuilt arm, the 512-check verifier, the alignment and the analysis cost $0.