Repository path: wiki/findings/results/RS-20260811g-carrier.md · rendered 2026-09-09
Page metadata (front matter)
| type | result |
|---|---|
| id | RS-20260811g-carrier |
| status | active |
| created | 2026-08-11 |
| updated | 2026-08-11 |
| senses | — |
| internal-judgment-only | true |
| provisional | true |
| links | workshop/experiments/E-20260811g-carrier/design.md, workshop/translations/kasagi/R06-v1/translation.md, workshop/translations/kasagi/collation.md, wiki/arms/ARM-carrier.md, wiki/findings/results/RS-20260811f-dakghar-address.md, wiki/findings/results/RS-20260808g-two-pasts.md, framework/v0.2/README.md, config/models.md, config/budget.md |
RS-20260811g-carrier — the primary is withheld by the gate this design promoted, and the gate's failure is the finding
Design E-20260811g-carrier, frozen at c34165a; the translation limb it hangs on
(T-kasagi-R06-v1, «Kaşağı» whole) was frozen at d58caa5 before the design was written, and
its contamination was measured after that commit and before any locus was chosen. Pre-run critic
NEEDS-AMENDMENT, 8 findings, 1 BLOCKING; five accepted, three accepted in part with the
overrules written (design §9). Verifier 665 checks, 0 failures, 6 mutation tests, 6 caught.
Cost $0.290330128 of a declared ceiling of $1.46.
1. The question and the one-sentence answer
RS-20260811f found a Bengali honorific verb ending reaching English at zero while its own
control — the same social relation carried by an address noun — reached English at 0.8333.
This design asked whether that is general: is what predicts arrival the carrier (free word
against bound morpheme) rather than the content? — and crossed carrier with content inside one
Turkish story, so that language, work, passage, hands, arbiters, question and statistic are all
constants.
The answer is that this run cannot say, and the reason is worth more than the numbers it
withheld: three arbiters reading the Turkish itself agreed the two variants differed at only
0.4444 for the bound edits and 0.5556 for the free ones, against a registered floor of 0.667.
G1 — the source-parity gate this design promoted from checksum to load-bearing precisely because
everything turned on it — fails, and every primary is withheld.
The withheld numbers point exactly where the conjecture predicted, which is the situation in which a registered gate is worth having and the one in which it is most tempting to argue around. They are reported in §4, unlicensed, and nothing is released.
2. What passed, and it is not nothing
| gate | bar | value | |
|---|---|---|---|
G2 — a bound morpheme English must carry (Bilmiyorum → Biliyorum) |
≥ 0.75 | 1.0000 | PASS, 6 of 6 cells |
G3 — the floor, one hand against itself on identical input |
≤ 0.40 | 0.1304 | PASS, and below the critic's 0.25 review threshold |
G4 — the added-word control the critic bought (— Gel buraya, hadi!; Hemen kasabaya at gönderildi.) |
≤ 0.50 | 0.0000 | PASS, 11 of 11 cells SAME |
G1 — source parity, level clause |
both ≥ 0.667 | GRAM 0.4444, LEX 0.5556 | FAIL |
G1 — source parity, gap clause |
|GRAM − LEX| ≤ 0.25 | 0.1111 | pass |
G2 and G4 together kill both cheap explanations of the pattern in §4 before it is read.
A bound morpheme is not invisible to these hands and arbiters — negation crossed at ceiling, 6 of 6.
And adding a free word does not by itself make an arbiter say the two Englishes differ — the
semantically inert additions crossed at exactly zero, 11 of 11. The critic's SERIOUS 3, which
bought G4, bought the run's most useful number.
3. Why G1 failed, and why dropping the obvious culprit does not rescue it
The three arbiters do not agree about the Turkish.
| seat | slug | SOURCE all |
SOURCE GRAM |
SOURCE LEX |
CROSS |
FLOOR |
|---|---|---|---|---|---|---|
J1 |
openai/gpt-5.6-terra |
0.500 | 0.333 | 0.667 | 0.125 | 0.125 |
J2 |
google/gemini-3.6-flash |
0.917 | 0.833 | 1.000 | 0.458 | 0.125 |
J3 |
deepseek/deepseek-v4-pro |
0.083 | 0.167 | 0.000 | 0.087 | 0.143 |
J3 returned SAME at 11 of 12 source items, including every one of the address-term edits
that J1 and J2 both describe correctly and in detail. It is the seat that is not reading the
Turkish.
And removing it does not save the gate. With J3 dropped — a post-hoc analysis, unlicensed,
computed after the gate failed, registered in no design, and used to release nothing — the source
parity comes back GRAM 0.5833, LEX 0.8333. The bound arm is still below 0.667. G1 fails on
two seats as it fails on three, so the withholding is not one seat's artifact and the primary
stays withheld.
What G1's failure actually says, per window (DIFFERENT verdicts out of three seats, reading
the Turkish):
| window | content | GRAM |
LEX |
|---|---|---|---|
W1 father → his groom, Gel → Gelin / + ağam |
social | 1/3 | 2/3 |
W2 child → the servant, ağlıyorsun → -sunuz / + abla |
social | 2/3 | 2/3 |
W3 groom → the master's son, yap → yapın / + beyim |
social | 0/3 | 2/3 |
W4 haberi yoktu → yokmuş / + anlaşılan |
epistemic | 1/3 | 1/3 |
W5 gönderildi → gönderilmiş / + anlaşılan |
epistemic | 1/3 | 2/3 |
W6 vardı → varmış / + anlaşılan |
epistemic | 3/3 | 1/3 |
Two things in that table are worth more than the pooled means.
- The social 2sg → 2pl edits are faint even in the Turkish. At
W3no seat of three saw anything when an old groom switched fromyaptoyapınaddressing his master's small son. The critic's BLOCKING 1 predicted the opposite problem — that these edits would be loud pragmatic violations — and on this evidence they are not loud, they are barely visible. - The one bound edit every seat saw is the mirative
varmış(W6, 3/3), which is also the one bound edit whose free-word counterpart they mostly missed (anlaşılan, 1/3). The epistemic column is not ordered the way the social column is, and the critic's SERIOUS 5 — that-mIşis polysemous — cuts both ways rather than one.
4. The withheld numbers, unlicensed, reported because a withheld primary that is never printed cannot be checked
None of this is licensed by any gate. It is printed so a successor can see what a working instrument would have been measuring, and it supports no claim.
GRAM (bound) |
LEX (free word) |
difference | |
|---|---|---|---|
CROSS, all six windows |
0.0556 | 0.4111 | +0.3556 — below its own registered bar of 0.40, bootstrap 90% interval [0.189, 0.544], which straddles the bar |
SOCIAL (W1–W3) |
0.0556 | 0.4444 | +0.3889 |
EPISTEMIC (W4–W6) |
0.0556 | 0.3778 | +0.3222 |
FLOOR (one hand against itself) |
— | — | 0.1417 over the four floor windows |
So even had G1 passed, P1 would have failed its own threshold (0.3556 against 0.40) with an
interval the data cannot resolve — the outcome amendment A5 of RS-20260811f's critic was written
to make visible. The design did not get to a positive result and then lose it to a gate; it got to
an underpowered one.
P3's shape is the one thing here that replicates rather than merely repeats. The bound arm
crossed at 0.0556, below the floor of 0.1417 — a bound mark reaches English at less than
the rate one hand differs from itself given byte-identical input. That is the same shape as
RS-20260811f's Bengali null (CROSS 0.0000, FLOOR 0.0000) in a second language family, and it
is withheld on the same gate as everything else.
And the two GRAM cells that did come back DIFFERENT are false positives, which only the
what_differs text shows. Of 24 CROSS-GRAM cells, 2 were DIFFERENT, and reading them:
W2/T-a/J1: "Version 2 says Hasan's ghost appears, implying he is dead, whereas Version 1 says only Hasan's image appears."W6/T-a/J2: "Version 1 refers to grooming multiple horses, whereas Version 2 refers to grooming a single horse."
Neither is the manipulation; both are ordinary paraphrase drift elsewhere in the passage. Read
by reasons rather than by counts, the bound arm arrived at 0 of 24. This is why P4 was
registered to require a named carrier, and it is a general lesson about this instrument: a binary
same/different verdict without its reason will silently count drift as signal.
5. What the translated English actually did — description, no gate
These are the hands' own words, and they are the most informative thing in the run.
The bound edits produce byte-identical English. — Gel buraya! and — Gelin buraya!:
A (as printed, 2sg) |
GRAM (2pl) |
LEX (+ ağam) |
|
|---|---|---|---|
T-a kimi-k3 |
— Come here! | — Come here! | — Come here, ağam! |
T-b grok-4.5 |
— Come here! | — Come here! | — Come here, my agha! |
And the free words arrive by being carried over, not by being translated into an English device.
That was not the conjecture's expectation and it changes how the mechanism should be described.
ağam became a transliteration in one hand and a calque in the other. At W3, beyim
became "my little master" and "young master" — English has no honorific system to put it in, so
both hands used the one thing English does have in that position: a noun in the vocative slot.
At W5 and W6, anlaşılan became "Apparently", in the sentence-adverb slot.
What a free word buys is not a matching category in the target. It is a position. English has nowhere to put a second-person plural agreement suffix and no honorific paradigm to receive
ağam— but it has a place at the end of a call where a noun of address goes, and a place at the front of a sentence where an adverb goes, and a word can be set down in either. A suffix has nowhere to be set down.
That is a description of eight renderings, not a finding, and it is exactly what a successor with a working source gate would be built to test.
One bound edit did leave a trace, in one hand only. At W5, gönderildi → gönderilmiş moved
T-b from "A horse was sent to town" to "A horse had been sent to town" — the evidential
pushed into the English pluperfect — while T-a wrote "A horse was sent to town" both times. One
hand, one site; named here because P3's per-window rule would have named it and because it is the
only candidate in the run for a bound mark relocating.
6. The translation limb, and the contamination measurement
T-kasagi-R06-v1: Ömer Seyfettin's «Kaşağı» (1919) whole, 887 Turkish words → 1,500
English, R06 single pass, log D1–D18 frozen with the draft. The project's nineteenth
source language and its first Turkic one. A boy breaks his mother's currycomb, blames his small
brother, watches his father slap and banish him, and is still rehearsing the confession the morning
the brother dies of diphtheria.
Contamination suspected, and the boundary is one word. Before translating, the lead saw
search-result titles including "The Currycomb (Kaşağı)" and a two-sentence plot blurb; no
sentence of any English rendering was read. Measured after the freeze against the only freely
reachable English (a 2020 web version, itself visibly machine-produced and made from a third
Turkish recension): 16 shared 7-grams, 0 twelve-grams, 0 fifteen-grams, longest common run 11
tokens — clean, by tools/dependence_check.py --unicode.
The copy-text finding is separate and travels further than this run does. collation.md: the
two freely reachable Turkish e-texts of this story are not two witnesses to one text but two
different modernisations — 826 tokens against 819, agreement 0.852, 103 divergence runs — and
at [5] they modernise in opposite directions (uyumlu in the conservative-looking witness
against ahenkli in the school one). No reachable text is the 1919 text. All seven of this
design's edited strings were checked in both recensions before the freeze and all seven are present
verbatim, so nothing measured here turns on a word the two texts disagree about.
7. Limits, and what this licenses
It licenses nothing about carriers. G1 failed; P1, P2, P3 and P4 are all withheld.
- Two hands and three arbiters, one of which is measurably not reading the source language.
- Six windows. Even a passing gate would have left the primary underpowered — it was.
- One language, one story, one translator's set of loci, chosen by the lead from its own frozen log; the lead is not blind to them.
- Variant B is not attested text in either arm, which is what
G1existed to make safe. LEXedits add a word andGRAMedits substitute one (design §3).G4at 0.0000 bounds that rival explanation but does not remove it: an inert added word is not the same as a meaningful one.- Free-word carriage and explicit marking are not separable here (critic
A4). "The carrier predicts arrival" and "explicitness predicts arrival" remain two readings of the same pattern, and §5's slot description is a third. - Three cells are void and none is imputed —
J3at a 900-token cap, all three spending the whole cap on hidden reasoning and returning null content. Note (bmb), on the same slug S157 recorded it on. 147 of 150 cells parsed, 0.980, above the 0.90 failure bar. They were not re-dispatched: the primary was already withheld and a re-dispatch would have introduced a second cap into one seat's cells for no gain. - No jury scored anything, so nothing here says whether any of it costs a reader anything.
- Tier D is NOT PASSED; every evaluative sentence on this page is
internal-judgment-onlyandprovisional.
8. What a successor needs, and it is a gate rather than a design
The design is sound and the instrument under it is not. Three things, in order:
- A source-side arbiter panel that can read the source.
G1should be run as its own cheap stage before any hand is paid, with seats admitted on it;J3's 0.083 would have excluded it at a cost of three calls. The admission rule must be registered before dispatch — which is exactly why §3's drop-J3figure releases nothing here. - Bound edits with a demonstrated source-side signal.
W3scored 0/3 in the Turkish. A successor picks loci that clear the parity bar first, and there is no reason it must build them from a story it also wants to translate. - More than six windows, or a graded response rather than a binary, since the binary saturates the free arm and floors the bound one.
9. Cost and reconciliation
201 bodies. $0.290330128 of a declared ceiling of $1.46 — 19.9%.
| stage | calls | cost |
|---|---|---|
pre-run critic, qwen/qwen3.7-max (reserve slug, neither hand nor arbiter), cap 12,000 |
1 | $0.020465625 |
| Stage T, 2 hands × 25 | 50 | $0.113725100 |
| Stage J, 3 arbiters × 50 items | 150 | $0.156139403 |
waste — 3 dead bodies, all deepseek/deepseek-v4-pro, finish_reason: length, null content, 900 reasoning tokens each |
— | $0.004260050, 1.47% of spend |
Key reconciliation: 94.050841400 → 94.341171516, delta 0.290330116 against a per-response sum of 0.290330128 — 1.2 × 10⁻⁸ apart, the third exact reconciliation in a row. Between-session drift into this session's opening snapshot was +0.008957900, the sixth consecutive session to show it.