Repository path: workshop/experiments/E-20260811g-carrier/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260811g-carrier |
| status | frozen |
| created | 2026-08-11 |
| updated | 2026-08-11 |
| senses | — |
| links | workshop/translations/kasagi/R06-v1/translation.md, workshop/translations/kasagi/collation.md, wiki/arms/ARM-carrier.md, wiki/findings/results/RS-20260811f-dakghar-address.md, wiki/findings/results/RS-20260808g-two-pasts.md, framework/v0.2/README.md, config/models.md, config/budget.md |
E-20260811g — the carrier, not the content: a 2 × 2 inside one Turkish story
Frozen 2026-08-11, before any call. The translation limb it hangs on (T-kasagi-R06-v1, the
whole story) was frozen and committed at d58caa5 before this design was written, and its
contamination was measured after that commit and before any locus was chosen.
1. Where the question came from
It came from translating, at log decision D6, and from a control in the previous session's run.
Every second-person form in «Kaşağı» is sen; siz never occurs. D6 decided to compensate for
none of it and to log eighteen sites as losses. That is the same decision ARM-dakghar's V5 made
about Bengali, and RS-20260811f then measured it: CROSS 0.0000, 60 of 60 cells SAME — the
honorific verb ending reached English at exactly zero.
But that run's own added control did reach English, at 0.8333. C-c changed the same social
relation by changing an address term — a free word — and the arbiters saw it. Same play, same
seats, same call, same question; the only difference was what the Bengali used to carry the mark.
The conjecture this design tests: what predicts arrival is the carrier, not the content. A distinction the source expresses in a free word reaches English; the same distinction expressed in a bound morpheme does not.
Two things stop that being established, and the design is built against both. (a) It has been seen only in the social domain, so English does not do social is a rival explanation of equal standing — so content is crossed with carrier. (b) The free-word arm was a positive control chosen to be easy and was never matched against the bound arm as an equal arm — so here both arms are primaries, in the same windows, against the same baseline rendering.
2. Question
Holding the language, the work, the passage, the translating hands, the arbiters, the question and the statistic constant, does a one-word change to a passage of Turkish reach an independent English hand's rendering more often when the word is free than when it is bound — and does that hold for an epistemic distinction as well as a social one?
3. Materials — six windows, three variants each, one word apart
Every variant is built by exact string substitution from the frozen copy-text
workshop/translations/kasagi/source-trwikisource.txt (tr.wikisource revid 172873), by
build_items.py, which asserts that each replacement fires exactly once in its window.
Each window has an A (as printed), a GRAM (the mark carried by a bound morpheme) and a LEX
(the same mark carried by a free word). A is shared by both contrasts, so GRAM and LEX are
measured against the identical baseline rendering.
SOCIAL — the mark elevates the footing between two speakers
| window | units | site (as printed) | GRAM (bound) |
LEX (free word) |
|---|---|---|---|---|
W1 |
[23]–[29] | — Gel buraya! (father → his groom) |
— Gelin buraya! (2sg → 2pl) |
— Gel buraya, ağam! (+ honorific address) |
W2 |
[50]–[58] | — Niye ağlıyorsun? diye sordum. (child → the servant) |
— Niye ağlıyorsunuz? … |
— Niye ağlıyorsun, abla? … |
W3 |
[2]–[9] | — Hadi yap! derdi. (groom → the master's son) |
— Hadi yapın! derdi. |
— Hadi yap, beyim! derdi. |
EPISTEMIC — the mark makes the narrator's access indirect
| window | units | site (as printed) | GRAM (bound) |
LEX (free word) |
|---|---|---|---|---|
W4 |
[36]–[41] | Zavallının bir şeyden haberi yoktu. |
… haberi yokmuş. (-DI → -mIş) |
Zavallının anlaşılan bir şeyden haberi yoktu. |
W5 |
[49] | Kasabaya at gönderildi. |
Kasabaya at gönderilmiş. |
Anlaşılan kasabaya at gönderildi. |
W6 |
[20b] | … penceresiz küçük bir odası vardı. |
… odası varmış. |
Anlaşılan ahırın köşesinde Dadaruh'un … odası vardı. |
[20b] is the sub-span of the printed paragraph [20] from Ben bir gün yalnız başıma kaldım. to
Az daha sevincimden haykıracaktım. — declared here because it is the one window that is not a whole
printed unit.
C-NEG — the positive control, and it is a bound morpheme
| window | site | edit | why |
|---|---|---|---|
W1 |
— Bilmiyorum, dedi. |
— Biliyorum, dedi. |
one bound morpheme (the negative -mA-), and English must carry it. If bound morphemes were simply invisible to these hands, this would fail too. It is the design's answer to the GRAM arm just changes less of the surface. |
LEXNULL — the added-word control, added on the critic's SERIOUS 3
| window | site | edit | why |
|---|---|---|---|
W1 |
— Gel buraya! |
— Gel buraya, hadi! |
one free word, carrying neither a social nor an epistemic mark |
W5 |
Kasabaya at gönderildi. |
Hemen kasabaya at gönderildi. |
one free word, temporal, carrying neither mark |
Together C-NEG and LEXNULL bracket the artifact. C-NEG is a bound edit that must arrive;
LEXNULL is a free edit that should not. If both behave, neither bound is invisible nor free
is always visible can explain P1.
Both recensions agree at every locus
collation.md establishes that the two reachable Turkish e-texts are different modernisations.
All seven edited strings were checked in the MEB school recension before the design was frozen and
all seven are present verbatim; only surrounding words differ (Hadi/Haydi, Niye/Neden,
haykırdı/bağırdı, birdenbire/aniden). No measurement here turns on a word the two texts
disagree about.
The asymmetry this design does not remove, stated before the run
GRAM edits substitute one word; LEX edits add one. So a LEX variant is one word longer and
its English will usually be one phrase longer. This is registered as the leading rival explanation
of any positive result — arbiters notice added length — and the design's defence against it is
G1 below plus C-NEG, not a claim that it has been eliminated. Both arms are one-word edits and
both add a mark rather than remove one; that is the extent of the matching.
4. Procedure
Stage T — the independent hands. Two non-Anthropic panel seats, neither the lead and neither an arbiter, are each given one variant at a time, alone, with no mention of politeness, evidentiality, address, or that an experiment exists. The prompt asks for an English translation of the passage and nothing else.
T-a=moonshotai/kimi-k3,T-b=x-ai/grok-4.5(config/models.mdP4, P3).- Per hand: 6 windows × 3 variants (
A,GRAM,LEX) + 2LEXNULL+ 1C-NEG+ 4A-repeats in independent calls = 25 calls; 50 hand calls in all. - The
A-repeat exists to measure the floor: how much one hand differs from itself on identical input. No hand sees two variants of a window in one context.
Stage J — the arbiters. Three non-Anthropic seats, none of them a translating hand, are shown two passages at a time and asked one fixed question, blind to the item's class and to the design.
J1=openai/gpt-5.6-terra,J2=google/gemini-3.6-flash,J3=deepseek/deepseek-v4-pro(P1, P2, P5). No seat both translates and arbitrates.- The question is deliberately symmetric across the two contents, so that neither arm is favoured by what the arbiter is asked to look for:
"Below are two versions of the same short passage. Ignore differences of wording, phrasing or style that make no difference to the sense. Do the two versions convey the SAME thing, or something DIFFERENT, about (a) what is asserted or implied about the people and what happens, (b) how the speaker stands socially to the person spoken to or spoken about, and (c) what the speaker claims to know and how they claim to know it? A difference of wording or style alone is not a difference."
Answer as JSON: {"verdict": "SAME"|"DIFFERENT", "what_differs": "..."}.
- Item classes, 50 items × 3 seats = 150 calls:
- SOURCE — the two Turkish variants (A vs GRAM, A vs LEX). 6 × 2 = 12.
- CROSS — hand h's English of A against hand h's English of the variant. 6 × 2 × 2 hands
= 24.
- LEXNULL — the added-word control, W1 and W5 only, 2 × 2 hands = 4 (critic A3).
- FLOOR — hand h's English of A against hand h's English of A again, on W1, W2,
W4, W5 only, 4 × 2 = 8 (narrowed from six windows to pay for LEXNULL, critic A3).
- CONTROL — hand h's English of W1 A against hand h's English of W1 C-NEG. 2.
- Item order is shuffled under a seed fixed in the built file; which passage is shown first is
randomised per item and recorded.
Measure. d = 1 if the seat returns DIFFERENT, 0 if SAME. A cell that does not parse is
void and reported, never imputed.
5. Registered gates and predictions
Written before any call. Every gate is a withholding gate: if it fails, the primary is not read.
G1— source parity, and it is load-bearing here, not a checksum. OnSOURCEitems, theGRAMandLEXcontrasts must both come backDIFFERENTat mean ≥ 0.667, and their means must differ by ≤ 0.25. If the free-word edits change the Turkish more than the bound edits do, an English difference measures the size of the edit and not the carrier. This is the gateRS-20260811fdemoted because nothing there turned on it; here everything does.G2— the bound-morpheme positive control.C-NEGonCONTROL, mean ≥ 0.75. A bound morpheme that English must carry has to arrive, or the GRAM arm measures the arbiters' blindness to short edits rather than anything about carriers.G3— the floor. OnFLOORitems, meandmust be ≤ 0.40. The run is void only ifFLOOR ≥ CROSSoverall.G4— the added-word control, ADDED on the critic's SERIOUS 3, and it is the gate that matters most.LEXNULLonCROSS, meand≤ 0.50.LEXNULLadds one free word carrying neither a social nor an epistemic mark (— Gel buraya, hadi!;Hemen kasabaya at gönderildi.). If adding any word makes an arbiter call the two Englishes different,P1measures added-word salience and not the carrier.P1— the primary, and it can fail in either direction. Over the six windows, onCROSS: meand(LEX) − meand(GRAM) ≥ 0.40 ⇒ the carrier predicts arrival. A bootstrap 90% interval over windows is reported alongside; if it straddles 0.40 the verdict says so and the threshold is not treated as a boundary the data can resolve.P2— the crossing, which is what makesP1more than a restatement ofRS-20260811f.P1's difference must hold in both contents separately:LEX − GRAM ≥ 0.25withinSOCIALand withinEPISTEMIC. If it holds inSOCIALonly, the rival explanation — English does not do social relations, and this is not about carriers — survives and the result page says so as the primary reading.P3— the null half, registered so that it can fail. Over the six windows,d(GRAM) −d(FLOOR) ≤ 0.20 ⇒ a bound mark reaches English at the rate one hand differs from itself, replicatingRS-20260811fin a second language family. If any single window reachesd(GRAM) −d(FLOOR) ≥ 0.50 it is named and the null is stated as predominantly, not universally.P4— carriers named, not just counted. A positiveP1is only reported as relocation if the arbiters'what_differstext names an English device (an adverb, a modal, a vocative, a title, a choice of verb) at ≥ half theLEXwindows that drove it. A numerical difference with no nameable carrier is reported as unexplained.
Failure criteria. The run is void if: fewer than 90% of cells parse; any hand returns a body that
is not a translation (a refusal, a commentary, a transliteration); or FLOOR ≥ CROSS overall.
6. What this design cannot establish
- Two hands is two hands, three arbiters is three arbiters. Any null is about these seats on this material, not about English.
- Six windows. The denominator is small and
P1's margin is not a confidence interval. - The arbiters share a training distribution with the hands. "An independent reader cannot tell them apart" means these arbiters.
- Variant B is not attested text, in either arm.
G1is what stops that being fatal. - One language. Even a clean crossing licenses a statement about Turkish→English, and the arm page binds step 2's wording to that.
- The added-word asymmetry of §3 is bounded, not removed.
- Nothing here scores the lead's own translation, which no seat sees; the lead never judges its own translation (charter §5).
- Free-word carriage and explicit marking are not separable here (critic
A4). A free word that carries evidentiality or deference is by that fact more explicit than a suffix. "The carrier predicts arrival" and "explicitness predicts arrival" are two readings of the same result and this design does not choose between them. - Every
GRAMedit is a pragmatically loud violation and everyLEXedit is pragmatically consistent (criticA1), becausesizoccurs nowhere in this story. This biases against the registered prediction and is registered here before the run.
7. Declared deviations
- No
senses:are invoked. This is a descriptive measurement of what survives into English, not an evaluation. Any evaluative sentence on the result page carriesinternal-judgment-only. - The work was chosen for its grammar. Turkish was selected because it grammaticalises both an
epistemic and a social distinction that English lacks. The windows and the edits were chosen
after the translation and its log were frozen and committed at
d58caa5, from the log's ownD6,D7andD8; the lead is not blind to them. The design's protection is that the lead neither translates for the measurement nor arbitrates it. EPISTEMIC-LEXuses one device three times (anlaşılan);SOCIAL-LEXuses three (ağam,abla,beyim). The social device has to fit its speaker pair and no single Turkish word does all three. Recorded, not corrected.
8. Pre-flight cost estimate
Built from max_tokens, not from expected answer length (note (abc)).
| stage | calls | prompt (est.) | max_tokens |
worst case |
|---|---|---|---|---|
pre-run critic, qwen/qwen3.7-max (reserve slug, not a seat) |
1 | ~6,000 | 12,000 | $0.062 |
Stage T, T-a moonshotai/kimi-k3 ($3/$15) |
25 | ~400 | 400 | $0.180 |
Stage T, T-b x-ai/grok-4.5 ($2/$6), reasoning low |
25 | ~400 | 800 | $0.140 |
Stage J, J1 openai/gpt-5.6-terra ($1/$6), reasoning off |
50 | ~500 | 900 | $0.295 |
Stage J, J2 google/gemini-3.6-flash ($1.5/$7.5), reasoning low |
50 | ~500 | 1,200 | $0.488 |
Stage J, J3 deepseek/deepseek-v4-pro (list $0.632/$1.263, priced at the 4× routing caution) |
50 | ~500 | 900 | $0.291 |
| 201 | $1.456 |
Declared ceiling: $1.46. UTC day 2026-08-11 stands at $3.497487166 of $5.00 before this run,
so the ceiling leaves $0.04 of the day's cap unspent even if every call bills its worst case.
A re-dispatch that would carry the run past the ceiling is not made: it is deferred to the next
UTC day and reported as a gap, never imputed. max_tokens is 400–1,200 because a cap prices the
estimate and must also be large enough to hold an answer (S148, S157, S160 all truncated a body).
9. Pre-run critic — adjudication
One pass, qwen/qwen3.7-max (reserve slug, neither a hand nor an arbiter), cap 12,000,
finish_reason: stop, $0.020465625. Verdict NEEDS-AMENDMENT, 8 findings — 1 BLOCKING,
4 SERIOUS, 3 ADVISORY. Full text at critic.md. Five accepted, three accepted in part with the
overrule written.
| # | finding | adjudication |
|---|---|---|
| B1 | W2 GRAM is confounded: a child switching to siz with his own nanny marks alienation or sarcasm, not deference, so a DIFFERENT verdict may be detecting this sounds wrong for this character. LEX (abla) reinforces the footing while GRAM violates it — an asymmetric contrast. |
ACCEPTED IN PART, and the remedy the critic asks for does not exist. siz occurs nowhere in this story, so there is no window where a 2pl edit is pragmatically unmarked, and swapping W2 would move the problem, not remove it — W1 (master → groom) and W3 (groom → master's child) carry the same asymmetry. Amendment A1: the asymmetry is registered here, before the run, as a direction-of-bias statement. Every GRAM edit is a pragmatically loud violation of an established footing; every LEX edit is pragmatically consistent. That biases against the registered prediction: a loud source change that still fails to reach English makes P3 stronger, not weaker, and a positive P1 obtained despite it is conservative. The result page states this as the first thing a reader should know about the social arm. |
| S2 | G1 is near-tautological: adding a content word or a morphological marker always changes the Turkish, so G1 will pass at ~1.000 and cannot show that the two edits are comparable in size. |
ACCEPTED IN PART. Overruled on removal, and the reason is written: G1's second clause (\|GRAM − LEX\| ≤ 0.25) is what makes it a parity gate, and its failure mode is real — finding S5 below is exactly the case where GRAM would come back below 0.667 in the source and the primary would be unreadable. Overruled on the proposed fix (a 0–3 magnitude scale on SOURCE items only): the design's comparability rests on one binary statistic used identically at every item class, and a different instrument on one class would break it. Amendment A2: every SOURCE item's what_differs text is reported verbatim on the result page, so that what the arbiter saw can be checked even though how much is not measured. |
| S3 | No LEX-NULL control. G1 shows the Turkish differs and C-NEG shows bound morphemes can travel, but nothing shows that adding any semantically inert free word does not by itself trigger DIFFERENT. Without it, "carrier effect" is not distinguishable from "added-word salience". |
ACCEPTED, and it is the amendment that matters. Amendment A3: LEXNULL added to W1 (social) and W5 (epistemic) — — Gel buraya, hadi! and Hemen kasabaya at gönderildi., one added free word each, carrying neither a social nor an epistemic mark. Promoted to a withholding gate G4. Paid for inside the frozen call budget by measuring FLOOR on four windows instead of six (W1, W2, W4, W5), so the run is still 50 items and 201 calls and the declared ceiling does not move. |
| S4 | anlaşılan centres the narrator's inference process explicitly; -mIş need not. A positive result may show that English likes explicit metacognition, not that it likes free words. |
ACCEPTED as a limit, not remediable by this design. A free word that carries evidentiality is more explicit — that is not a confound sitting beside the mechanism, it is a description of the mechanism, and the same is true of abla on the social side. Amendment A4: §6 gains limit 8 — free-word carriage and explicit marking are not separable here, and "the carrier predicts arrival" and "explicitness predicts arrival" are two readings of the same result. The result page carries both. |
| S5 | -mIş is polysemous: in literary narration it can be a plain narrative perfective rather than an evidential, so W5/W6 GRAM might score zero because the wrong reading won, not because bound marks fail. |
ACCEPTED, and it is already gated. G1 reads the Turkish: if arbiters given gönderildi against gönderilmiş return SAME, G1 fails on the EPISTEMIC-GRAM cell and the primary is withheld. Amendment A5: G1 is evaluated per content as well as pooled, and the EPISTEMIC-GRAM source cell is reported by name whatever it does. Overruled on the proposed extra validation call: it would duplicate a gate the design already pays for. |
| A6 | The three social LEX terms (ağam, abla, beyim) are three different treatments with different translation probabilities; at n = 3 one outlier could carry the arm. |
ACCEPTED. Amendment A6: per-window numbers are printed for every arm, and P3's single-window naming rule is extended to LEX — if one window drives the social LEX effect, the result page names it and states the effect as that window's. |
| A7 | G3's floor at ≤ 0.40 is permissive; a floor above 0.25 should trigger a review before arbitration. |
ACCEPTED as a reporting rule. Amendment A7: if FLOOR > 0.25 the result page says the instrument was noisy and discounts P1's margin accordingly. The gate itself stays at 0.40, which RS-20260811f's critic set for a stated reason (a floor raised by benign paraphrase costs power, not validity). |
| A8 | Naming dimensions (a)/(b)/(c) in the arbiter question primes the arbiter to scan for exactly what the LEX variants supply. |
ACCEPTED IN PART, and overruled on the proposed generic question. A generic "do these convey the same meaning?" is what RS-20260811f's ADVISORY 6 was written against, and here it would be worse: an unnamed dimension is one an arbiter may not look for, which penalises the subtle arm — GRAM — more than the explicit one, biasing toward the registered prediction rather than away from it. Naming all three is what makes the question symmetric between the two contents, which is its purpose. Amendment A8: one sentence added to the prompt before dispatch — "an extra word that adds nothing to the sense is not a difference" — which attacks the length artifact at the point where it would operate, at zero cost. |
What the critic bought. One withholding gate the design did not have (G4), a registered
direction-of-bias statement on the whole social arm (A1), a per-content reading of G1 (A5),
and a limit (A4) that changes how a positive result may be worded. The declared ceiling does not
move: A3 was paid for out of FLOOR breadth, not out of budget.