Repository path: workshop/experiments/E-20260812d-slot-or-carrier/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260812d-slot-or-carrier |
| status | frozen |
| created | 2026-08-12 |
| updated | 2026-08-12 |
| senses | — |
| links | workshop/translations/shukuhai/R06-v1/translation.md, workshop/translations/shukuhai/source-aozora-60665.txt, wiki/arms/ARM-carrier.md, wiki/findings/results/RS-20260811g-carrier.md, wiki/findings/results/RS-20260811f-dakghar-address.md, framework/v0.2/README.md, config/models.md, config/budget.md |
E-20260812d — the slot, not the carrier: five cells inside one Japanese story
Frozen 2026-08-12, before any call. The translation limb it hangs on — T-shukuhai-R06-v1,
Kikuchi Kan's 「祝盃」 (1920) whole, 2,816 Japanese characters to 1,513 English words — was frozen
and committed at 48f2856 before this design was written and before any window was chosen.
1. Where the question came from, and why it is not the question step 1 asked
ARM-carrier step 1 (E-20260811g, RS-20260811g-carrier) asked whether the carrier predicts
arrival: does a distinction the source puts in a free word reach English while the same distinction
in a bound morpheme does not? Its source-parity gate G1 failed and every primary was withheld.
But its §5 — eight renderings, described, gated by nothing — said something the carrier conjecture does not say:
What a free word buys is not a matching category in the target. It is a position. English has nowhere to put a second-person plural agreement suffix and no honorific paradigm to receive
ağam— but it has a place at the end of a call where a noun of address goes, and a place at the front of a sentence where an adverb goes, and a word can be set down in either. A suffix has nowhere to be set down.
That is a different predictor, and on the Turkish materials it is not separable from the
carrier, because there every free word happened to have an English slot and every bound morpheme
happened not to. This design separates them. It is the successor RS-20260811g §8 specified,
and it takes all three of §8's requirements: (1) a source-side panel admitted on a rule registered
before dispatch; (2) manipulations whose source-side signal is established before any hand is
paid, with selection done on that signal; (3) more than six windows, and a graded response instead
of a binary.
2. Question
When a source marks something English does not grammaticalise, what predicts whether an independent English hand puts in anything a reader can recover — the carrier the source uses (free word against bound morpheme), or whether English has a position the content can be set down in?
The two answers are separable because Japanese supplies all four combinations.
| English has a position | English has no position | |
|---|---|---|
| bound morpheme | BS — inferential 〜らしい suffixed to a narration predicate |
BN — plain → polite 〜です/ます on a speech predicate |
| free word | FS — inferential どうやら prefixed to a narration predicate · FV — address term 君/おい ⇄ 先生 |
FN — first person 僕 ⇄ 俺 |
Three contrasts, each holding everything else still. They are the design, not the 2 × 2:
FVvsFN— carrier constant (a free word), content constant (how the speaker stands to the person he is addressing), only the slot varies. English has a vocative; English has one I.FSvsBS— content constant (how directly the narrator claims to know), slot constant (English keeps a sentence-adverb position open), only the carrier varies.FNvsBN— content constant (the speaker's footing), slot constant (none), only the carrier varies.
The carrier account predicts FV > BN, FN > BN, FS > BS, and FN ≈ FV.
The slot account predicts FV > FN, FS ≈ BS, FN ≈ BN, and both FS and BS high.
They disagree on FV vs FN and on FS vs BS, which is why those two contrasts are the primary.
3. The wire between the limbs, in one sentence
The translation limb rendered the whole story under a regime that forbids revision, and its log
records at D9, D10 and D20 three marks the translator could not carry — the sentence-final
particles, the flat politeness level, and the 俺 of [43]; the study limb takes the same prose,
makes each of those the only thing that moves between two versions of a passage, and asks whether
independent hands carry them. The translation generated the losses; the study measures them.
The lead's own English at each selected site is printed on the result page as description with no gate. The lead is not a hand, does not arbitrate, and its renderings enter no statistic.
4. Materials — twenty candidate windows, six control items, one story
Copy-text workshop/translations/shukuhai/source-aozora-60665.txt, units [1]–[45], built by
build_items.py, which applies every edit as an exact string replacement and asserts that the
replaced string occurs exactly once in its window. Variants are never typed twice: variant A is
a verbatim slice of the copy-text at 18 of 20 windows. No paragraph is used by two items —
asserted in the builder; 41 of 45 paragraphs are used, each once.
Each window is one to three consecutive paragraphs. The five cells, four candidate windows each:
| cell | carrier | slot | content | operation | sites |
|---|---|---|---|---|---|
BS |
bound | yes | epistemic | append 〜らしい to the main predicate | [3] [6-7] [22] [25-26] |
FS |
free | yes | epistemic | insert どうやら before the subject | [4] [12] [23-24] [32] |
BN |
bound | no | social | plain → polite predicate, one token | [19-21] [30-31] [37-38] [11] |
FN |
free | no | social | 僕 ⇄ 俺, one token | [43-44] [13-14] [17-18] [9-10] |
FV |
free | yes | social | address term: pronoun 君/おい ⇄ name + 君 | [27-28] [33-34] [15-16] [39-40] |
Surface-magnitude matching, which step 1 did not have (RS-20260811g limit 5: "LEX edits add
a word and GRAM edits substitute one"). Within each contrast the two arms use the same
operation: BS and FS both add one element to an unhedged narrative assertion; BN and FN
both substitute one token inside a speech. FV substitutes one token at FV1/FV2 and, at
FV3/FV4, substitutes one token between two variants that both carry an inserted address term —
flagged BOTH-EDITED in items.json and named again in §7.
The four cells' content is not fully crossed with the slot factor and cannot be. BS/FS are
epistemic and BN/FN/FV are social. That is exactly why FV exists: it is social content in a
free word with an English position, so the slot factor is tested inside one content type at
FV vs FN, and the carrier factor inside one content type twice, at FS vs BS and FN vs
BN.
Controls
| id | kind | rated on | what it is |
|---|---|---|---|
CP1 |
C-POS |
FACT | [1] 五六年前 → 十五六年前 — a number English must carry |
CP2 |
C-POS |
FACT | [36] 気持が分った → 気持が分らなかった — a negation |
CP3 |
C-POS |
FACT | [45] ソーダ水 → 葡萄酒 — a lexical fact |
CN1 |
C-NULL |
SOCIAL | [2], byte-identical pair |
CN2 |
C-NULL |
SOCIAL | [29] つい/\ → ついつい — the iteration mark spelled out, semantically inert |
CN3 |
C-NULL |
EPISTEMIC | [41], byte-identical pair |
The C-NULL items are rated on the very questions the cells are rated on. A false-alarm bound
measured on a different question would not bound anything. Two of the three are byte-identical, so
no prompt anywhere in this run may assert that the two versions differ — see §5.
5. Procedure
Every raw body is written to runs/ before anything is computed from it (note (bco)). One item per
call, never batched; independent calls are dispatched concurrently, which is the standing reading of
judgment is never parallelized — that rule bars a seat from disposing of several judgments inside
one call, and every judgment here is one seat, one item, one call, as at E-20260811g.
Presentation order is fixed before dispatch by sha256("E-20260812d-slot-or-carrier|<item id>")
mod 2, computed in build_items.py and stored as flip in items.json. The same flip governs the
item on both the source side and the English side.
No prompt states that the two versions differ. Every rating prompt says they may or may not.
Stage 0 — admission, and it runs before any hand is paid
All five panel seats (config/models.md v1) rate the six control items on the source side, in
Japanese: 3 C-POS on FACT, 3 C-NULL on the SOCIAL and EPISTEMIC questions. 30 calls.
Admission rule, registered here before dispatch (strengthened on critic finding 4; the mean conditions are replaced by min/max conditions so that one extreme score cannot carry a seat in):
min(C-POS) ≥ 2, andmax(C-NULL) ≤ 1, andmean(C-NULL) ≤ 0.50, andmin(C-POS) > max(C-NULL).
Assignment rule, registered here before dispatch. Among admitted seats let
margin = mean(C-POS) − mean(C-NULL). The two seats with the largest margin are the hands; the
remaining three admitted seats are the arbiters. Ties break by slug, ascending. The rationale is
written before the numbers exist: the hand's task is to read Japanese and write English, the
harder reading; the arbiter's Stage-E task uses no Japanese at all, and its Stage-S task is the
screen it has just passed.
Contingency, registered. 5 admitted → 2 hands, 3 arbiters. Exactly 4 admitted → 2 hands, 2
arbiters, and the loss of the third vote is declared on the result page. Fewer than 4 admitted →
the run does not proceed and ARM-carrier closes retired, saying that the conjecture is
unreachable with this apparatus, which is the branch the arm's step 2 names.
Stage S — the source side, arbiters only
Each arbiter rates all 20 candidate windows in Japanese on that window's declared question.
60 calls (3 × 20). The C-NULL and C-POS values are carried over from Stage 0 for the admitted
arbiters; they are not re-run.
Prompt shape, with <dim> filled from the item:
Below are two versions of a short passage from a Japanese short story. They may or may not differ. Question: how large is the difference between them in
? 0no difference ·1a slight difference ·2a clear difference ·3a large difference. Answer as JSON:{"score": <0-3>, "what": "<one sentence on what differs, or 'nothing'>"}.
<dim> is one of exactly three strings, fixed here:
- SOCIAL — how the speaker positions himself in relation to the person he is speaking to
- EPISTEMIC — how directly the narrator claims to know what he reports
- FACT — what the passage states as fact
Selection — registered rule, applied after Stage S and before Stage T
For each cell, SRC(window) = mean over admitted arbiters. The three windows with the highest
SRC are selected; the fourth is dropped. Fifteen windows go forward. Ties break by window id.
Stage T — the hands
Each hand translates each selected variant, in its own call, blind: it is shown one passage, is not told a second version exists, and is not told what is being studied.
Translate the following passage from a Japanese short story into English. Render it as it stands. Give only the translation, with no comment.
Calls: 2 hands × (15 windows × 2 variants + 3 C-POS × 2 variants) = 72, plus 6 floor calls —
each hand translates variant A of BS3, FN3 and FV2 a second time, in a fresh call.
78 calls. temperature 0.0 throughout, as at E-20260811g.
Stage E — the English side, arbiters only, no Japanese shown
Each arbiter rates, for each hand, the pair of English renderings of that window, on the same question the window was rated on in Japanese. It is shown two English passages and nothing else. 3 arbiters × (15 + 3 + 3) items × 2 hands = 126 calls.
The floor items are the two renderings of the same variant by the same hand.
6. Registered gates and predictions
All statistics are means of integer scores on 0–3.
Gates — every one of these is evaluated before any primary is read
| gate | what it bounds | bar |
|---|---|---|
G0 |
admission (§5) | ≥ 4 seats admitted, else the run does not proceed |
G1a |
source visibility, per cell — a manipulation invisible in Japanese cannot show anything about English | cell-mean SRC ≥ 1.50; a cell below it has its English numbers withheld |
G1b |
source parity across the five cells — gates the two main effects only | max − min of the five cell-mean SRC ≤ 0.60 |
G1c |
pairwise source parity — gates each two-cell contrast separately | |SRC(X) − SRC(Y)| ≤ 0.60 for the pair each of P3, P4, P5 reads |
G2 |
the English instrument can see a difference it must see | mean ENG over C-POS ≥ 2.00 |
G3 |
floor — one hand against itself on identical input | mean ENG over floor items ≤ 0.60 |
G4 |
source false-alarm, recomputed over admitted arbiters only | mean SRC over C-NULL ≤ 0.60 |
G5 |
parse rate | ≥ 0.90 of dispatched bodies yield an integer score |
Parity withholds narrowly, not globally — this is where the design deliberately departs from
step 1, whose whole-run withholding on one parity gate produced a session with no measurement in
it. G1b gates only P1 and P2, the two five-cell main effects. G1c gates each
two-cell contrast on its own two cells, so a parity failure between, say, BS and FS cannot
withhold P3, which does not read those cells. Every difference statistic is printed with both
cell source means beside it, whether or not its gate held.
The transmission ratio ENG/SRC is not used (critic finding 3, accepted in full): on a bounded
0–3 integer scale a ratio is unstable near the floor and can inverse the apparent direction of an
effect. A contrast whose G1c fails is printed and marked PARITY-UNMET, and licenses nothing.
G1a withholds per cell, because a mark nobody can see in the source is not a test of anything.
If G2, G3 or G4 fails, every primary is withheld: those bound the instrument, not the
materials.
Predictions, registered
| statement | fires at | |
|---|---|---|
P1 |
slot main effect — mean ENG over {BS,FS,FV} minus mean over {BN,FN} |
≥ 0.75 |
P2 |
carrier main effect — mean ENG over {FS,FN,FV} minus mean over {BS,BN} |
≥ 0.75 |
P3 |
the decisive social contrast — FV − FN, carrier and content constant |
≥ 0.75 |
P4 |
the decisive epistemic contrast — |FS − BS|, content and slot constant |
≤ 0.50 |
P5 |
secondary, descriptive — the no-slot carrier contrast, |FN − BN| |
≤ 0.50 |
P5 is demoted to a descriptive check on critic finding 6, accepted: with both cells at the
floor it holds trivially and confirms nothing, so it may not carry a reading on its own. Registered
exploratory attached to it: if FN and BN both exceed 1.00, the arbiters' what texts for
those cells are read and reported — RS-20260811g §4 established that on this instrument the
reason text, not the count, is where a spurious arrival shows itself.
The reading rule, fixed before the run:
P1andP3fire,P2does not,P4andP5hold → the predictor is the position, not the carrier.framework/v0.2§2 gains that statement with its scope: one language, one story, two hands, three arbiters, fifteen windows.P2fires andP1does not → the carrier conjecture is supported andARM-carrier's original statement stands, on one language.- Both fire → they are not separated by this design; both are reported and the refusal is
written.
P3andP4are then reported as the only evidence bearing on which. - Neither fires → the cell means are reported and nothing is written into the framework.
A 90% bootstrap interval over windows is reported for P1, P2 and P3, 10,000 resamples,
seeded by sha256(design id). An interval that straddles a bar is reported as straddling it; the
bar is not moved (RS-20260811g §4).
7. What this design cannot establish, stated before the run
- One language, one story, one author, one period. Nothing here licenses a statement about
English in general, and
ARM-carrier's own constraint — two conditions is not a theory of English — binds the wording whatever the numbers do. - Slot availability is the lead's judgment, made before the run and not measured. That English has a vocative and no second I is not in dispute; that a sentence-adverb position is "the same kind of thing" as a vocative is a claim this design assumes rather than tests.
- Content is confounded with slot across the epistemic/social divide. Only
FVvsFNbreaks it, and that is one contrast on three windows. FV3andFV4compare two edited variants, neither attested.FV1andFV2are attested atA. If theFVcell's result rests onFV3/FV4the result page must say so, and the per-window numbers are printed so it can be seen.FV's B variant names a person, and a name is referential as well as relational. The swap is pronoun ⇄ name + 君 — licit peer address in 1920 Tokyo either way — so the pragmatic-anomaly confound is gone, but a rival reading survives and is stated here: English may carryFVbecause a name is propositional content a translator must keep, not because a footing has a position to sit in. This design cannot separate those two, and neither couldRS-20260811f'sC-ccontrol or Turkishağam.らしいandどうやらare not synonyms (critic finding 2, accepted in part). They are the nearest free/bound pair Japanese offers for inference-from-evidence, and they collocate — どうやら〜らしい is the ordinary form, which is the basis for treating the content as constant. They are not identical:らしいleans evidential-hearsay,どうやらconjectural.P4's 0.50 bar therefore tests practical equivalence, not synonymy, and the per-window source means are printed so a reader can see whether the two arms were in fact matched. The design does not restrictBSto non-past predicates as the critic proposed: past-predicateらしいis the form this very story attests at[15], and moving off it would make the manipulation less natural, not more.- The assignment rule may select hands for confidence rather than granularity (critic finding 4, residual). Admission was strengthened to min/max conditions, which removes the single-extreme case; the remaining concern is recorded here and every seat's six control scores are printed on the result page so the assignment can be audited.
- The copy-text has one witness (
T-shukuhai-R06-v1§Copy-text). No reading is collated. - No jury scored quality. Nothing here says whether any of it costs a reader anything.
- Tier D is NOT PASSED. Every evaluative sentence on the result page is
internal-judgment-onlyandprovisional. - Two hands. A hand-level effect and a general effect are not separable at n = 2, and the per-hand numbers are printed.
8. Pre-flight cost estimate
Today's earlier sessions spent $0.17249665 (S163 $0.0235655, S164 $0.06981315, S165 $0.0791180); the day's headroom is $4.82750335.
| stage | calls | cap out | worst case |
|---|---|---|---|
pre-run critic (qwen/qwen3.7-max, reserve slug, neither hand nor arbiter) |
1 | 14,000 | $0.075 |
| Stage 0 — 5 seats × 6 items | 30 | 400 | $0.09 |
| Stage S — 3 arbiters × 20 | 60 | 400 | $0.28 |
| Stage T — 2 hands × 39 | 78 | 700 | $0.35 |
| Stage E — 3 arbiters × 42 | 126 | 400 | $0.58 |
| routing headroom (S022: a slug can bill ~4× list through provider choice) | — | — | $0.73 |
| declared ceiling | 295 | $2.10 |
Ceiling raised from $1.60 to $2.10 on critic finding 5, accepted: the original headroom was ~17% of a 295-call run priced off a table that S022 showed can be wrong by 4× through provider routing alone. Hard stop, registered: if cumulative spend after Stage S exceeds $0.60, Stage T is dispatched with one hand only and every two-hand claim is withdrawn on the result page.
Worst cases are built from max_tokens, not from an expected length — note (abc). The stage costs
above price the dearest admissible assignment (arbiters = moonshotai/kimi-k3,
google/gemini-3.6-flash, openai/gpt-5.6-terra), because the assignment is not known until
Stage 0 has run.
Stage gating. Stage 0 is dispatched first and costs under $0.10; if G0 fails the run stops
there. Stage S is dispatched next; if G1a withholds three or more cells the run stops before any
hand is paid, since the design would then have nothing left to compare.
9. Pre-run critic — adjudication
One independent adversarial pass over the frozen design, items.json and build_items.py, run
before any Stage-0 call, on qwen/qwen3.7-max — the reserve slug, neither a hand nor an
arbiter. Verdict NEEDS-AMENDMENT, six findings, two BLOCKING. Full text: critic.md.
Five accepted in full, one accepted in part with the overrule written. Every amendment below
was made before Stage 0 was dispatched.
| # | tag | finding | disposition |
|---|---|---|---|
| 1 | BLOCKING | FV3/FV4 put 先生 in a peer's mouth: a pragmatic anomaly whose salience would be indistinguishable from a slot effect, and P3 — the decisive contrast — rests on FV |
Accepted in full. The FV swap is now pronoun 君/おい ⇄ name + 君, licit peer address in every window and in both variants. items.json rebuilt; §4 and §7.5 rewritten |
| 2 | BLOCKING | らしい (evidential) and どうやら (conjectural) are different epistemic sub-categories, so FS vs BS does not hold content constant |
Accepted in part, §7.6. The reframing is taken and P4 is restated as a test of practical equivalence rather than synonymy; per-window source means will be printed. The proposed repair — restrict BS to non-past predicates — is overruled: past-predicate らしい is the form this story attests at [15], and the substitute would be less natural Japanese, which would show up in SRC as a larger difference and confound the thing the repair was meant to protect |
| 3 | SERIOUS | the G1b ratio fallback ENG/SRC is unstable on a bounded integer scale and can invert an effect |
Accepted in full. The ratio is struck. Parity now withholds narrowly: G1b gates only the five-cell main effects, and a new G1c gates each two-cell contrast on its own two cells |
| 4 | SERIOUS | admission on means is carried by one extreme score; assignment by margin may select confidence rather than granularity | Accepted in part. Admission is now min/max: min(C-POS) ≥ 2, max(C-NULL) ≤ 1, mean(C-NULL) ≤ 0.50, min(C-POS) > max(C-NULL). The proposed sd caps are overruled — an sd over three items is too noisy to exclude a seat on — and the residual concern is written into §7.7 with every seat's six scores to be printed |
| 5 | MINOR | the ceiling leaves too little headroom for a 295-call run under provider routing | Accepted in full. Ceiling $1.60 → $2.10, plus a registered hard stop at $0.60 after Stage S |
| 6 | MINOR | P5 holds trivially whether both cells floor or both rise |
Accepted in full. P5 demoted to a descriptive check with a registered exploratory read of the arbiters' reason texts |
Cost of the critic pass: $0.027901100, finish_reason: stop, one call, no re-dispatch.