Repository path: workshop/experiments/E-20260811h-domestication-channel/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260811h-domestication-channel |
| status | frozen |
| created | 2026-08-11 |
| updated | 2026-08-11 |
| senses | cultural-mediation, perceived-source-carriage |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-realia-channel.md, wiki/findings/results/RS-20260811b-realia-channel.md, workshop/translations/jutrenje/R06-v1/translation.md, workshop/experiments/E-20260811b-realia-channel/materials/realia.json, wiki/goodness-senses.md, framework/v0.2/README.md, config/models.md, config/budget.md |
E-20260811h — which road out of a culture-bound word places the English
ARM-realia-channel step 2, the arm's closing step. Frozen before any API call.
1. The question, and why it is not step 1's question
RS-20260811b (step 1) deleted the source-culture world from seven translated passages and asked
three seats which national variety of English the prose was written in. The verdict did not move:
+0.048, P = 0.3125, and the pre-run critic's equivalence margin was not met either, so the run
concluded no evidence of leakage, and not enough evidence of separability.
But it found something it had not designed for. Of the twelve realia cues the seats gave for the prose question, every one was one of three words — halfpennies, smock, councillor — and every one is the translator's own Anglicisation of a foreign thing. The genuinely foreign words in the same corpus — taiga, yamen, Sanzu-no-Kawa, the Dragon Throne — were cited 118 times for where is this set and not once for whose English is this.
That is a conjecture with a predictor in it, and it is about the oldest decision in translation:
A culture-bound item has three roads out of it — carry the source word over, substitute the nearest thing in the target's own world, or generalise to a location-free phrase. The conjecture is that only the second road moves where a reader places the English itself, and that the first road, which is every bit as conspicuously foreign, does not.
If that holds, a translator who domesticates in order to make a translation read naturally in
English is not producing neutral English. They are producing English that reads as belonging to
a particular country — a cost wiki/goodness-senses.md's cultural-mediation does not name, and
one framework/v0.2 §7 currently has no measurement of.
What this unit teaches about translating literature (the subject rule, wiki/tracks.md): which
of the three standard treatments of a culture-bound item changes a reader's placement of the
translator's English, as against their placement of the story's world. It is a question about a
translator's choice, not about this project's instruments.
2. Three changes from step 1, each forced by something step 1 found
- Three arms, not two. Step 1's
KEEPconflated transfer with domestication — Garnett's halfpennies sat in the same arm as Shaw's Sanzu-no-Kawa.TRA/DOM/NEUseparates them, and the whole finding lives in the contrast between the first two. - Proper names are held constant and excluded from the manipulation. Step 1's spans included
St Petersburg, Yakov, the great Lena, Yamashiro-Ya, so its
MUTEarm could be read as relocating the story rather than generalising its furniture. Here nothing moves the story out of Serbia or Russia, andC5checks that. - Forced pairwise comparison, not an absolute verdict. Step 1's §9.1 is explicit that at n = 7 passages only an effect consistent in 6 of 7 could have reached P = 0.05. A within-pair forced choice removes between-item variance and yields up to 48 binary judgments per contrast. This is a weaker claim than step 1's and is registered as such (§8, limit 1): it measures whether two renderings are distinguishable in a direction, not whether a verdict changes. Step 1 already supplies the absolute-verdict null.
3. Materials
Translation limb: T-jutrenje-R06-v1 — Laza Lazarević, «Први пут с оцем на јутрење» (1879),
§I and §II, 1,519 Serbian words → 1,948 English, frozen at commit 287916d before this design
existed. The project's twentieth source language and its first South Slavic one. Contamination
not measurable — no English rendering of the work is reachable — and declared as an assertion
on the artifact; the within-passage design does not turn on it (§8, limit 6).
The wire between the limbs, in one sentence. The translator's R06 log records, at each
culture-bound site, the two renderings the translator did not choose — the source word carried
over, and the nearest English-domestic article — and the study limb puts all three to independent
readers to find out which of them changes where they place the English itself.
Eight passages, four forms each, built by code.py, which is deterministic and makes no API
call:
| passage | hand | source lang | words | sites | STRONG sites |
|---|---|---|---|---|---|
W1 |
lead R06 |
Serbian | 148 | 16 | 4 |
W2 |
lead R06 |
Serbian | 144 | 6 | 5 |
W3 |
lead R06 |
Serbian | 116 | 7 | 6 |
W4 |
lead R06 |
Serbian | 259 | 5 | 3 |
W5 |
lead R06 |
Serbian | 221 | 5 | 3 |
W6 |
lead R06 |
Serbian | 347 | 3 | 1 |
A |
Garnett | Russian | 118 | 3 | 3 |
B |
Field | Russian | 128 | 4 | 3 |
A and B are E-20260811b's passages A and B in their US-spelled base form, with new DOM
renderings added to spans that experiment had already declared. A carries halfpennies itself —
the word that generated the conjecture is in the corpus, in the hand that wrote it.
The four forms. TRA (source word carried over) · DOM (nearest article of the English
domestic world) · NEU (location-free English phrase) · SHAM (NEU with a
location-neutral-to-location-neutral substitution at every site, matched site for site).
brit is frozen in the translator's log, not assigned later. STRONG = the domesticating
rendering names a specifically British thing (sovereign, shilling, firkin, hogshead,
Michaelmas, pitch-and-toss, indentures, counterpane, rushlight, billycock, on tick,
Civil Service, Whitehall). WEAK = English-domestic with no particular nation in it. 28 of 49
sites are STRONG.
Orthography is forced to AMERICAN in every form of every passage by one mechanical map applied
identically to all four (code.py SPELLING). No "reads more British" verdict in this run can be
produced by a spelling; G4b asserts no British spelling survives anywhere.
4. Procedure
Each pair shows two forms of the same passage, labelled A and B, in both orders. Two questions, each elicited on its own call:
A7 — both prompts are frozen verbatim here, and one pair gets one question on one call. The
wrapper is run.py TEMPLATE; {question} is one of exactly these two strings:
BRIT— "Ignoring what the story is about and where it is set, and judging only the English itself -- its diction and its writing -- which version reads more as though it were written by a British writer for British readers?"FOREIGN— "Which version is more clearly set outside the English-speaking world?"
Each is followed by "Answer NEITHER if there is no difference." and by the required one-line JSON
shape {"answer": "A"|"B"|"NEITHER", "cue": "<quoted exactly from the passage>", "confidence": 0-3},
then the two passages under --- A --- and --- B ---.
A5 — every asterisk is stripped from every form (code.py deitalic(), critic finding F5).
Markdown italics marked the TRA arm and only the TRA arm, so a seat could have read the source
words as quotation rather than as prose. No arm is now typographically marked.
Six contrasts, all eight passages, three seats = 64 pairs, 192 calls, plus 6 duplicate pairs
(18 calls) for G5. C1 and C2 carry the claim and run in both orders on all eight
passages (16 pairs each). C3, C4, C5, C6 run in one order each, alternating by
passage so that order is balanced 4/4 inside every control contrast at half the calls. This is
a scale-down from both-orders-everywhere, forced by the two dead critic bodies in §7, and it is
declared rather than absorbed: the controls are order-balanced but not order-paired, so an order
effect large enough to move a control by more than the 4/4 balance absorbs would not be visible
inside that contrast. G2 measures position preference across all five contrasts and is what would
catch it.
| id | comparison | question | what it is for |
|---|---|---|---|
C1 |
DOM vs NEU |
BRIT |
the primary |
C2 |
TRA vs NEU |
BRIT |
the equivalence half — a foreign word is equally marked and should not move this |
C3 |
SHAM vs NEU |
BRIT |
edit-presence control |
C4 |
TRA vs NEU |
FOREIGN |
positive control — the seats must be able to see the foreign loading when asked about the world |
C5 |
DOM vs NEU |
FOREIGN |
confound check — domesticating must not move the narrative setting |
C6 |
DOM vs TRA |
BRIT |
the asymmetry, tested directly (A3, critic F3) |
Seats. P1 openai/gpt-5.6-terra, P3 x-ai/grok-4.5, P2 google/gemini-3.6-flash
(reasoning: {effort: low}, as E-20260811b bound it). deepseek/deepseek-v4-pro is not
used: it failed step 1's manipulation check at 0.692 and quoted strings absent from the passage in
4 of 80 cues, and it is note (bmb)'s repeat offender on finish_reason: length. temperature
0, max_tokens 250, one licensed re-dispatch at 800 for any body returning length.
Judgment is not parallelised across seats within a pair; the lead judges nothing (charter §5).
5. Predictions, registered here and in the frozen translator's log (D10 Q1, Q2)
Rates are over decided judgments (NEITHER excluded), pooled across passages, orders and
eligible seats.
A4 — the inferential unit is the PASSAGE, not the call. Two orders of one passage, three
fixed seats at temperature 0, and six passages from one hand are not independent observations. Every
primary is decided on the per-passage rate, by an exact one-sided sign test over the eight
passages (minimum attainable P = 1/256 = 0.0039; 7 of 8 gives 0.035) plus a cluster bootstrap over
passages. Pooled call-level rates are reported as descriptive only and as overstating precision.
A8 — the primary reading is cue-attributed. A judgment counts toward a causal claim only if
its cited cue lies inside a manipulated span and differs between the two forms shown. The
all-judgments reading is the sensitivity analysis, and both are printed. Note (bmh): a
same/different verdict must be read, not counted.
A10 — minimum denominator 24 decided judgments for any equivalence verdict; below it the
verdict is INCONCLUSIVE, which is neither a pass nor evidence against.
H1—C1picksDOMoverNEUonBRITin ≥ 7 of 8 passages, and the pooled decided rate is ≥ 0.75 (descriptive).H2(the equivalence half, andH1is not reportable without it) —C2's rate forTRAhas a 95% Clopper–Pearson interval inside [0.35, 0.65].H3—C3's rate forSHAMinside [0.35, 0.65].H4(positive control) —C4picksTRAonFOREIGNin ≥ 7 of 8 passages.H5—C5's rate forDOMinside [0.35, 0.65].H7(THE PRIMARY, added asA3on critic finding F3) —C6,DOMvsTRAdirectly onBRIT, picksDOMin ≥ 7 of 8 passages. This is the asymmetry itself, and the conclusion is tied to it.H1andH2remain, but aC6null withholds the finding whateverC1andC2do.H6(dose, fromQ2) — DESCRIPTIVE ONLY, dropped from the confirmatory set asA9(critic F9):STRONGcount is confounded with total sites, passage length, hand, work, language and the held-constant St George material. The per-passage table and the Spearman ρ are reported with those confounds named, and no threshold is registered.
The asymmetry is the finding, and A2 narrows what may be said about it. H1 alone is
near-tautological: the DOM arm contains Michaelmas, sovereign, Whitehall, and the question
asks whether the prose sounds British. What is not tautological is H7: that a foreign word,
equally conspicuous and equally a substitution, is not read as evidence about the English while
a target-culture word is. The reportable claim is therefore an effect of overt target-culture
substitution, and the craft consequence is that the domesticating road necessarily reaches for
target-culture words, which are then read as evidence about the prose. If H1 passes and H2
fails high, the result is "any culture-marked substitution moves it" and the domestication claim is
withheld.
A1 — and the claim is never that the English moved while the world stayed put. Substituting
the target culture's institution is domestication; Michaelmas for Đurđevdan changes the
furniture and the English together. C5 checks only that the narrative setting — place, proper
names, events — did not move, and that is all it is said to check.
A6 — the archaism discriminator, on critic finding F6. C3 is not matched to DOM for
archaism or institutional colour, and a matched control cannot be built here because most archaic
English institutional vocabulary is nation-marked. Instead, every cited cue is classified against
the frozen brit flag. If DOM cues concentrate on WEAK sites (hogshead, firkin,
pelisse — old and institutional but not national) rather than STRONG ones (sovereign,
Michaelmas, Whitehall), the rival explanation "archaic institutional diction reads British" is
live and the domestication claim is withheld.
6. Gates, and what each withholds
| gate | bar | withholds if failed |
|---|---|---|
G1 returns |
100% of 240 bodies | nothing; void cells reported, never imputed |
G2 position preference |
no seat picks the A slot on > 0.70 of its decided judgments | that seat, from all primaries |
G3 cue verbatim |
≥ 0.90 of a seat's quoted cues occur in the passage shown | that seat, from all primaries |
G4 build |
a forms byte-identical outside declared sites · b no British spelling in any form · c all four forms distinct |
the run |
G5 duplicate stability — a REPORT, not a gate (A8) |
on 6 byte-identical duplicate pairs, a seat repeats its answer on ≥ 4 | nothing. Byte-identical repeats at temperature 0 are a weak reliability test and are reported as such |
G6 NEITHER rate |
reported per seat and per contrast | nothing; a seat at > 0.60 contributes little and is said to |
G7 read, not counted — PRIMARY (A8) |
a judgment counts toward a causal claim only if its cited cue lies inside a manipulated span and differs between the two forms | the counted reading, which becomes the sensitivity analysis. Note (bmh), S161: 2 of 24 arrivals were paraphrase drift, visible only in free text |
H4 is the gate on H2. If the seats cannot see the foreign loading even when asked about the
world, then C2's null means "these seats notice nothing", not "foreign words do not place the
English", and H2 is withheld.
7. Cost
Pre-flight, built from max_tokens and the exact input length of every pair (note (abc): price
the worst case from the cap the request permits, never from an assumed output length). Printed by
run.py --dry-run and recorded in the result page against the billed actual.
Declared ceiling: $1.15, inside a UTC-day headroom of $1.212182706 at session start.
Prices from config/models.md; the arithmetic is printed by run.py --dry-run and is:
| line | worst case |
|---|---|
study, 64 pairs × 3 seats, max_tokens 450 (incl. C6) |
$0.747534 |
G5 duplicates, 6 pairs × 3 seats |
$0.070087 |
| re-dispatch contingency, 20 length-deaths at 800 | $0.129000 |
pre-run critic, openai/gpt-5.6-terra, spent |
$0.031341 |
| already spent — two dead critic bodies, see below | $0.153746 |
| total | $1.131721 |
The design was scaled down mid-build and this is why. The pre-run critic was first sent to
moonshotai/kimi-k3, chosen because it is not one of the three judging seats. It returned
finish_reason: length with null content and 2,497 reasoning tokens at a 2,500 cap, and again
with null content and 5,997 reasoning tokens at a 6,000 cap — $0.053388 + $0.100358 =
$0.153746 spent for no body. This is note (bmb)'s failure mode on a third slug and note
(abc)'s arithmetic for the sixth session running. Two consequences, both taken before dispatch:
- The judges'
max_tokenswas raised from 250 to 450, becausegoogle/gemini-3.6-flashis also a hidden-reasoning seat and 250 would have exposed 62 of its calls to the same death. C3,C4,C5were cut from two orders to one order each, balanced by passage, to pay for (1) and for the burned $0.153746 inside the day's headroom.
And the critic seat changed to openai/gpt-5.6-terra, which IS one of the three judging seats.
That is a real weakening of the critic's independence from the jury and is declared here rather
than in the result: the only two panel seats outside this jury are moonshotai/kimi-k3, which
would not return a body twice, and deepseek/deepseek-v4-pro, which is excluded from the jury for
exactly the same failure mode. The critic does not judge any item; what is lost is the guarantee
that no model both criticised the design and answered under it.
8. Limits, registered before the run
- A forced pairwise choice is a weaker instrument than an absolute verdict, and a
C1pass licenses only these seats can tell which of two renderings reads more British, not a reader's placement of a translation changes. Step 1 measured the stronger claim and found nothing. DOMis the ceiling of domestication, not typical practice — every culture-bound item in the passage at once. A pass at the ceiling licenses nothing about a translator who domesticates four sites in forty.- Six of eight passages are one hand, one work, one language. They are not eight independent
observations.
AandBare reported separately as well as pooled, and the per-passage table is the honest unit. DOMrenderings were written by the lead, in the frozen log, and no second annotator scored their Britishness. Step 1 asked for two blinded annotators with a codebook and could not get them; neither can this.- The estimand is an instructed task — can these three seats separate the channels when told
to — inherited unchanged from step 1's amendment
A18. Nothing here is a claim about human readers, and the word reader does not appear in any reported claim. - Contamination is not measurable on the Serbian (no reachable English rendering). The design is within-passage, so anything recalled sits in all four arms equally and cannot produce a difference between them; the declaration is still an assertion and says so on the artifact.
- St George is in this story twice — the icon, and Đurđevdan as the hiring day — and he is
England's patron. He is held constant in every arm and so cannot create a difference, but he
raises the floor on
W3andW6. W6is 347 words carrying oneSTRONGsite. It is kept deliberately, as the low end of the dose gradientH6predicts on, and it is the passage most likely to returnNEITHER.
9. Failure criteria
G4fails → the run does not dispatch.H4fails →H2and therefore the asymmetry claim are withheld; only descriptive rates are reported.H7(C6) fails → the finding is withheld whateverH1andH2do (A3).DOMcues concentrate onWEAKsites → the finding is withheld (A6).- fewer than 24 decided judgments in a contrast → that contrast's verdict is INCONCLUSIVE (
A10). G2orG3fails on a seat → that seat is excluded from every primary and the primaries are reported with and without it, as step 1 reportedL2.- Any primary reaching its bar only after excluding a seat is reported as withheld, with both readings printed.
- If
C1andC2are both high and within 0.10 of each other, the finding is "any culture-marked substitution" and the domestication claim is not made.