Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260811h-domestication-channel/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260811h-domestication-channel
statusfrozen
created2026-08-11
updated2026-08-11
sensescultural-mediation, perceived-source-carriage
internal-judgment-onlytrue
provisionaltrue
linkswiki/arms/ARM-realia-channel.md, wiki/findings/results/RS-20260811b-realia-channel.md, workshop/translations/jutrenje/R06-v1/translation.md, workshop/experiments/E-20260811b-realia-channel/materials/realia.json, wiki/goodness-senses.md, framework/v0.2/README.md, config/models.md, config/budget.md

E-20260811h — which road out of a culture-bound word places the English

ARM-realia-channel step 2, the arm's closing step. Frozen before any API call.

1. The question, and why it is not step 1's question

RS-20260811b (step 1) deleted the source-culture world from seven translated passages and asked three seats which national variety of English the prose was written in. The verdict did not move: +0.048, P = 0.3125, and the pre-run critic's equivalence margin was not met either, so the run concluded no evidence of leakage, and not enough evidence of separability.

But it found something it had not designed for. Of the twelve realia cues the seats gave for the prose question, every one was one of three words — halfpennies, smock, councillor — and every one is the translator's own Anglicisation of a foreign thing. The genuinely foreign words in the same corpus — taiga, yamen, Sanzu-no-Kawa, the Dragon Throne — were cited 118 times for where is this set and not once for whose English is this.

That is a conjecture with a predictor in it, and it is about the oldest decision in translation:

A culture-bound item has three roads out of it — carry the source word over, substitute the nearest thing in the target's own world, or generalise to a location-free phrase. The conjecture is that only the second road moves where a reader places the English itself, and that the first road, which is every bit as conspicuously foreign, does not.

If that holds, a translator who domesticates in order to make a translation read naturally in English is not producing neutral English. They are producing English that reads as belonging to a particular country — a cost wiki/goodness-senses.md's cultural-mediation does not name, and one framework/v0.2 §7 currently has no measurement of.

What this unit teaches about translating literature (the subject rule, wiki/tracks.md): which of the three standard treatments of a culture-bound item changes a reader's placement of the translator's English, as against their placement of the story's world. It is a question about a translator's choice, not about this project's instruments.

2. Three changes from step 1, each forced by something step 1 found

  1. Three arms, not two. Step 1's KEEP conflated transfer with domestication — Garnett's halfpennies sat in the same arm as Shaw's Sanzu-no-Kawa. TRA / DOM / NEU separates them, and the whole finding lives in the contrast between the first two.
  2. Proper names are held constant and excluded from the manipulation. Step 1's spans included St Petersburg, Yakov, the great Lena, Yamashiro-Ya, so its MUTE arm could be read as relocating the story rather than generalising its furniture. Here nothing moves the story out of Serbia or Russia, and C5 checks that.
  3. Forced pairwise comparison, not an absolute verdict. Step 1's §9.1 is explicit that at n = 7 passages only an effect consistent in 6 of 7 could have reached P = 0.05. A within-pair forced choice removes between-item variance and yields up to 48 binary judgments per contrast. This is a weaker claim than step 1's and is registered as such (§8, limit 1): it measures whether two renderings are distinguishable in a direction, not whether a verdict changes. Step 1 already supplies the absolute-verdict null.

3. Materials

Translation limb: T-jutrenje-R06-v1 — Laza Lazarević, «Први пут с оцем на јутрење» (1879), §I and §II, 1,519 Serbian words → 1,948 English, frozen at commit 287916d before this design existed. The project's twentieth source language and its first South Slavic one. Contamination not measurable — no English rendering of the work is reachable — and declared as an assertion on the artifact; the within-passage design does not turn on it (§8, limit 6).

The wire between the limbs, in one sentence. The translator's R06 log records, at each culture-bound site, the two renderings the translator did not choose — the source word carried over, and the nearest English-domestic article — and the study limb puts all three to independent readers to find out which of them changes where they place the English itself.

Eight passages, four forms each, built by code.py, which is deterministic and makes no API call:

passage hand source lang words sites STRONG sites
W1 lead R06 Serbian 148 16 4
W2 lead R06 Serbian 144 6 5
W3 lead R06 Serbian 116 7 6
W4 lead R06 Serbian 259 5 3
W5 lead R06 Serbian 221 5 3
W6 lead R06 Serbian 347 3 1
A Garnett Russian 118 3 3
B Field Russian 128 4 3

A and B are E-20260811b's passages A and B in their US-spelled base form, with new DOM renderings added to spans that experiment had already declared. A carries halfpennies itself — the word that generated the conjecture is in the corpus, in the hand that wrote it.

The four forms. TRA (source word carried over) · DOM (nearest article of the English domestic world) · NEU (location-free English phrase) · SHAM (NEU with a location-neutral-to-location-neutral substitution at every site, matched site for site).

brit is frozen in the translator's log, not assigned later. STRONG = the domesticating rendering names a specifically British thing (sovereign, shilling, firkin, hogshead, Michaelmas, pitch-and-toss, indentures, counterpane, rushlight, billycock, on tick, Civil Service, Whitehall). WEAK = English-domestic with no particular nation in it. 28 of 49 sites are STRONG.

Orthography is forced to AMERICAN in every form of every passage by one mechanical map applied identically to all four (code.py SPELLING). No "reads more British" verdict in this run can be produced by a spelling; G4b asserts no British spelling survives anywhere.

4. Procedure

Each pair shows two forms of the same passage, labelled A and B, in both orders. Two questions, each elicited on its own call:

A7 — both prompts are frozen verbatim here, and one pair gets one question on one call. The wrapper is run.py TEMPLATE; {question} is one of exactly these two strings:

Each is followed by "Answer NEITHER if there is no difference." and by the required one-line JSON shape {"answer": "A"|"B"|"NEITHER", "cue": "<quoted exactly from the passage>", "confidence": 0-3}, then the two passages under --- A --- and --- B ---.

A5 — every asterisk is stripped from every form (code.py deitalic(), critic finding F5). Markdown italics marked the TRA arm and only the TRA arm, so a seat could have read the source words as quotation rather than as prose. No arm is now typographically marked.

Six contrasts, all eight passages, three seats = 64 pairs, 192 calls, plus 6 duplicate pairs (18 calls) for G5. C1 and C2 carry the claim and run in both orders on all eight passages (16 pairs each). C3, C4, C5, C6 run in one order each, alternating by passage so that order is balanced 4/4 inside every control contrast at half the calls. This is a scale-down from both-orders-everywhere, forced by the two dead critic bodies in §7, and it is declared rather than absorbed: the controls are order-balanced but not order-paired, so an order effect large enough to move a control by more than the 4/4 balance absorbs would not be visible inside that contrast. G2 measures position preference across all five contrasts and is what would catch it.

id comparison question what it is for
C1 DOM vs NEU BRIT the primary
C2 TRA vs NEU BRIT the equivalence half — a foreign word is equally marked and should not move this
C3 SHAM vs NEU BRIT edit-presence control
C4 TRA vs NEU FOREIGN positive control — the seats must be able to see the foreign loading when asked about the world
C5 DOM vs NEU FOREIGN confound check — domesticating must not move the narrative setting
C6 DOM vs TRA BRIT the asymmetry, tested directly (A3, critic F3)

Seats. P1 openai/gpt-5.6-terra, P3 x-ai/grok-4.5, P2 google/gemini-3.6-flash (reasoning: {effort: low}, as E-20260811b bound it). deepseek/deepseek-v4-pro is not used: it failed step 1's manipulation check at 0.692 and quoted strings absent from the passage in 4 of 80 cues, and it is note (bmb)'s repeat offender on finish_reason: length. temperature 0, max_tokens 250, one licensed re-dispatch at 800 for any body returning length.

Judgment is not parallelised across seats within a pair; the lead judges nothing (charter §5).

5. Predictions, registered here and in the frozen translator's log (D10 Q1, Q2)

Rates are over decided judgments (NEITHER excluded), pooled across passages, orders and eligible seats.

A4 — the inferential unit is the PASSAGE, not the call. Two orders of one passage, three fixed seats at temperature 0, and six passages from one hand are not independent observations. Every primary is decided on the per-passage rate, by an exact one-sided sign test over the eight passages (minimum attainable P = 1/256 = 0.0039; 7 of 8 gives 0.035) plus a cluster bootstrap over passages. Pooled call-level rates are reported as descriptive only and as overstating precision.

A8 — the primary reading is cue-attributed. A judgment counts toward a causal claim only if its cited cue lies inside a manipulated span and differs between the two forms shown. The all-judgments reading is the sensitivity analysis, and both are printed. Note (bmh): a same/different verdict must be read, not counted.

A10 — minimum denominator 24 decided judgments for any equivalence verdict; below it the verdict is INCONCLUSIVE, which is neither a pass nor evidence against.

The asymmetry is the finding, and A2 narrows what may be said about it. H1 alone is near-tautological: the DOM arm contains Michaelmas, sovereign, Whitehall, and the question asks whether the prose sounds British. What is not tautological is H7: that a foreign word, equally conspicuous and equally a substitution, is not read as evidence about the English while a target-culture word is. The reportable claim is therefore an effect of overt target-culture substitution, and the craft consequence is that the domesticating road necessarily reaches for target-culture words, which are then read as evidence about the prose. If H1 passes and H2 fails high, the result is "any culture-marked substitution moves it" and the domestication claim is withheld.

A1 — and the claim is never that the English moved while the world stayed put. Substituting the target culture's institution is domestication; Michaelmas for Đurđevdan changes the furniture and the English together. C5 checks only that the narrative setting — place, proper names, events — did not move, and that is all it is said to check.

A6 — the archaism discriminator, on critic finding F6. C3 is not matched to DOM for archaism or institutional colour, and a matched control cannot be built here because most archaic English institutional vocabulary is nation-marked. Instead, every cited cue is classified against the frozen brit flag. If DOM cues concentrate on WEAK sites (hogshead, firkin, pelisse — old and institutional but not national) rather than STRONG ones (sovereign, Michaelmas, Whitehall), the rival explanation "archaic institutional diction reads British" is live and the domestication claim is withheld.

6. Gates, and what each withholds

gate bar withholds if failed
G1 returns 100% of 240 bodies nothing; void cells reported, never imputed
G2 position preference no seat picks the A slot on > 0.70 of its decided judgments that seat, from all primaries
G3 cue verbatim ≥ 0.90 of a seat's quoted cues occur in the passage shown that seat, from all primaries
G4 build a forms byte-identical outside declared sites · b no British spelling in any form · c all four forms distinct the run
G5 duplicate stability — a REPORT, not a gate (A8) on 6 byte-identical duplicate pairs, a seat repeats its answer on ≥ 4 nothing. Byte-identical repeats at temperature 0 are a weak reliability test and are reported as such
G6 NEITHER rate reported per seat and per contrast nothing; a seat at > 0.60 contributes little and is said to
G7 read, not counted — PRIMARY (A8) a judgment counts toward a causal claim only if its cited cue lies inside a manipulated span and differs between the two forms the counted reading, which becomes the sensitivity analysis. Note (bmh), S161: 2 of 24 arrivals were paraphrase drift, visible only in free text

H4 is the gate on H2. If the seats cannot see the foreign loading even when asked about the world, then C2's null means "these seats notice nothing", not "foreign words do not place the English", and H2 is withheld.

7. Cost

Pre-flight, built from max_tokens and the exact input length of every pair (note (abc): price the worst case from the cap the request permits, never from an assumed output length). Printed by run.py --dry-run and recorded in the result page against the billed actual.

Declared ceiling: $1.15, inside a UTC-day headroom of $1.212182706 at session start. Prices from config/models.md; the arithmetic is printed by run.py --dry-run and is:

line worst case
study, 64 pairs × 3 seats, max_tokens 450 (incl. C6) $0.747534
G5 duplicates, 6 pairs × 3 seats $0.070087
re-dispatch contingency, 20 length-deaths at 800 $0.129000
pre-run critic, openai/gpt-5.6-terra, spent $0.031341
already spent — two dead critic bodies, see below $0.153746
total $1.131721

The design was scaled down mid-build and this is why. The pre-run critic was first sent to moonshotai/kimi-k3, chosen because it is not one of the three judging seats. It returned finish_reason: length with null content and 2,497 reasoning tokens at a 2,500 cap, and again with null content and 5,997 reasoning tokens at a 6,000 cap — $0.053388 + $0.100358 = $0.153746 spent for no body. This is note (bmb)'s failure mode on a third slug and note (abc)'s arithmetic for the sixth session running. Two consequences, both taken before dispatch:

  1. The judges' max_tokens was raised from 250 to 450, because google/gemini-3.6-flash is also a hidden-reasoning seat and 250 would have exposed 62 of its calls to the same death.
  2. C3, C4, C5 were cut from two orders to one order each, balanced by passage, to pay for (1) and for the burned $0.153746 inside the day's headroom.

And the critic seat changed to openai/gpt-5.6-terra, which IS one of the three judging seats. That is a real weakening of the critic's independence from the jury and is declared here rather than in the result: the only two panel seats outside this jury are moonshotai/kimi-k3, which would not return a body twice, and deepseek/deepseek-v4-pro, which is excluded from the jury for exactly the same failure mode. The critic does not judge any item; what is lost is the guarantee that no model both criticised the design and answered under it.

8. Limits, registered before the run

  1. A forced pairwise choice is a weaker instrument than an absolute verdict, and a C1 pass licenses only these seats can tell which of two renderings reads more British, not a reader's placement of a translation changes. Step 1 measured the stronger claim and found nothing.
  2. DOM is the ceiling of domestication, not typical practice — every culture-bound item in the passage at once. A pass at the ceiling licenses nothing about a translator who domesticates four sites in forty.
  3. Six of eight passages are one hand, one work, one language. They are not eight independent observations. A and B are reported separately as well as pooled, and the per-passage table is the honest unit.
  4. DOM renderings were written by the lead, in the frozen log, and no second annotator scored their Britishness. Step 1 asked for two blinded annotators with a codebook and could not get them; neither can this.
  5. The estimand is an instructed task — can these three seats separate the channels when told to — inherited unchanged from step 1's amendment A18. Nothing here is a claim about human readers, and the word reader does not appear in any reported claim.
  6. Contamination is not measurable on the Serbian (no reachable English rendering). The design is within-passage, so anything recalled sits in all four arms equally and cannot produce a difference between them; the declaration is still an assertion and says so on the artifact.
  7. St George is in this story twice — the icon, and Đurđevdan as the hiring day — and he is England's patron. He is held constant in every arm and so cannot create a difference, but he raises the floor on W3 and W6.
  8. W6 is 347 words carrying one STRONG site. It is kept deliberately, as the low end of the dose gradient H6 predicts on, and it is the passage most likely to return NEITHER.

9. Failure criteria