Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260813a-world-dose/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260813a-world-dose
statusfrozen
created2026-08-13
updated2026-08-13
sensescultural-mediation, perceived-source-carriage
internal-judgment-onlytrue
provisionaltrue
linkswiki/arms/ARM-dose.md, wiki/findings/results/RS-20260812e-dose.md, workshop/experiments/E-20260812e-dose/design.md, workshop/translations/stachomazochtra/R06-v1/translation.md, workshop/translations/stachomazochtra/contamination.md, wiki/goodness-senses.md, framework/v0.2/README.md, config/models.md, config/budget.md

E-20260813a-world-dose — does keeping some of the source words keep any of the world?

ARM-dose step 2, the step that closes the arm. Frozen before dispatch. code.py is frozen with it and is part of the design.

1. The question, and why it is the one contrast left

RS-20260812e laddered domestication rung by rung and found the prose channel graded at every rung: one Anglicised item beats none (28/28), half beats one (30/30), all beats half (30/30), all beats one (60/60). On the world channel — which version is more clearly set outside the English-speaking world? — it ran only two contrasts, K7 (D1 vs D0) and K8 (DA vs D0), and both came back 0 of 30: the untouched page is chosen as more clearly foreign against the one-item page just as completely as against the fully domesticated one.

Its own §6.3 states the defect in terms:

"DA vs D1 on the world question is the one contrast this design should have carried and does not … so this run establishes that the world cost is present and complete against the untouched page at one item, and not that it stops growing."

framework/v0.2 §7.9 was written from that gap and currently tells a translator "there is no measured dose at which the setting is left alone." That sentence is true and its natural reading is not. A reader takes it to mean the setting is an all-or-nothing thing, spent at the first Anglicised word — which would make partial domestication pointless on this channel and would be a real piece of advice. The run that produced it cannot distinguish that from the opposite: that the world, like the prose, is a slope, and every item you keep keeps some of it.

And there is a second question the prose channel has an answer to and the world channel does not. K4 established that the dose-1 prose effect is specific to the English domestic article and not to the missing foreign word: at one site, D1 (guineas) is read as more British than T1 (gold pieces) at 56 of 58. The same contrast on the world question has never been run, and it is the one a translator actually faces. Faced with zlaté, three roads: carry it over, Anglicise it, or find a location-free phrase. wiki/goodness-senses.md names all three under cultural-mediation. If the world cost at one site is caused by losing the source word, the third road buys nothing and is a false comfort. If it is caused by gaining an English one, the third road is real.

Two primaries, therefore, and the second is the one whose answer is not guessable:

  1. K9 — does the world cost grow with dose? DA against D1, world question.
  2. K10 — is the world cost at one item specific to the English domestic article? D1 against T1, world question, same site, minimal pair.

The unit's one sentence (continue-prompt.md §4.5): it teaches what a translator's story's setting costs, item by item, and whether reaching for a location-free phrase instead of an English one protects it. That is about translating, not about the project's apparatus.

2. Materials

2.1 The ten frozen windows — the primary set

Loaded verbatim from E-20260812e-dose/materials/items.json, not rebuilt. All fifty forms of the ten windows (H1–H4, W1–W5, B) are read out of that frozen file byte for byte, so every text this run shows a seat on those windows is the identical string E-20260812e showed, including the seeded choice of which site is dose 1. code.py asserts the identity by SHA over each form.

Rebuilding them was rejected for a specific reason: E-20260812e's code.py draws all its permutations from one random.Random(20260812) stream consumed in sorted window-name order, so inserting new windows whose names sort before W would silently change W1–W5's permutations and K11 would stop being a replication of K8. Loading the frozen artifact removes the possibility.

win hand source lang words manipulable sites dose ladder (1 / h / N)
H1–H4 lead Czech 198–311 7 / 8 / 9 / 7 as E-20260812e
W1–W5 lead Serbian 116–259 16 / 6 / 7 / 5 / 5 as E-20260812e
B Field (published) Russian 128 4 as E-20260812e

2.2 The five new windows — a replication set, reported separately

Cut from T-stachomazochtra-R06-v1 — Alexandros Papadiamantis, «Ἡ Σταχομαζώχτρα» (1889), the whole story, Modern Greek → English, the project's first modern-Greek source — translated by the lead this session under R06 and frozen at commit 8698001 before this design was written, with its 40-row culture-bound site table written in the translator's log as the translation was written.

win what words manipulable sites dose ladder (1 / h / N)
S1 the neighbour's astonishment 225 6 1 / 3 / 6
S2 the trades 298 9 1 / 5 / 9
S3 the cold house 340 7 1 / 4 / 7
S4 the priest at the door 211 7 1 / 4 / 7
S5 the shop 114 9 1 / 5 / 9

Built by the same code path, the same five forms, the same MIN_SITES = 4 rule, from a separate RNG stream, random.Random(20260813), consumed in sorted order over S1–S5 only.

Why they are a replication set and not part of the primary. ARM-dose step 2 was written down before this translation existed, so the lead knew, while making the Greek site table, that a world-channel dose contrast was coming. T-hastrman-R06-v1 was not in that position. The Greek windows are therefore analysed separately and never pooled into a primary verdict; a pooled fifteen-window figure is reported as descriptive only. This is declared on the translation artifact in the same words.

Contamination. T-stachomazochtra-R06-v1 is contamination: suspected, measured: workshop/translations/stachomazochtra/contamination.md records 91 / 14 / 5 shared 7-, 12- and 15-grams and a 17-token longest common run against Elizabeth Key Fowden's 2007 rendering, verdict DEPENDENT? — and 0 shared 12-grams inside all five windows, longest window run 10 tokens. This design does not turn on independence: every form compared is an edited copy of the same text, so what that text shares with Fowden it shares in both arms of every pair.

2.3 The five forms

Unchanged from E-20260812e §2, and the new windows are built by the same functions:

form dose build
D0 0 every site TRA — the source word carried over
D1 1 perm[0] → DOM, every other site TRA
DH ⌈N/2⌉ perm[:h] → DOM, the rest TRA
DA N every site DOM
T1 — perm[0] → NEU, every other site TRA — the minimal-pair control

Orthography is forced to American in every form by the same mechanical map, applied identically to all five, so that no verdict can be produced by a spelling. The Greek translation is written in British spelling and the map covers it.

3. Procedure

Forced pairwise comparison, prompt template and question strings byte-identical to E-20260811h's and E-20260812e's — that identity is what makes K11 a replication rather than a similar-looking new measurement. A seat sees two versions of one window labelled A and B, is told they differ only in a handful of words, is asked one question, and returns one line of JSON: the answer, the single word or phrase that decided it quoted exactly, and a confidence 0–3. NEITHER is available.

id high-dose form low-dose form question windows orders role
K9 DA D1 FOREIGN all 15 both PRIMARY — does the world cost grow with dose?
K10 D1 T1 FOREIGN all 15 both PRIMARY — is it the domestic article or the missing source word?
K11 DA D0 FOREIGN all 15 one the gate — must replicate RS-20260812e K8
K12 DH D1 FOREIGN 10 frozen one the middle rung, direction only
K13 DA DH FOREIGN 10 frozen one the top rung, direction only
K14 DA D0 BRIT 5 Greek one the new material's gate — replicates K3 in a third language pair

Rate always means the fraction of decided, cue-attributed judgments that chose the more-domesticated form, exactly as in RS-20260812e, on both questions. On FOREIGN a rate below 0.5 therefore means less domestication read as more foreign, which is the expected direction; K7 and K8 were 0.000 under this convention.

Cue attribution is the primary reading (E-20260811h A8): a judgment counts toward a causal claim only if its quoted cue lies inside a site that differs between the two forms shown. Counted (unattributed) rates are reported alongside and are descriptive.

The inferential unit is the window (E-20260811h A4). Pooled rates with Clopper–Pearson intervals are descriptive and overstate precision; the sign test over windows is the inference.

Seats. P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash (reasoning: {effort: low}), P3 x-ai/grok-4.5 at default effort, no reasoning parameter — the configuration RS-20260812e §7.1 collapsed to after its dispatch race, so this run is comparable to it by construction. Temperature 0. Judgment is not parallelised across seats within a pair; the lead judges nothing (charter §5), and none of the fifteen windows is any seat's own output.

The pre-run critic is qwen/qwen3.7-max, which is NOT one of the three judging seats. moonshotai/kimi-k3 is not asked, per note (bhf) rule (iii) and note (bmb)'s cost finding.

4. Predictions and verdicts, registered

G — the gate. K11 must reproduce RS-20260812e K8 on the ten frozen windows: D0 chosen as more clearly foreign in ≥ 9 of 10 frozen windows (equivalently, DA chosen in ≤ 1). If K11 fails, every other verdict in this run is withheld, because a world-channel manipulation that does not work at full dose cannot be read at partial dose.

P6 (K9) — THE FIRST PRIMARY, an equivalence test, on the ten frozen windows.

P7 (K10) — THE SECOND PRIMARY, also an equivalence test, on the ten frozen windows.

Note that P6 and P7 are not symmetric in interest. P6's GRADED outcome is the one the picture predicts; its value is that SPENT was reachable and registered, and that the framework's current sentence is ambiguous between them. P7 has no predicted direction — the prose-channel answer (K4, specific at 0.966) does not transfer, because the two questions are about different things, and the run is designed so that a null on P7 is as publishable as a hit.

P8 (K12, K13) — the middle and top rungs, direction only. Reported as rates and window counts on the ten frozen windows. No equivalence verdict, by the minimum-denominator rule below. Registered reading: if P6 returns GRADED, both should sit below 0.5; if either sits at or above 0.5 while K9 is below 0.35, the world ladder is not monotone, which is a finding against the slope story and is reported as one.

P9 (K14) — the new material's gate. DA chosen on BRIT in ≥ 4 of 5 Greek windows. PASS → RS-20260812e's core prose result replicates in a third language pair, and the Greek windows may be carried into the descriptive pooled figures. FAIL → the Greek windows are reported in full and excluded from every pooled figure, and the failure is the finding.

The replication set, registered in advance. K9, K10 and K11 on the five Greek windows are computed and reported separately, with their own rates, intervals and window counts. Agreement is not required for the primaries to stand and disagreement does not overturn them; a divergence is reported as a divergence and referred to the fact that the Greek site table was made with this contrast known.

Minimum denominator, computed before dispatch, per note (bhr). An equivalence verdict requires ≥ 30 decided, cue-attributed judgments. K9 and K10 each run both orders on ten frozen windows across three seats = 60 dispatches each, so a 50% loss still clears the bar. K11–K14 run one order and cannot support an equivalence verdict at any loss rate, so none is registered on them and none may be given afterwards.

Secondary, descriptive only: K9 and K10 split by the lead site's frozen STRONG/WEAK flag. Two WEAK windows out of ten cannot support a verdict and none is registered.

5. Gates

gate bar if it fails
G1 returns every pair dispatched; empty and truncated bodies counted and reported, never imputed reported in the result's limits
G2 position preference ≤ 0.70 on the A slot, per seat, over the whole run that seat is dropped from primaries; both figures printed
G3 cue verbatim ≥ 0.90 of quoted cues occur verbatim in one of the two texts shown, per seat, over decided judgments (E-20260812e amendment A1) that seat is dropped from primaries; both figures printed
G4a–e build byte-identity outside sites · no British spelling in any form · five forms pairwise distinct · D1/T1 differ at exactly one site · ladder strictly nested run in the build, on the five new windows; the ten frozen ones are asserted byte-identical to E-20260812e's by SHA instead
G5 duplicates a seeded 8-pair subsample re-dispatched byte-identically a report, not a gate
G6 NEITHER reported per seat —

Note (blf) is a rule of this run, not a hope. The dispatch script is launched detached from the first attempt, never in the foreground. RS-20260812e lost 30.6% of its spend to two racing foreground-and-restarted dispatchers and RS-20260812i lost $0.062 to a foreground dispatch killed by the harness's 120-second tool timeout after it had billed. run.py writes one body per file, refuses to start if another instance holds its lock file, and never overwrites an existing body.

6. What this cannot establish

  1. Nothing about human readers. Three LLM seats, an instructed task, temperature 0. The word reader is not used in any claim.
  2. No goodness sense is scored. Tier D is NOT PASSED; every sentence is provisional.
  3. "More clearly set outside the English-speaking world" is one question, asked one way. It is not a measurement of where a reader thinks the story is set; it is a forced comparison between two pages that differ in a handful of words.
  4. K11 is near-tautological in the way K3 was, and is used only as a gate.
  5. Which site is dose 1 is a seed's choice, fifteen times over.
  6. The STRONG/WEAK flags are one annotator's, in all three frozen logs, with no second coder.
  7. The Greek site table was built with this contrast known. §2.2. It is why those five windows are not in any primary.
  8. DA and D1 differ by more than a count on some windows. DA also removes the internal inconsistency of a page that transfers eight items and Anglicises one. Whatever a seat is responding to when it prefers D1 as more foreign, this design cannot separate fewer English words from a more mixed page.

7. Cost — pre-flight

Worst case is built from the caps run.py actually sends, and from the one place a cap is known not to bind (note (abc), third mode). RS-20260812e §7.1 measured x-ai/grok-4.5 returning a mean of about 1,400 completion tokens against a max_tokens of 500, because on xAI that parameter does not bound reasoning tokens. P3 is therefore priced at 2,000 completion tokens, not at its 500 cap. P1 and P2 are priced at their caps (350 / 1,000), which they have been observed to respect.

Pairs: K9 15×2 + K10 15×2 + K11 15 + K12 10 + K13 10 + K14 5 = 100 pairs × 3 seats = 300 bodies, plus G5 8 duplicate pairs × 3 = 24.

python3 run.py --dry-run prints the exact figure from the frozen items.json; the table is filled in from that output before dispatch and is reproduced in the result page.

Declared ceiling: $1.80, inside a UTC-day headroom of $5.00 — 2026-08-13 is a fresh day and this is its first session. If the dry run prices above the ceiling, K13 is dropped first and K12 second, and the drop is recorded here before dispatch rather than in the analysis.

7.1 Amendment A1 — the pre-committed drop fired, and the ceiling is raised with it

Written before dispatch, after the dry run and before any judging call.

The dry run priced the full six-contrast design at a worst case of $2.921030, above the $1.80 ceiling. §7's own rule was applied, in the order it names:

  1. K13 (DA vs DH, FOREIGN) dropped. 10 pairs.
  2. K12 (DH vs D1, FOREIGN) dropped. 10 pairs.
  3. G5 duplicates reduced from 8 pairs to 4 — one per surviving contrast, which is what the subsample was for.

What the drop costs, stated now rather than discovered later: this run can order the world ladder between dose 1 and dose N and cannot order its intermediate rungs. If P6 returns GRADED, the claim available is more Anglicised items read as less clearly foreign than fewer, not the world ladder is monotone at every rung. The prose channel has the rung-by-rung result (RS-20260812e K1, K5, K6); the world channel will not, and framework/v0.2 may not borrow it across.

The reduced design prices at $2.299413 worst case — still above $1.80, and the ceiling is raised to $2.30, with the reason written here before dispatch. The reason is not that the run grew: it did not, it shrank by 20 pairs. It is that this session prices P3 more conservatively than any previous one. RS-20260812e §7.1 measured x-ai/grok-4.5 returning a mean of about 1,400 completion tokens against a max_tokens of 500, because on xAI that parameter does not bound reasoning; every earlier estimate in this project priced that seat at its cap and therefore under-priced it — that is note (abc)'s third mode. Pricing it at 2,000 completion tokens adds $1.0739 to the worst case on its own, and is the correct number. The expected actual is near $1.2, by the ratio RS-20260812e actually realised (324 bodies for $1.0013 per-response at comparable prompt lengths); 252 calls here against 324 there.

$2.30 is 46% of a fresh UTC day's $5.00. It is declared here as a hard stop: if the running per-response total passes it, dispatch halts and the analysis runs on whatever is complete, with the shortfall reported.

8. Amendment A2 — the pre-run critic returned NEEDS-REDESIGN, and it was right

critic.md, qwen/qwen3.7-max, $0.05861709, finish_reason: stop, run before any judging call. Three BLOCKING findings, three non-blocking. Every one is acted on below, and the two that matter change what this run measures. Sections 1, 3 and 4 above are superseded where they conflict with this section; nothing has been dispatched under the superseded version.

A2.1 — finding 1, BLOCKING, ACCEPTED: DA vs D1 was a token count, not a measurement

"K9 compares DA (0 foreign words) against D1 (N-1 foreign words) … the model will trivially pick D1 because it literally contains foreign words, measuring only its ability to see non-English text rather than the 'world cost' of domestication."

This is correct and it is the same defect that killed S171's registered primary one session ago (note (bmv): a primary that is true by construction predicts nothing). Asking which of two pages is "more clearly set outside the English-speaking world" when one has six Greek words in it and the other has none is not an experiment.

The fix is the critic's own, and it makes the run better than the design it replaces. A sixth form is added:

form build
NA every manipulable site NEU — the page fully rendered in location-free English

DA and NA carry zero source words each and differ only in the kind of English that replaced them: the English domestic article against the location-free phrase. On S5, DA reads a pinch of snuff … his breeches … his nightcap … Mr. Margaritis's counting-house; NA reads a pinch of powdered tobacco … his baggy trousers … his woolen cap … Margaritis's shop. Neither is Greek.

K9 becomes DA vs NA, FOREIGN, both orders, all 15 windows.

A2.2 — what the run now measures, which is a better question than the one it started with

K9 and K10 are now the same contrast at two doses: the English domestic article against the location-free phrase, at one site (K10, D1 vs T1, against a background that still carries the source words) and at every site (K9, DA vs NA, against a background that carries none).

Rate now means: the fraction of decided, cue-attributed judgments choosing the ENGLISH-DOMESTIC-ARTICLE form as more clearly set outside the English-speaking world. P6 and P7 are re-registered on that meaning:

The joint reading, registered now: P6 SPECIFIC with P7 NOT SPECIFIC → the effect is real but needs many sites, and one careful choice buys nothing measurable. Both SPECIFIC → it is there at a single word. Both NOT SPECIFIC → on the world channel the three roads of cultural-mediation collapse to two, transfer and everything-else, which would be the most consequential thing this arm could return.

What is given up, and it is the arm's original wording. ARM-dose step 2 asked whether the world cost grows with dose. This run does not answer that and the honest reason is that the question as posed cannot be answered by this instrument: on the world question, "more domesticated" and "fewer foreign tokens" are the same variable, so any dose ladder built from D0…DA measures the token count. That finding goes in the result page and closes the arm; it does not become a new arm.

A2.3 — finding 2, BLOCKING, PARTIALLY ACCEPTED, with the disagreement written down

"K10's signal is drowned out by the remaining source words … likely resulting in NEITHER or random guesses."

The critic proposes moving the minimal pair onto a fully domesticated background — which is K9, and is now in the run. K10 is kept as it stands, and the reason is evidential: RS-20260812e K4 is the identical pair of forms, judged on the other question, and returned 56 of 58 = 0.966 with two NEITHERs in the whole run. The instrument demonstrably resolves this one site against this background. Whether it resolves it on the world question is exactly what is unknown.

So the critic's prediction is registered as a named outcome before dispatch: if K10 returns NOT SPECIFIC or a high NEITHER rate while K9 returns SPECIFIC, that is the critic's finding 2 confirmed, it is reported in those words, and no rescue is attempted.

A2.4 — finding 3, BLOCKING, ACCEPTED, and its arithmetic was checked and is right

"At n=60, the 90% Clopper-Pearson interval falls entirely inside [0.35, 0.65] only for k ∈ {29, 30, 31, 32}."

Recomputed: at n = 60 the reachable set is k ∈ 25…35 for [0.30, 0.70] and k ∈ 28…32 for [0.35, 0.65] — 11 of 61 possible outcomes against 5 of 61. The critic is right that a true rate near 0.40 could never have reached an equivalence verdict. The equivalence window is widened to [0.30, 0.70], registered here before dispatch, and both reachable ranges are printed by analyse.py so the choice is auditable. The SPECIFIC region is k ≤ 11 of 60 (rate ≤ 0.183).

A2.5 — a defect this cross-check found, in code inherited from two earlier runs

Checking the critic's arithmetic meant computing the intervals a second way, and the two ways disagreed. clopper_pearson as inherited from E-20260811h and carried into E-20260812e is wrong for k ≤ 1 when n is above about 50 — clopper_pearson(1, 60, 0.10) returned [1.000000, 1.000000] and clopper_pearson(0, 60, 0.10) returned an upper bound of 1.000000 against a true 0.048703. The cause is a missing symmetry guard: the continued fraction for the incomplete beta converges only for x < (a+1)/(a+b+2), and the guard was never written.

No published figure is false. Every k ≤ 1 figure in RS-20260811h and RS-20260812e is at n = 30, where the routine is correct; all five published intervals were re-derived by the exact binomial and agree to three decimals. But K9 and K10 here run at n = 60, and K7/K8 came back at 0 of 30 — so this run could very easily have landed in the broken region and published a false interval. Repaired in analyse.py with the guard; verify.py computes the same bounds by bisection on the exact binomial tail and the two now agree on all 272 (k, n) pairs tested. This is method work done as a gate inside the unit it blocks (continue-prompt.md §4.5), not as a unit of its own. Method note (bmy).

A2.6 — finding 4, NON-BLOCKING, ACCEPTED: K11 is a positive control, not a gate

"K11 … can only fail if the model is completely broken."

True. K11 is relabelled a positive control and is no longer described as a gate whose failure withholds the primaries. It stays in the run because it replicates RS-20260812e K8 on byte-identical materials and extends it to the Greek windows, which is worth its 15 pairs. The gates that can actually fail are G2 (position preference), G3 (cue verbatim) and K14 — the Greek windows' DA vs D0 on BRIT, which will fail if the Greek DOM renderings do not read as British. If K11 nevertheless comes back at anything other than near-zero, the run is stopped and reported as an instrument failure.

A2.7 — finding 5, NON-BLOCKING, ACCEPTED and hardened

"the authors can 'rescue' a failed or inconclusive primary by pointing to the replication set."

Registered, categorically: the five Greek windows are descriptive only. They may not overturn a primary, may not rescue an INCONCLUSIVE one, and may not supply a causal conclusion of their own. If the frozen ten return INCONCLUSIVE, the run reports INCONCLUSIVE, whatever the Greek windows say.

A2.8 — finding 6: no leakage found in the build

Recorded. The critic examined americanise, deitalic and the skeleton check and found the build clean. NA is built through the same functions and re-asserted by G4a, G4c, G4g and, for the ten frozen windows, by G4h — the splice that produces NA from the frozen DA is validated by re-splicing the same spans with the TRA strings and asserting byte-identity with the frozen D0, which it achieves on all ten.

A2.9 — the run as dispatched

id domestic-article form comparator question windows orders pairs role
K9 DA NA FOREIGN all 15 both 30 PRIMARY — at every site
K10 D1 T1 FOREIGN all 15 both 30 PRIMARY — at one site
K11 DA D0 FOREIGN all 15 one 15 positive control, replicates K8
K14 DA D0 BRIT 5 Greek one 5 gate on the new material

80 pairs × 3 seats = 240 bodies, plus 4 duplicate pairs × 3 = 12. Cost unchanged; the dry run is re-priced against the amended items.json before dispatch and its output is reproduced in the result page.