Repository path: workshop/experiments/E-20260812e-dose/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260812e-dose |
| status | frozen |
| created | 2026-08-12 |
| updated | 2026-08-12 |
| senses | cultural-mediation, perceived-source-carriage |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-dose.md, wiki/findings/results/RS-20260811h-domestication-channel.md, workshop/translations/hastrman/R06-v1/translation.md, workshop/translations/jutrenje/R06-v1/translation.md, workshop/experiments/E-20260811h-domestication-channel/design.md, config/models.md, config/budget.md, wiki/goodness-senses.md, framework/v0.2/README.md |
E-20260812e-dose — how many words does it take?
ARM-dose step 1. Frozen before dispatch. code.py and materials/items.json are frozen with
it and are part of the design.
1. The question
RS-20260811h established that a page whose culture-bound items have been replaced by English
domestic articles reads as British against the same page carrying the source words, at 24 of 24
cue-attributed judgments and 8 of 8 passages. Its DOM arm substitutes every such item at once.
Its own §6.5 says what that leaves undone: "A translator who domesticates four sites in forty is not
described here."
This run asks where on the road between the two ends the effect actually sits — and whether the first step already spends all of it.
Two things are asked at once, because one is meaningless without the other:
- Dose. Domesticate one item; half the items; all the items. Does the judgment move in proportion, or is it at ceiling from the first word?
- Specificity. If one substituted word is enough, is it enough because the word is British —
or because there is one fewer foreign word on the page?
RS-20260811h'sC2makes this a live rival and not a quibble: on that run a location-free rendering read more British than a transferring one, at 34 of 40. So "one fewer foreign word" is a mechanism the project has already measured, and a dose-1 effect must be shown not to be it.
2. Materials
Ten windows, each built into five forms differing only inside sites declared in a frozen translator's log.
| win | hand | source lang | words | sites | manipulable | dose ladder (1 / h / N) | lead site |
|---|---|---|---|---|---|---|---|
H1 |
lead | Czech | 198 | 7 | 7 | 1 / 4 / 7 | cylindr → beaver / tall hat (STRONG) |
H2 |
lead | Czech | 289 | 9 | 8 | 1 / 4 / 8 | Bruská brána → town bar / gate in the walls (STRONG) |
H3 |
lead | Czech | 311 | 10 | 9 | 1 / 4 / 9 | kupec → greengrocer / shopkeeper (STRONG) |
H4 |
lead | Czech | 245 | 7 | 7 | 1 / 4 / 7 | zlaté → guineas / gold pieces (STRONG) |
W1 |
lead | Serbian | 148 | 16 | 16 | 1 / 8 / 16 | pačaluci → wide gaiters / wide legs (WEAK) |
W2 |
lead | Serbian | 144 | 6 | 6 | 1 / 3 / 6 | petačka → firkin / small cask (STRONG) |
W3 |
lead | Serbian | 116 | 7 | 7 | 1 / 4 / 7 | fes → billycock / cap (STRONG) |
W4 |
lead | Serbian | 259 | 5 | 5 | 1 / 2 / 5 | dukat → sovereign / gold piece (STRONG) |
W5 |
lead | Serbian | 221 | 5 | 5 | 1 / 2 / 5 | peć → range / stove (STRONG) |
B |
Field (published) | Russian | 128 | 4 | 4 | 1 / 2 / 4 | chinovnik circle → departmental circle / working circle (WEAK) |
H1–H4 are new, cut from T-hastrman-R06-v1 — Jan Neruda, «Hastrman» (1878), the whole story
translated by the lead this session under R06 and frozen at commit d528410 before this design
existed, with its 33-site culture-bound table written in the translator's log as the translation
was written. W1–W5 and B are E-20260811h's own windows, rebuilt from its frozen sites.json
rather than restated.
Two windows are dropped by a rule fixed before the build, not by inspection: a window needs
≥ 4 manipulable sites to carry a three-point ladder, and A (3) and W6 (3) do not have them.
This is declared here because W6 is RS-20260811h's striking one-site window and its absence is
therefore not an accident to be discovered later.
Manipulable means the site's TRA and DOM renderings are distinct strings after asterisks are
stripped. Two H sites are not (Turnov; on the second floor — the frozen log's note 1 names both
in advance); they are held at TRA in every form and are constants.
The five forms
Over one seeded permutation of each window's manipulable sites (SEED = 20260812, code.py):
| form | dose | build |
|---|---|---|
D0 |
0 | every site TRA — the source word carried over |
D1 |
1 | perm[0] → DOM, every other site TRA |
DH |
⌈N/2⌉ | perm[:h] → DOM, the rest TRA |
DA |
N | every site DOM — RS-20260811h's ceiling arm |
T1 |
— | perm[0] → NEU, every other site TRA — the control |
perm[0] is the first site in the permutation whose TRA, DOM and NEU are three distinct
strings, so that T1 is never degenerate. Eight of the ten lead sites are flagged STRONG in
their frozen logs and two (W1, B) WEAK — a property of the seed, recorded now.
Orthography is forced to American in every form by one map applied identically (G4b), so that no
"reads more British" verdict can be produced by a spelling. The new translation is written in
British spelling and the map covers it; residual British lexis in the untouched material is a
constant across all five forms of a window and raises every form's floor equally.
3. Procedure
Forced pairwise comparison. A seat is shown two versions of one window, labelled A and B, told they
differ only in a handful of words, asked one question, and required to return one line of JSON:
the answer, the single word or phrase that decided it, quoted exactly, and a confidence 0–3.
NEITHER is available. Prompt template and question strings are byte-identical to
E-20260811h's — the whole point of reusing them is that K3 below is a replication.
| id | high-dose form | low-dose form | question | orders | role |
|---|---|---|---|---|---|
K1 |
D1 |
D0 |
BRIT | one | the bottom of the curve — does one word move it? |
K2 |
DA |
D1 |
BRIT | both | THE PRIMARY — does going the rest of the way add anything? |
K3 |
DA |
D0 |
BRIT | one | the gate — positive control, replicates RS-20260811h C6 |
K4 |
D1 |
T1 |
BRIT | both | THE PRIMARY CONTROL — same site, domestic against location-free |
K5 |
DH |
D1 |
BRIT | one | secondary, the middle of the curve |
K6 |
DA |
DH |
BRIT | one | secondary, the top of the curve |
K7 |
D1 |
D0 |
FOREIGN | one | the world channel at dose 1 |
K8 |
DA |
D0 |
FOREIGN | one | the world channel at dose N |
- BRIT — "Ignoring what the story is about and where it is set, and judging only the English itself — its diction and its writing — which version reads more as though it were written by a British writer for British readers?"
- FOREIGN — "Which version is more clearly set outside the English-speaking world?"
Rate always means the fraction of decided, cue-attributed judgments that chose the more-domesticated form, so a rate above 0.5 always means "more domestication was noticed".
Cue attribution is the primary reading (E-20260811h A8): a judgment counts toward a causal
claim only if its quoted cue lies inside a site that differs between the two forms shown. Counted
(unattributed) rates are reported alongside and are descriptive.
The inferential unit is the window (E-20260811h A4). Pooled rates with Clopper–Pearson
intervals are descriptive and overstate precision; the sign test over ten windows is the inference.
Seats. P1 openai/gpt-5.6-terra, P3 x-ai/grok-4.5, P2 google/gemini-3.6-flash
(reasoning: {effort: low}). deepseek/deepseek-v4-pro is not in the jury, for
E-20260811b's reason. Temperature 0. Judgment is not parallelised across seats within a pair;
the lead judges nothing (charter §5), and none of the ten windows is any seat's own output.
The pre-run critic is qwen/qwen3.7-max, which is NOT one of the three judging seats.
RS-20260811h §7.2 had to declare that its critic was also a judge; this run does not, and the
reserve slug is used precisely so that it need not. moonshotai/kimi-k3 is not asked, per note
(bhf) rule (iii) — it has returned a dead body from hidden reasoning at S106, S123, S127, S128 and
S162 and the rule is to change the seat, not the ceiling.
4. Predictions and verdicts, registered
G — the gate. K3 must reproduce RS-20260811h C6: DA chosen over D0 on BRIT in
≥ 9 of 10 windows (one-sided sign P ≤ 0.05 at 9/10 = 0.0107). If K3 fails, every other
verdict in this run is withheld, because a manipulation that does not work at full dose cannot be
read at partial dose.
P1 (K1) — does one word move it? Bar: pooled cue-attributed rate ≥ 0.75 and ≥ 8 of
10 windows favouring D1. PASS → one substitution is enough to move the judgment. FAIL → it is
not, and the curve has a floor above dose 1.
P2 (K2) — THE PRIMARY, an equivalence test.
- 90% Clopper–Pearson interval for the pooled cue-attributed rate inside [0.35, 0.65] →
SATURATED: the first domesticated item spends the whole price, and partial domestication is
not a partial choice.
- interval entirely above 0.65 → GRADED: more domestication reads as more British, and a
translator who Anglicises some items pays part of the price.
- otherwise → INCONCLUSIVE, stated as such and not relaxed.
P3 (K4) — THE PRIMARY CONTROL, also an equivalence test.
- interval entirely above 0.5 with a point estimate ≥ 0.65 → the dose-1 effect is specific to
the English domestic article, and P1 may be attributed to domestication.
- interval inside [0.35, 0.65] → NOT SPECIFIC: P1, if it passed, measures one fewer
foreign word, and no claim about domestication may be built on it.
- otherwise → INCONCLUSIVE.
Minimum denominator, computed before dispatch, per note (bhr). An equivalence verdict (P2,
P3) requires ≥ 30 decided, cue-attributed judgments. K2 and K4 each run both orders on ten
windows across three seats = 60 dispatches each, so a 50% loss to NEITHER, unattributed cues
and dead bodies still clears the bar. K1, K3, K5–K8 run one order = 30 dispatches each,
which cannot support an equivalence verdict at any loss rate — so none is registered on them,
and none may be given afterwards. This is RS-20260811h §5's failure written into the design instead
of discovered in the analysis.
P4 (K5, K6) — the middle of the curve, direction only. Reported as rates and window counts.
No equivalence verdict. Registered reading: if P2 returns SATURATED, then K5 and K6 should both
sit near chance; if either is clearly above 0.65 while K2 is not, the ladder is not monotone and
that is a finding against the saturation story, reported as such.
P5 (K7, K8) — the world channel, direction only. RS-20260811h §1.4 found that
domesticating the furniture moves the world as well as the prose. Prediction: K7 and K8 both
below 0.5. The dose comparison is the point: if the prose effect saturates (P2 SATURATED) but
K7 > K8 — one domesticated item costs less world than all of them — then the two channels have
different dose curves, and a translator's mixed choice buys something after all: the same prose
placement for less world. That combination is the one result of this run that would change practice,
and it is registered here before dispatch so that it cannot be told as a story afterwards.
Secondary, descriptive only: K1 and K4 split by the lead site's frozen STRONG/WEAK flag
(8 / 2). Two windows cannot support a verdict and none is registered.
5. Gates
| gate | bar | if it fails |
|---|---|---|
G1 returns |
every pair dispatched; empty and truncated bodies counted and reported, never imputed | reported in the result's limits |
G2 position preference |
≤ 0.70 on the A slot, per seat, over the whole run | that seat is dropped from primaries; both figures printed |
G3 cue verbatim |
≥ 0.90 of quoted cues occur verbatim in one of the two texts shown, per seat | that seat is dropped from primaries; both figures printed |
G4a–e build |
byte-identity outside sites · no British spelling · five forms pairwise distinct · D1/T1 differ at exactly one site · ladder strictly nested |
already run: 50 checks, all PASS |
G5 duplicates |
a seeded 8-pair subsample re-dispatched byte-identically | a report, not a gate |
G6 NEITHER |
reported per seat | — |
Note (bmb) is fixed in this runner and the fix is the reason it is written down. E-20260811h
lost 14 bodies to finish_reason: length because run.py tested if content: and truncated content
is not empty. Here re-dispatch fires on empty content OR finish_reason == "length", and the
max_tokens caps are raised to 350 / 1000 / 500 (P1 / P2 / P3) — P2 carries hidden reasoning
inside its cap and was the seat that truncated.
6. What this cannot establish
- Nothing about human readers. Three LLM seats, an instructed task. The word reader is not
used in any claim (
E-20260811hA18). - No goodness sense is scored. Tier D is NOT PASSED; every sentence here is
provisional. K3is near-tautological in the same wayRS-20260811hH1was, and is used only as a gate.- Six of ten windows are one hand. Four are new and in a new language pair; one is a published English hand. The dose result will still be a result about ten passages.
- The
STRONG/WEAKflags are one annotator's, in both frozen logs, with no second coder. - Which site is dose 1 is a seed's choice. Ten windows means ten independent draws, which is the
only defence offered; a different seed would put different words in the
D1slot. - The lead knew a manipulation of culture-bound items was coming while translating «Hastrman» —
ARM-dosewas constituted first. It did not know the doses, the contrasts, or which sites the permutation would pick. Declared on the translation artifact and repeated here. T-hastrman-R06-v1's contamination is declared and not measured, because the only English renderings of «Hastrman» are in copyright and none is legitimately reachable. This design does not turn on independence: every form compared here is an edited copy of the same text, so whatever it shares with Pargeter or Heim, it shares in both arms of every pair.
7. Cost — pre-flight
100 pairs × 3 seats = 300 bodies, plus 8 duplicate pairs × 3 = 24. Worst case is built from
max_tokens and the exact prompt length of every pair, per note (abc).
python3 run.py --dry-run prints, from the frozen items.json:
| line | worst case |
|---|---|
| study, 100 pairs × 3 seats, caps 350 / 1000 / 500 | $1.575703 |
G5 duplicates, 8 pairs × 3 seats |
$0.126056 |
re-dispatch contingency (10% of P2 bodies dying once and re-sent at 2× cap) |
$0.150000 |
pre-run critic, qwen/qwen3.7-max, cap 16,000 |
$0.096000 |
| total worst case | $1.947759 |
Declared ceiling: $2.00, inside a UTC-day headroom of $4.455707 at the time of writing
($0.544293044 already spent today across four sessions). The caps in run.py are the caps this
table is priced from — the check note (abc) exists for, and the one S166 had to make mid-run.