Repository path: workshop/experiments/E-20260811b-realia-channel/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260811b-realia-channel |
| status | frozen |
| created | 2026-08-11 |
| updated | 2026-08-11 |
| track | T3 |
| senses | cultural-mediation, perceived-source-carriage |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-realia-channel.md, workshop/translations/fengbo/R06-v1/translation.md, workshop/translations/fengbo/contamination.md, wiki/findings/results/RS-20260811-floor.md, workshop/experiments/E-20260811-floor/code.py, wiki/goodness-senses.md, config/models.md, config/budget.md |
E-20260811b — the two channels a translated page is located by
ARM-realia-channel step 1. Frozen before any judge call and before any material was built.
Tier D is NOT PASSED: every jury figure this design can produce is an LLM-panel-perceived
judgement about English and no goodness sense is scored.
1. Question
A translated page is located twice. It is located by the world it describes — roubles, a samovar, a queue, a yamen, Akaki Akakievitch — and it is located by the English it is written in — colour against color, towards against toward. The translator controls the first almost completely and the second hardly at all.
Can a reader tell them apart?
Specifically: does removing the source-culture words change a reader's judgement of whose English the page is written in, when not one letter of spelling has changed? And the mirror: does changing the spelling change a reader's judgement of where the story is set, when not one culture-bound word has changed?
2. Why it is a translation question and not an instrument question
The subject rule (wiki/tracks.md, continue-prompt.md §4.5) requires one sentence saying what
this teaches about translating literature or evaluating translations. The sentence:
It measures what the oldest decision in translation — keep the source's word or replace it — does to a reader, on a channel the translator did not intend to touch.
RS-20260811-floor §6 met this phenomenon as a contaminant of its own instrument and filed it as a
limit. This design turns it round and makes it the object. What comes out lands on
wiki/goodness-senses.md's cultural-mediation, whose entire content is culture-bound items and
whose only evidence to date is about whether raters sort such decisions into the right box, not
about what the decisions do.
3. Design — a 2×2 crossed within passage
Two factors, both applied to the same passage, so that everything else about it — its author, its subject, its period, its difficulty, and any recognition a judge has of it — is a constant and cannot produce an effect.
| factor | levels | what changes |
|---|---|---|
W (world) |
K keep · M mute |
each declared realia span is replaced by a generic English equivalent. Nothing else changes |
P (prose) |
B british · U us |
each declared orthographic span is swapped to the other side of the same word. Nothing else changes |
Four forms per passage: KB KU MB MU. Plus two controls (§5).
3.1 Passages — eight, selection rule stated before selection
From the eighteen texts of E-20260811-floor/materials/corpus.json, already on disk and already
public domain, plus the translation limb T-fengbo-R06-v1 frozen this session at 79835e3.
Selection rule, applied in this order and mechanically where it can be:
- The text must be translated narration (or, for the two originals in the corpus, excluded — an English-original passage has no source-culture world to mute).
- A candidate window is a contiguous run of 110–150 words beginning at a paragraph boundary.
- The window must contain ≥ 2 orthographic spans that are members of
code.py's_PAIRSinvertible list, so thatPis manipulable. - The window must contain ≥ 4 realia spans, so that
Wis manipulable. - The earliest qualifying window in the text is taken. No window is chosen for what it says.
- One passage per hand, and no more than two per source language.
Realia spans are annotated by the lead, because no word list can find them. This is declared as a
non-blind step and is bounded three ways: the annotation is committed as data before any judge is
called; it is made without knowing which of QP/QW any judge will answer first; and every
substitution is mechanically checked against §5's conditions, which forbid the annotation from
touching the orthographic channel at all.
3.2 The two elicitations, in separate calls
The whole design turns on the two questions not contaminating each other, so a judge never sees both about the same passage in the same context. One item, one call, stateless.
QP — the prose question. Which national variety of English is this passage written in?
BRITISH, AMERICAN or CANNOT-TELL. Then quote up to three words or phrases from the passage that
decided it. Then rate 0–6 how strongly the writing itself — spelling, idiom, phrasing — is
marked as one national variety. The prompt states, in terms, that subject matter and recognition
of the text are not grounds and that the answer in that case is CANNOT-TELL — the same
instruction RS-20260811-floor used and which did not hold; it is repeated here so that the
comparison to that run is like for like, and its failure is not this design's primary.
QW — the world question. Where is this passage set? Name a country, or CANNOT-TELL. Then
quote up to three words or phrases that decided it. Then rate 0–6 how strongly the passage's
subject matter locates it in a particular country or culture.
3.3 Judges
Three seats, three labs, per config/models.md. L1 x-ai/grok-4.5 (P3) · L2
deepseek/deepseek-v4-pro (P5) · L3 google/gemini-3.6-flash (P2), the last with
reasoning: {"effort": "low"} — note (bmb), which cost S156 45.7% of its spend in dead bodies
and is applied here from the first dispatch. The pre-run critic is openai/gpt-5.6-terra (P1)
and judges nothing. z-ai/glm-5.2 is excluded as a seat on the measurement in
RS-20260811-floor §6: it flipped its verdict on three of four sentences it saw twice unchanged.
4. Predictions, registered
| id | prediction | statistic | bar |
|---|---|---|---|
P1 |
The world leaks into the prose judgement. Muting the realia raises CANNOT-TELL on QP, though no spelling changed |
CANNOT-TELL rate on QP, M minus K, pooled over 8 passages × 2 orthographies × 3 seats |
≥ +0.15, and an exact paired permutation test over passages at P < 0.05 |
P2 |
The prose does not leak into the world judgement. Flipping the orthography leaves where is this set alone | (a) share of QW country answers that change between B and U within a passage-and-W cell; (b) mean loc rating difference |
(a) < 0.10 and (b) CI includes 0 |
P3 |
The leak is asymmetric. | P1's effect size minus P2(b)'s, on the two 0–6 ratings |
P1's nat shift strictly larger |
P4 |
Manipulation check — the seats do read spelling. On K items, the QP verdict follows the orthography shown |
share of committed QP verdicts matching the shown side, on K items |
≥ 0.70 |
P5 |
The translator's frozen prediction is worth something. On the lead's passage, QW quotes land on sites the translator's log marked predicted-placing: yes more often than on sites it marked no |
rate per site class, reported with an exact test | reported, not gated — n is one passage |
P1 is the primary. If P1 fails, the finding is that the two channels are separable and the
translator is free; that is a first-class result and is written as such.
5. Controls and gates
| gate | what it is for | bar | consequence of failure |
|---|---|---|---|
G1 returns |
100% of cells after at most one licensed re-dispatch round | missing cells are reported, never imputed | |
G2 duplicate stability |
the control without which no flip can be read. Four passages' KB form is presented a second time, byte-identical, under a different item id. A seat that does not repeat its own QP verdict is measuring noise |
a seat must repeat on ≥ 3 of 4 | that seat is disqualified from P1, P2, P3, its numbers reported and excluded. RS-20260811-floor lost a seat exactly here |
G3 sham edit |
the control that separates the world was removed from the text was edited. On four passages, a SB form replaces the same number of realia spans with different but equally source-marked realia (roubles → kopecks, samovar → ikon-corner) |
if P1's effect appears on SB as strongly as on MB, P1 is withheld — the cause is editing, not the world |
stated in the result, primary withheld |
G4 span validity |
every derived form differs from its base only in the declared spans | checked character by character, all forms | any failure voids that passage |
G5 the lead's passage |
contamination on T-fengbo-R06-v1 is suspected at a 26-token run (contamination.md) |
the lead's passage is X1, never pooled into any figure characterising published hands, and supports no claim requiring an independent hand |
structural, not a test |
G6 manipulation |
P4 |
P1's null is uninterpretable on a seat that cannot see spelling at all: such a seat is reported and excluded from P1 |
stated in the result |
6. Materials, counts and cost
| passages | 8 |
| forms | KB KU MB MU on all 8 = 32; SB on 4 = 4; duplicate KB on 4 = 4 |
| items | 40 |
| questions | 2, in separate calls |
| seats | 3 |
| calls | 40 × 2 × 3 = 240 |
| prompt size | ~450 tokens/call |
max_tokens |
700 per call, and note (abc) governs: the ceiling below is built from that cap, not from an expected answer length |
Pre-flight ceiling. Worst case at max_tokens on the most expensive seat's list price, times a
4× routing allowance (the provider caution in config/models.md): 240 × (450 × $3/M + 700 ×
$15/M) ≈ $0.09 × 4 ≈ $0.35 for the dearest seat if it were all three; the realistic mixed figure
is well under that. Plus one critic pass at a 12,000 cap — S156's first critic pass truncated at
6,000 and had to be re-run, so the cap starts at 12,000 here — at ≈ $0.10.
Declared ceiling: $1.20. Today's UTC ledger stands at $0.329161470 of $5.00 (S156), so headroom is $4.67 and this fits without deferral.
7. Procedure
code.pybuildsmaterials/passages.json: windows selected by §3.1, orthographic spans located by importingE-20260811-floor/code.py's_PAIRSandmarks(), realia spans read from a committed annotation file, all forms generated,G4asserted.- Pre-run critic (
openai/gpt-5.6-terra, cap 12,000) over this design and the built materials. Findings adjudicated in writing incritic.md; amendments committed before any judge is dispatched. run.pydispatches 240 calls, one item per call, raw request and response JSON preserved per call,usage.costandproviderread off every response.analyse.pycomputes every figure in §4 and §5.verify.pyrecomputes every number the result page reports, from the raw files, by a different route, and includes mutation tests that must be caught.
8. Failure criteria, written before the run
G2fails on two or more seats → the run cannot supportP1and the result says so.G3fires →P1withheld.P4< 0.70 on all three seats → the design measured seats that do not read orthography;P1's direction is uninterpretable and is reported as such.P1fails → the channels are separable, reported as the finding, not as a disappointment.
9. What this design cannot do
- It is a forced-choice task, not reading.
RS-20260811-floor§9.5 applies unchanged: a judge asked whose English is this is doing something no reader of a novel does. - Eight purposive passages, five source cultures, two of them Russian. Every figure is corpus-specific.
MUTEis a caricature of domestication. A real domesticating translator does not replace every culture-bound noun with a generic; the manipulation is the extreme end of a decision that is normally made site by site.- The realia annotation is the lead's, and the lead also wrote one of the eight passages.
G5bounds the second; nothing bounds the first except the mechanical checks inG4. - LLM seats are not readers. Nothing here is licensed as a claim about human reading.
10. Order of operations, recorded because S156's was not clean
T-fengbo-R06-v1 and its translator's log were committed at 79835e3, before this file
existed. This file is committed before code.py runs, before any material is built, and before
any judge is called. Everything that changes after this commit is recorded as a numbered amendment
in critic.md with its direction of bias, and the amendments are committed before dispatch.