Repository path: framework/v0.3/entries/HB-realia.md · rendered 2026-09-09
Page metadata (front matter)
| type | entry |
|---|---|
| id | HB-realia |
| status | active |
| created | 2026-09-05 |
| updated | 2026-09-05 |
| senses | cultural-mediation, naturalness, perceived-source-carriage, consistency |
| pairs | RU→EN, PL→EN, JA→EN, ZH→EN, SR→EN, CS→EN, EL→EN, NL→EN |
| provisional | true |
| internal-judgment-only | true |
| links | framework/v0.3/README.md, framework/v0.2/README.md, wiki/goodness-senses.md, wiki/findings/results/RS-20260811b-realia-channel.md, wiki/findings/results/RS-20260811h-domestication-channel.md, wiki/findings/results/RS-20260812e-dose.md, wiki/findings/results/RS-20260813a-world-dose.md, wiki/base/anchors/A-shaw-spider-thread/A-shaw-spider-thread.md, wiki/base/anchors/A-garnett-vanka/A-garnett-vanka.md, workshop/translations/max-havelaar-i-c/R04-v1/translation.md, workshop/canon/max-havelaar-i-c/manifest.md, wiki/plan.md |
Culture-bound words: keeping the world without giving away where the prose is written
Standing. Evidence classes X1a (two Tier 1 precedent anchors, published translations read closely
against source, internal-judgment-only) and X2 (four machine-scored experiments, three LLM seats
on instructed forced-choice tasks, recomputed from raw outputs). No X3: Tier D is NOT PASSED, so
nothing here is a jury's word that one rendering carries the source's world, or reads more
naturally, than another — only whether a strategy is present and what a seat's forced choice does
with it. Consolidated 2026-09-05 (S249) from v0.2 §7.5, §7.6, §7.9, §7.10.
1. The problem
A source text is full of words that carry a piece of its own culture along with their sense — coins, ranks, institutions, foods, customs, proverbs, place names, units of measure. A target reader may not own any of that furniture. At every such site the translator picks one of a small set of moves — carry the foreign word over, reach for the nearest article of the target's own world, describe it in location-free terms, gloss it, or cut it — and the evidence below says that choice is doing two separate jobs at once, not one: it decides how native the English sounds (which country's English this reads as), and, independently, it decides how foreign the story's own setting still reads. A translator who reaches for a natural-sounding target-culture substitute "to make the prose read well" is, on this evidence, also quietly relocating the story — and no strategy measured here buys the native-sounding prose without paying for the setting.
2. What published translators do
| pair · work | hands | measured | result | source |
|---|---|---|---|---|
| JA→EN Akutagawa 「蜘蛛の糸」 | Shaw 1930 | one class of Buddhist-afterworld realia, read whole against source | 4 different strategies in 1,000 words: substitute (極楽→"Paradise"), transparent calque (血の池→"Pond of Blood"), copy-opaque (三途の河→"Sanzu-no-Kawa"), copy-opaque+incidental scaffold (針の山→"the needle of the dread Hari-no-Yama") |
A-shaw-spider-thread |
| RU→EN Chekhov «Ванька» | Garnett 1922 | realia handling, whole 1,250-word story, read against source | same range used inconsistently — two dogs (Каштанка, Вьюн) handled two different ways in one sentence; a new handling, measure conversion (пудового сома→"a forty-pound sheat-fish"); overlap-cheap sites (a last, an icon) against copy-opaque residue (a proverb, a folk custom) |
A-garnett-vanka |
| RU, PL, JA, ZH→EN, 7 passages (Garnett/Field/Hapgood/Shaw/Morri published; Benecke & Busch published Polish; lead ZH) | machine census, 3 seats | does deleting every culture-bound item change the verdict on whose English this is | no: +0.048, CI −0.069 to +0.164, P = 0.3125, n = 7 (registered bar ≥ +0.15, P < 0.05, not reached) — but the evidence seats cite for that verdict is 13.5% realia when the world is present, 0% when it is muted, and every one of those twelve cues is the translator's own domestication (halfpennies, smock, councillor), never a retained foreign word (taiga, yamen, Sanzu-no-Kawa, cited 118 times for the world question and never for the prose one) | RS-20260811b |
| SR→EN Lazarević (fresh lead rendering) + RU→EN Garnett, Field (published) | 3 seats, 8 passages, 49 sites | forced choice: TRANSFER vs DOMESTICATE vs NEUTRALISE (location-free), plus an edit-matched sham | DOMESTICATE beats TRANSFER on "whose English" at 24/24, 8/8 passages (P = 0.0039) and beats NEUTRALISE at 48/48; TRANSFER beats NEUTRALISE at only 6/40 = 0.150 (reverses a naive expectation — carrying the source word over reads as less British than a location-free English phrase); one nation-marked substitution in 347 words flips 6/6 judgments; asked instead which page is more clearly set outside England, seats choose the domesticated page at only 1/17 | RS-20260811h |
| CS→EN Neruda «Hastrman» (fresh lead rendering) + 5 of the Serbian windows + 1 published English hand | 3 seats, 10 windows, dose ladder (none / one / half / all sites domesticated) | is there a middle position, or does one item already saturate the prose channel | nothing saturates: one beats none 28/28, half beats one 30/30, all beats half 30/30, all beats one 60/60 (sign P = 0.00098); but the effect is already at ceiling from one item (1.000, same as the full-substitution page) — one word in 245 ("Three guineas!" for "Three zlaté!") was enough; a same-site control (guineas vs gold pieces) confirms it is the domestic article, not the missing foreign one, that does this (56/58 = 0.966) | RS-20260812e |
| CS, SR, EL→EN, 15 windows, 4 source languages | 3 seats | DOMESTICATE (full) vs NEUTRALISE (full) — both pages carry zero source words; which reads as more foreign-set | the location-free page wins at 42/49 cue-attributed judgments, 30/33 at a single differing site (both below the registered equivalence boundary, though the ten-window sign test itself does not reach P < 0.05) — except at a proper name: Prague's Malá Strana→"the Lesser Town" (a transparent calque) reads as more foreign-set than the location-free "this quarter," the one window-level reversal in the run | RS-20260813a |
The regularities, each with what it was measured on:
- What leaks into a judgment of the English is the translator's own domestication, never the
retained foreign word. All twelve realia-cues for "whose English" in a seven-passage census are
the translator's own Anglicisations; the genuinely foreign items in the same corpus are cited 118
times for the world question and not once for the prose one. (
RS-20260811b; RU, PL, JA, ZH.) - Put to a direct forced choice, domestication is the strongest lever this project has measured
for making prose read as target-nation writing — one nation-marked substitution in 347 words
flips every judgment; eight such words push the effect to a ceiling that eight more cannot beat.
(
RS-20260811h,RS-20260812e; SR, RU, CS.) - Transferring the source word over does not merely fail to place the prose in the target nation —
it places it the other way. A page carrying the source's own foreign words reads as less
British than the same page in location-free English (6/40 against a registered floor), even though
the same seats correctly call the transferred words foreign when asked about the world.
(
RS-20260811h.) - Domestication is not all-or-nothing, and nothing saturates on the prose channel — but one item is
already most of the effect. A ten-window dose ladder (none/one/half/all) is monotonic all the way
to "all," and yet the single-item rung already sits at the same ceiling as full substitution,
confirmed to be the domestic article doing it and not the absent foreign one (a same-site control
at 0.966). (
RS-20260812e.) - Domestication moves the world question too, in the opposite direction, at every dose measured
— there is no dose that leaves the setting alone. Asked which page is more clearly set outside
England, seats pick the domesticated page at only 1 of 17 against the untouched original, and the
fully transferred page beats both the one-item and the fully domesticated pages at 30/30 on the
same question. (
RS-20260811h,RS-20260812e.) - Once the source word is already gone, the kind of English that replaces it still decides where
the story reads as set. A location-free phrase ("a gate in the walls") reads as more foreign-set
than a domestic substitute ("the town gate") even though both pages carry zero source words — the
third road is a real, distinguishable position, not a synonym for domestication once the foreign
word is gone. (
RS-20260813a.) - The world-channel finding reverses at a proper name, and the mechanism is legible. A calqued
toponym (Malá Strana→"the Lesser Town") keeps the place, because it Englishes the words without
deleting the referent; a location-free phrase ("this quarter") deletes the place outright. Calque
and location-free are opposites at a name, not neighbours on the same scale they occupy at a common
noun. (
RS-20260813a.) - No competent published translator holds one policy for a coherent class of culture-bound items.
Shaw handles Buddhist-afterworld place names four different ways in 1,000 words; Garnett handles two
named dogs two different ways in one sentence. This is normal published practice, not a lapse — but
it is measurably a
consistencycost: the reader's map of the fictional world stops cohering. (A-shaw-spider-thread,A-garnett-vanka.) - Where the source and target cultures already overlap, the fork is nearly free; it bites only at
the residue that does not overlap. A cobbler's last is "a last," an icon is "the ikon," a
matins service is "the midnight service" — no decision worth logging. The fork reappears, and goes
copy-opaque by default, at the proverb, the folk custom, the private joke that has no target-culture
counterpart at all. (
A-garnett-vanka.) - A silent unit conversion is itself a domesticating move, and the most invisible one measured.
Recalibrating a пуд into "forty pounds" is neither substitution, calque, nor copy — it reads as
if no decision was made at all, which is exactly what makes it worth deciding on purpose.
(
A-garnett-vanka.)
3. What this project's own practice found
internal-judgment-only throughout; every item below is process observation from inside the
translator or the experiment-builder, not a scored result.
- Constructing the dose ladders required freezing a realia annotation before any design existed
(the
STRONG/WEAKtags on 49 Serbian sites and 33 Czech sites): a live decision about which items a translator's own log calls load-bearing, made once, by one coder, and then held fixed through every downstream manipulation — the annotation itself is a craft act this project has not yet had an independent second coder check at scale (§6). T-max-havelaar-i-c-R04-v1(NL→EN, this session's translation limb, 403 Dutch words): applying §5's guidance below in session surfaced a site the census-style experiments above do not have a category for — a bare, unglossed place name that is doing social work no dictionary entry recovers (Driebergen, where a comfortable Amsterdam merchant of 1860 retires). Transferring it plain, per item 4's proper-name rule, produces exactly the copy-opaque residue §2 items 8–9 catalogue in Shaw and Garnett — the loss is not hypothetical, it is what this session's own rendering does. Also confirmed at the point of choosing: two ordinary-seeming professional words (doctor and apothecary) are a domestication decision in miniature, since either obvious English alternative (chemist, druggist) is a nation-marked substitute of exactly the kind item 2 above says a single instance of can flip a reader's sense of whose English the page is.
4. The options
At a culture-bound site the moves are transfer (keep the source word, marked or not), substitute / domesticate (the nearest article of the target's own world), neutralise (a location-free description), gloss, calque (translate the foreign structure literally), measure conversion (silently recalibrate a quantity), or cut. The evidence says the option set splits cleanly by two things:
- Whether the site is a common noun/institution or a proper name. At a common noun, transfer and neutralise both keep the setting foreign; domestication alone relocates the prose. At a name, the pattern inverts: a transparent calque of the name keeps the place, and it is neutralising — deleting or paraphrasing it away — that erases the setting. (items 6–7.)
- Whether the two cultures already overlap at that site. Where they do, the choice is close to free and any competent move reads the same (item 9). Where they do not, every published hand measured reaches for copy-opaque by default (items 8–9), which is cheap to write and expensive to a reader who wants the sense as well as the string.
No option measured here buys both: a translator gets native-sounding prose or a preserved sense of place, and the trade is visible at a single word (item 5). Partial domestication is a real, graded middle position on the prose channel (item 4) but not a discount on the world cost, which appears to move as an early, not a late, expense.
5. Guidance
For a translator
- Enumerate every culture-bound site before opening any other translation, and split each one into
proper name or common noun/institution before deciding anything else — the two options below
answer to opposite rules. —
evidenced (RU, PL, JA, ZH, SR, CS, EL). - Decide, site by site, whether you are optimising for native-sounding prose or for a preserved
sense of place — not both. Reaching for the nearest target-culture article "to make it read
naturally" is a decision to relocate the story, not a neutral polish; it is the single strongest
lever this project has measured for changing a reader's sense of which nation wrote the page, and
it costs the setting almost completely (1 of 17). —
evidenced (SR, RU, CS). - At a common noun or institution where you want to keep the setting without manufacturing
unreadable opacity, prefer a location-free description over a domestic substitute. It reads as
more clearly foreign-set than a domestic article, at the cost of a few extra words and a
periphrastic feel — a real middle road, not a compromise that collapses into domestication once
the source word is gone. —
evidenced (CS, SR, EL). - At a proper name, that rule reverses: do not neutralise or paraphrase it away. A location-free
phrase at a name deletes the place; a transparent calque of the name's own structure keeps it while
still Englishing the words, and is the one move that has read as more foreign-set than the
alternative in this project's evidence. —
evidenced (EL), n = 1 window, flagged. - Do not expect a single policy for a coherent class of items to hold, and do not read another
translator's inconsistency as incompetence — the strongest published hands measured (Shaw,
Garnett) both cross strategies within one short text at the same class of item. If a coherent
fictional world matters to your purpose, that coherence is a decision to make deliberately, not a
default the class of item will hand you. —
evidenced (JA, RU), descriptive. - Do not labor the sites where source and target cultures already overlap — a shared institution,
food, or unit is nearly free to render either way, and no evidence here distinguishes the choices
at such a site. Spend the decision budget on the residue that does not overlap, where a translator
defaults to copy-opaque unless a deliberate choice is made otherwise. —
evidenced (RU), descriptive. - Partial domestication is a real middle position, but do not expect it to buy you a discount on
the world cost. Every rung of a dose ladder from one item to all of them is measurably distinct
on the prose channel; on the world channel, even the lightest dose measured already reads as less
foreign-set than the untouched page. —
evidenced (CS, SR). - A silent unit or measure conversion is a domestication like any other — decide it on purpose.
It is the easiest of the strategies to make invisibly, which is exactly why it should not be a
habit. —
evidenced (RU), n = 1.
For a pipeline
- Enumerate: detect culture-bound spans in the source (named entities, measures, institutions,
customs, foods) before any target text exists, and classify each as proper name or common
noun/institution. —
untested. - Classify overlap: for each span, whether the target culture has a ready near-equivalent
(overlap) or not (residue) — from a declared reference inventory for the pair. —
untested. - Set the policy parameter: prose-nativeness vs. world-placement priority, from declared
purpose (
wiki/goodness-senses.md§declared-purpose). Default per class from §4: common noun/residue → location-free description; common noun/overlap → either strategy, unconstrained; proper name → transparent calque where the name's structure allows one, transfer otherwise; never domesticate a proper name. —untested. - Render under the policy; flag for human review any single domesticating substitution on an
otherwise placeless page, since one item has been shown sufficient to flip a prose-nativeness
verdict at ceiling. —
untested. - Check: count spans by strategy used against the declared policy; flag inconsistent handling of
sites in the same coherent class (two names of the same kind handled two different ways) as a
consistencyrisk unless the inconsistency is declared deliberate; report the per-class counts. —untested.
Human entry points. Step 3's dial (which channel this rendering is optimising for) is a purpose decision for a person; step 4's single-substitution flag and any site the classifier cannot place in either overlap class are the places to route to a person. Without one, step 3 defaults to location-free description for common-noun residue and transfer for names — the reading that keeps the world at the cost of some readability, on the reasoning that a lost sense of place is not independently correctable later in a pipeline the way a domestic-sounding phrase can be if the purpose turns out to want it.
Application. This session's translation limb (T-max-havelaar-i-c-R04-v1, Multatuli's Max
Havelaar, NL→EN, a fresh 403-word passage, the "Lukas paragraphs") applies items 1–4 and 8 at six
decision sites: an occupational term calqued rather than promoted or Americanised (item 1's split);
an institution transferred bare because the source itself supplies no marking device (de Bank); a
currency retained and Anglicised rather than converted (item 8); a bare place name transferred with no
added gloss, and declared rather than fixed as the passage's one uncorrected loss (item 4, at a
common noun's cost rather than a name's, since the site does not calque); and a profession word chosen
specifically to avoid an unintended nation-marking substitution (item 2). Followability log in the
translation's own file — every item above decided something, at no cost beyond the choosing, except
the Driebergen site, whose cost is a genuine, unrecovered loss of connotation rather than a
followability tax.
6. Not evidenced, and open
- All quantitative evidence here is X2: three LLM seats, instructed forced-choice tasks, at ceiling on most contrasts, so magnitude is not measured — only direction and, on the dose ladder, ordering. Tier D is NOT PASSED; nothing here licenses a claim about what a human reader notices or prefers.
- The world-channel dose-response is explicitly unmeasured. Whether the world cost grows with the number of items domesticated, rather than being paid in full at the first one, is recorded as a question this design cannot reach (note (bmx)), not as an open gap the next run should just re-try.
- The realia and
STRONG/WEAKannotations are the lead's own, one coder throughout, checked against an independent annotator only once, at 0.517 agreement. - Language coverage skews toward the languages this project has already worked in (Russian three times, Serbian, Czech, Polish, Japanese, Chinese, modern Greek, one Dutch passage added this session); no target other than English; no pair where the target culture is the less-furnished one.
- The proper-name reversal (§2 item 7, §5 item 4) rests on one window — one Prague street name in one experiment. Whether it generalises past toponyms with a translatable literal structure (would it hold for a personal name, a saint's day, an institution's proper name?) is untested.
- Nothing here separates "keeps more of the setting" from "reads more like a translation" on the
world channel — the location-free renderings measured are also the longer ones, and the seats'
quoted reasons are the periphrases themselves (
RS-20260813a§6). - Open: whether a light, selective purchase of the location-free strategy is distinguishable from
a plain domesticated page to an unprimed reader (no design has tried a partial dose of the third
road, only full domestication vs. full neutralisation); whether the name/common-noun split (§4)
holds for culture-bound items besides toponyms; whether the
Driebergen-style loss this session's translation limb surfaced — a bare toponym doing social work no gloss was supplied for — is common enough across this project's filed translations to be worth a census of its own.
7. Sources consumed
- v0.2 §7.5 (S157,
RS-20260811b-realia-channel), §7.6 (S162,RS-20260811h-domestication-channel), §7.9 (S167,RS-20260812e-dose), §7.10 (S172,RS-20260813a-world-dose) — all four read whole and consolidated; none was withdrawn or scoped by a later section, each explicitly builds on and extends the one before it. RS-20260811b-realia-channel§9 (limits),RS-20260811h-domestication-channel§6–§7 (limits and defects),RS-20260812e-dose§6–§7 (limits, the 30.6%-waste dispatch defect),RS-20260813a-world-dose§6–§7 (limits, the proper-name reversal's own provenance as post hoc).A-shaw-spider-thread(Tier 1 precedent anchor, JA→EN, S012, second-reader-verified S015) andA-garnett-vanka(Tier 1 precedent anchor, RU→EN, S013, second-reader-verified S015) — both read whole for the realia-handling catalogues in §2 and §3; their other findings (onstyle-correspondence,voice, footing) belong to other families and are not consolidated here.wiki/goodness-senses.md§cultural-mediation— the sense's own change log carries a synthesis of the same four experiments and the strategy-set additions (scaffold, transparent calque, measure conversion, re-foreignisation, preserved-inert, added realia); this entry draws its §2 table and regularities from the underlying result pages directly rather than restating that log.