Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: framework/v0.3/entries/HB-register.md · rendered 2026-09-09

Page metadata (front matter)
typeentry
idHB-register
statusactive
created2026-09-07
updated2026-09-07
sensesnaturalness, perceived-source-carriage, voice, style-correspondence
pairsDA→EN, JA→EN, RU→EN, IT→EN
provisionaltrue
internal-judgment-onlytrue
linksframework/v0.3/README.md, framework/v0.2/README.md, wiki/findings/results/RS-20260809g-device-cross.md, wiki/findings/results/RS-20260813g-register-quadrants.md, wiki/findings/results/RS-20260811-floor.md, wiki/findings/results/RS-20260810c-register-reach.md, wiki/findings/results/RS-20260810z-idiom-reach.md, wiki/findings/results/RS-20260814d-elevation-resolution.md, wiki/findings/results/RS-20260815-register-room.md, wiki/findings/results/RS-20260815b-fluent-carriage.md, wiki/findings/results/RS-20260815e-fluent-carriage-2.md, wiki/base/anchors/A-mchugh-presence/A-mchugh-presence.md, wiki/base/anchors/A-english-tale-register/A-english-tale-register.md, workshop/translations/malavoglia-i-conscription/R04-v1/translation.md, workshop/canon/malavoglia-i-conscription/manifest.md, wiki/plan.md

Where the source sits above or below its own neutral register, and English cannot decline to choose a place

Standing. Two published-hand or published-corpus measurements (an Andersen compounding census and a twelve-text nation-marking floor), five internal experiments on three language pairs (Danish→EN, Japanese→EN twice, Russian→EN), and one fresh translation limb (Italian→EN, this session), all internal-judgment-only except the published-corpus counts, which are mechanical. Tier D is NOT PASSED and is now EXHAUSTED (config/models.md): nothing here rests on a jury, and naturalness's three period/idiom-indexed register anchors are evidence about English, not scoring licences a translation can be judged against. Consolidated 2026-09-07 (S252) from v0.2 §7 (Q-e), §7.1, §7.2, §7.3, both §7.4 headings (a numbering defect: S156 and S178 share the number), §7.11, §7.13.

1. The problem

A source can sit above or below its own neutral written register at a given point — Andersen's narrator drops into a compounding, homely Danish; Sōseki's narration and Chekhov's peasant speech mark themselves off from the surrounding prose; Verga's narrator's voice merges with a Sicilian village's mocking, self-interested talk. Three different decisions hide inside "carry the register": elevation (does the English rise or fall where the source does, and by how much), placelessness (can a translator mark a register shift without also marking a nation, a region, or a period the source did not intend), and period (what point on English's own historical and generic register spectrum — Victorian period-idiom, contemporary vernacular, unmarked literary-contemporary, a genre's own purpose-register — does the translation actually land on, and does the translator know what that point is made of). This entry's evidence answers the second question with the most force and refuses a general answer to the first four times over before finding out why.

2. What published translators do

pair · material hands what was measured finding source
DA→EN · Andersen, «Flipperne» (word-formation, a source-ward register feature) 3 published (an unattributed Gutenberg #1597 hand, Mrs H. B. Paull 1888, H. L. Brækstad 1900) against 3 lead renderings (no rule; ennoblement; an explicit compound-retention rule) 16 frozen Danish compounds, two blind coders agreeing at 139/144 all three published hands carry exactly 3 of 16; the lead's unruled pass already exceeds them at 6/16; ennoblement carries 4/16; the explicit rule reaches 16/16 RS-20260813g-register-quadrants
EN (originals) · 12 published narrative texts, 1843–1930, of documented national provenance 12 hands (Russian-translation-heavy, plus originals) an orthographic nation-mark inventory (spelling pairs like colour/color), frozen before any text was read published English narration carries 1.24–5.33 marks per 1,000 words; first mark at a median of word 627; 10 of 12 texts entirely one national convention once marked (purity 1.00) RS-20260811-floor

The regularities:

  1. No published hand goes toward a source's marked register at anywhere near the rate an explicit rule can reach, and the published hands agree with each other more than any of them agrees with the source. Three independent Andersen translators, forty years and one continent apart, all sit at exactly 3 of 16 retained compounds; a rendering told in so many words to retain them reaches 16 of 16. A framework line of the form "where the source compounds, compound" would be a recommendation against the entire published record on this pair, not a distillation of it. (§7.4, S178.)
  2. Ordinary published English narration is never register- or nation-mark-free, and it does not commit early. All twelve texts eventually carry a national spelling convention, but the median first mark falls past word 600, and once a text commits it almost never mixes conventions. "Placeless English" is not a book with no marks; no such book was found. It is a book that delays, then commits, then stays committed. (§7.4, S156.)

3. What this project's own practice found

The lead's attempts to write a general register-carriage recommendation refused four times in sequence, each refusal on new, more specific grounds — a chain worth reading as a chain, because each entry believed the previous obstacle was the last one:

  1. First attempt (RS-20260809g-device-cross, §7.1): asked whether class-marked idiom or nonstandard spelling costs a register drop more, on Italian and Japanese narration. The comparison itself was withheld for want of power — but a side finding was licensed and it reversed the working assumption: nonstandard spelling is not a placeless device. Two independent judges placed a respelling-only rendering more often than a rendering using every located idiom available (0.57 and 0.54 against a 0.25 bar), and one named the marker unprompted: "an' (US Southern/rural)." A translator cannot adopt eye-dialect "without relocating a book"; the commonest English respelling reads as somewhere, not as nowhere.
  2. Second attempt (RS-20260810c-register-reach, §7.2): built the larger cell the first attempt said it needed (30 Japanese sentences, 8 arms, 3 seats). The comparison died a second time, now by the manipulation itself: given explicit written permission to use located idiom and class-marked grammar, two independent hands changed their own English at 4 of 60 opportunities (0.067); given permission to respell, at 32 of 60 (0.533). The two "ways down" are not two options a translator weighs — on this material, under minimal revision, one is barely taken and the other is taken half the time.
  3. Third attempt (RS-20260810z-idiom-reach, §7.3): asked whether the located idiom was unavailable in the material or merely unreachable by minimal revision, by writing it from the source instead. Refused a third time, and the obstacle moved into the instrument's own baseline: two independent judges placed the "placeless" arm itself — written under a rule whose entire content was no located lexical item, idiom, or grammatical form — at 0.03 and 0.20, on words its own rule did not catch (apologise, trodden, for a song, the state of me, had it brought home to me, give myself airs). There is no placeless English to measure a located device against; every earlier figure in this chain is a difference measured from a floor that was never neutral, and one of the leaking markers was a spelling family (-ise/-ize) with no neutral form at all.
  4. The floor itself, measured (RS-20260811-floor, §7.4/S156): §3's headline — "there is no placeless English" — is withdrawn as overstated and replaced with a number: a rule instructing no located lexis or grammar produced a text with 0.51 orthographic nation-marks per 1,000 words, below the least-marked of twelve published texts (1.24), and the one leak (apologise) was the single family an independent pre-run critic had flagged in advance as historically unsafe; removing that family took the count to zero. The claim survives narrowed: country-level orthography can be nearly emptied by a rule that watches spelling specifically; class, region and period markers (trodden, for a song, the state of me) are untouched by that same rule and were never claimed to be. A translator's ordinary attention does not do this on its own — a single-pass rendering with no register rule at all, by the same translator, carried 3.92 marks per 1,000 words, and thirteen nation-choices its own translator noticed making (a mechanical detector found six; the other seven — carriage, up train, backwards, conductor, porter, should, mandarins — are words no wordlist watches).
  5. A register policy's audible work is grammar and discrete tokens, not volume of rewriting (RS-20260815-register-room, §7.11, Chekhov «Злоумышленник», three lead renderings, three model seats — ARM-elevation-resolution step 2, after step 1, RS-20260814d-elevation-resolution on Andersen, failed its own calibration gate and withheld everything, at ceiling ambiguity: the gate passed at +1.00 in 56 of 56 cells while the quantities it was meant to separate sat at +0.125 to +0.70, so nothing about the interior of the register scale could be read. Step 2 fixed that by adding a third, deliberately intermediate rung.) A light, one-step register-raising rule read as raising the source most at its shortest low-register spans (+0.75 at one-to-six-word lines, +0.50 at 26–60-word speeches) — the opposite of "more text, more room to raise." One corrected pronoun case (Us, the people… to We, the people…, one token in fourteen) moved every reading seat a full point on its own; a census of contractions, logged before any reading existed, was independently named by the blind seats. A full ennoblement rule, in contrast, raised the passage everywhere uniformly — including ordinary, unmarked narration — because its rule set redefines the neutral itself rather than nudging the low points.
  6. A source-blind reader's judgment of "this follows the source" is substantially a judgment of how translated the English sounds (RS-20260815b-fluent-carriage, RS-20260815e, §7.13, Japanese→English, two independent builds). An arm that carries nothing and is simply worse English was chosen as the more source-following option over a clean flattening at 20 of 21 forced comparisons; set against genuine non-fluency, the same seats could not reliably tell carriage from clumsiness (11–13 of 21 indeterminate, across both builds). Shown the source alongside, the same pair, the same seats, the preference reversed to the carrying arm at 20 of
  7. A craft claim from the first build — "carrying a source's form does not require paying in fluency" — was withdrawn on the second, cleaner build: readers there actually preferred the flattened arm as English (15 of 21, a lean, not significant).
  8. This session (T-malavoglia-i-conscription-R04-v1, Verga's I Malavoglia, ch. I, 468 Italian words): applying item 4's finding directly — no eye-dialect or nonstandard spelling anywhere, register carried by word choice and syntax alone — the lead enumerated four kinds of register-marked site (extended free-indirect village mockery, a gnomic authorial aside, two adverbial reduplications, one directly quoted plain line) before drafting, per §5 item 1 below, and reports that naming the no-spelling constraint once removed it as a live choice at every later site — a followability observation, not a measurement. One device (adverbial reduplication) survived translation once (zitti zitti → "silent, silent") and was abandoned once in the same passage (mogi mogi → "crestfallen and shamefaced," not "downcast, downcast") on the craft judgment that a short adjective doubles as rhythm and a longer one doubles as a stutter — recorded, not measured. The single contraction in the whole rendering was reserved for the one directly quoted line, applying item 5's finding that a discrete token, not volume, carries a register signal.

4. The options

The evidence refuses a single register-carriage recommendation and instead separates the decision into pieces with their own, narrower answers:

5. Guidance

For a translator

  1. Enumerate every register-marked site in the source before opening any other translation, and decide your no-spelling-or-not policy once, in the abstract, before looking at a single site. This session's practice found that deciding the constraint first removed it from every later site's deliberation. — evidenced (IT→EN), craft observation.
  2. Do not treat nonstandard spelling as a placeless way to lower register. It is the device most readily taken and the one most reliably read as a specific place; if you use it, expect the passage to be read as regionally or nationally located, on evidence from two languages and three independent runs. — evidenced (JA→EN twice, IT→EN).
  3. Do not expect located idiom or class-marked grammar to be reached by minimal revision even under explicit permission — two independent hands used it at 1 in 15 opportunities where they used respelling at 1 in 2. If the register drop matters, it likely needs writing from the source rather than editing a placeless draft toward it. — evidenced (JA→EN).
  4. There is no neutral spelling or vocabulary to retreat to. A rule against locating idiom is not a rule against locating spelling, and a translator's unaided attention catches under half of the nation-choices a rule-following pass still leaves in a text. Decide your national convention deliberately, not by default. — evidenced (RU→EN, one translator's practice).
  5. To raise a register at the source's own low points cheaply, use a small number of discrete, countable tokens (a grammatical form, a contraction pattern) at those sites rather than rewriting volume. A single corrected token can move a reading further than a quarter of a passage rewritten. — evidenced (RU→EN, one story, three model seats).
  6. Do not expect a device that carries the source's form to read as fidelity unless the reader can see the source. A blind reader mostly reads it as bad English; if the point is to signal fidelity to a bilingual or facing-page reader, the device works; if the point is craft for a monolingual reader, it does not pay for itself and may cost preference (a lean toward the flattened alternative was measured, not just an absence of gain). — evidenced (JA→EN, two builds).
  7. When aiming at "unmarked" or "invisible" prose, know what that register positively contains — the absence of period lexis, era slang, dialect spelling, archaism and Latinate display; an unsignposted free indirect discourse; plain say-attribution; sparse, unemphatic figuration; standard contractions in exposition — rather than treating it as a default with no cost. It is one indexed register among several, not the absence of one. — evidenced as anchor catalogue, untested for scoring (Tier D exhausted; these anchors ground description, not a jury target).
  8. Do not aim at a register as a purpose in itself. These catalogues describe what a register is; whether a given translation should occupy it is a purpose decision, and a target register is itself a target-culture norm — rendering non-English tale material in the idiom of English nursery collections is a domesticating move, not a neutral default. — descriptive caution, internal-judgment-only.

For a pipeline

  1. Enumerate: locate every point where the source's register departs from its own local neutral (up or down), separately from any decision about nation or locale. — untested.
  2. Classify each site's available device class (respelling/eye-dialect; located idiom/class grammar; discrete grammatical marker; wholesale rewriting) and flag that no class in this evidence is nation-neutral. — untested.
  3. Set the policy parameter: for a targeted raise, restrict the device to a small set of discrete tokens at the source's marked sites (the evidenced Chekhov effect); for a full register programme, expect uniform elevation, not concentration at marked sites, and treat that as a declared cost rather than a side effect. For a lowering, prefer located idiom/grammar over respelling unless a specific national placement is wanted, and expect low reachability from a minimal-revision procedure — write it from the source if the drop matters. — untested as a pipeline step; derived from evidenced findings.
  4. Render, then run a nation-mark check against the floor inventory (published English narration: 1.24–5.33 marks/1,000 words, first commitment by roughly word 600, near-total one-sidedness once committed) — a zero-mark target is not the goal and is not what published English does; a consistent, deliberately chosen convention is. — untested.
  5. Check: if a source-ward device is present, is the source visible to the intended reader? If not, flag that the device's fidelity payoff is unmeasured for that reader and may read as unfluent English instead. — untested.

Human entry points. Step 3's policy choice (targeted vs. wholesale; which national convention; whether a purpose-register like tale-telling is wanted) is a purpose decision belonging to declared-purpose (wiki/goodness-senses.md) and should route to a person. Without one, default to a targeted, discrete-marker policy at source-marked sites only (the cheaper, better-evidenced move) and no respelling-based lowering (the worst-evidenced-for-placelessness device).

Application. This session (S252) applied items 1, 2 and 5 to a fresh 468-word extent of Verga's I Malavoglia ch. I (T-malavoglia-i-conscription-R04-v1): no eye-dialect anywhere; register carried by word choice and syntax at four site types; a single reserved contraction at the passage's one quoted line. Followability cost: real but small — one device (reduplication) needed two different solutions in the same short passage, a craft finding the log states but does not measure; deciding the no-spelling policy in advance cost nothing further at any site once decided.

6. Not evidenced, and open

7. Sources consumed