Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: framework/v0.3/entries/HB-emphasis-and-typography.md · rendered 2026-09-09

Page metadata (front matter)
typeentry
idHB-emphasis-and-typography
statusdraft
created2026-09-09
updated2026-09-09
sensesperceived-source-carriage, style-correspondence
pairsRU→EN, RU→DE, EN→FR, EN→JA, EN→RU, SV→EN
provisionaltrue
internal-judgment-onlytrue
linksframework/v0.3/README.md, framework/v0.2/README.md, wiki/findings/results/RS-20260812-unlicensed-typography.md, wiki/findings/results/RS-20260812b-emphasis-carriage.md, wiki/findings/results/RS-20260812g-foreign-italic.md, wiki/findings/results/RS-20260811c-source-beliefs.md, wiki/base/anchors/A-poe-emphasis-carriage/A-poe-emphasis-carriage.md, wiki/arms/ARM-emphasis-carriage.md, wiki/arms/ARM-source-beliefs.md, workshop/translations/korotenkiy-roman/R06-v1/translation.md, workshop/translations/black-cat/R04-v3/translation.md, workshop/translations/william-wilson/R06-v1/translation.md, workshop/translations/belfry/R06-v1/translation.md, wiki/goodness-senses.md, wiki/archive/method-notes-S032-S243.md

Marks the reader will read as the author's: punctuation and emphasis the source did not license

Standing. Written at the project's close (2026-09-09, S257) as a consolidation of the record without the translation limb the procedure's step 6 requires — no fresh passage was translated under this entry, so it stays status: draft and its §5 Application reads "not applied". The evidence is four arithmetic censuses in three runs over stored public-domain texts (no jury in any), one reader-side lure experiment on three prompted non-Anthropic seats, and four lead translations read as self-audits; the correspondence coding is the lead's (internal-judgment-only) except the emphasis type coding (blind, 65 of 68) and the 25 Russian foreign-word calls (25 of 25 on carried/not). Tier D is NOT PASSED and is EXHAUSTED (config/models.md, RS-20260906-tierD-verdict-v3): evidence classes X1a, X1b and X2 are used, X3 is not, and nothing here says one rendering is better than another on a jury's word. Consolidated from v0.2 §7.7, §7.8, §7.8a.

1. The problem

Every language boundary changes the marks: a Russian dialogue dash becomes an English quotation mark; suspension points fall to a house style; a word italicised because it is foreign to English lands in a script that already says "foreign", or in a language where it is not foreign; small capitals have no French or Japanese counterpart; French italicises titles by rule. The reader of the translation reads every mark as the author's — on the one measurement made, an unlicensed mark with more confidence than a licensed one. So the translator must decide, mark by mark, which are the author's, which a convention's on either side, and which their own; the record finds that attention alone cannot see this, that a count can, and that the count misreads two classes unless the source's devices are inventoried first.

2. What published translators do

pair · material hands measured finding source
RU→EN, RU→DE · Gogol «Шинель» (1842) whole, 10,023 RU words Hapgood 1886, Field 1916, Kassner 1911 (German) marginal totals of ! ? … — ; per 1,000 words; no alignment; descriptive only Gogol's 21 suspension points reach the reader as 0 in Hapgood and 0 in Field; Kassner keeps 19. Field 4.99 dashes/1,000 vs Hapgood 3.12; Hapgood 10.03 semicolons/1,000 vs Field 7.63; both English hands' ! and ? rates sit above the Russian's (2.89/2.89 → 3.29/3.95, 3.00/3.36) RS-20260812-unlicensed-typography §4
EN→FR, EN→JA · Poe, seven tales, 241 prose paragraphs, 74 source emphasis spans (36 STRESS, 25 FOREIGN, 6 SMALLCAP, 3 TITLE, 4 TERM), 244 target spans Baudelaire 1857 (italics), Sasaki 1930s (傍点) span-level correspondence, both directions reader's side: licensed 0.3333 of 126 marks (FR), 0.3305 of 118 (JA); without the one exposed tale 0.2453 / 0.2245. Source's side: STRESS retained 0.833 / 0.944; FOREIGN 0.320 / 0.000; SMALLCAP 0.167 / 0.500; TITLE 1.000 / 0.333; TERM 0.000 / 0.250; all 74: 0.568 / 0.527 RS-20260812b-emphasis-carriage §4; A-poe-emphasis-carriage
EN→RU · the identical 74 spans («Belfry»'s 7 unmeasurable: its Russian witness carries no mark at all) Balmont, Скорпион 1901–1912 same; the 25 FOREIGN calls blind triple-coded (25/25 on carried/not) STRESS 35 of 35; FOREIGN 0.368 on six tales (0.280 over all seven tales); Latin-script renderings 0 of 9 keep the mark, Cyrillic renderings 7 of 10; licensed 0.4414 of 111 marks, which is 0.8723 (1901 vol. I, 47 marks) / 0.1034 (1906 vol. VII, 58 marks) / 0.3333 (6 marks) by copy-text RS-20260812g-foreign-italic §3–§5

The regularities, stated as what the hands did:

  1. A whole class of marks can vanish from a translation without any decision being recoverable from the texts. Gogol's twenty-one suspension points reach zero in both English hands and nineteen in the German; translators' choice or publishers' style cannot be told apart, and to the reader it makes no difference. The marginals license one narrow claim: three hands of one source produce different amounts of expressive typography, so the amount the reader sees is not fixed by the source (RS-20260812 §4).
  2. Emphasis is not dropped in translation; it is flooded. About a third of the emphasis marks two published hands' readers see are licensed by any emphasis in Poe — two hands, two centuries, two devices, the same figure to two decimals (0.3333, 0.3305) — against 0.9425 for punctuation on the Russian pair (RS-20260812b §4.1). The agreement is pooled: per-tale licence runs 0.053–1.000 in French, 0.070–0.867 in Japanese; Sasaki adds 2 marks in all of «Usher» and 40 in «William Wilson» (§4.3).
  3. The source's stress arrives; what floods arrives beside it. Where Poe leans on an English word the hands lean with him — 0.833, 0.944, 1.000 over the 36 (35) stress italics. The unlicensed marks are titles italicised by French rule (nine in «Usher»'s library paragraph, twelve of Baudelaire's 84 additions in one paragraph), English and slang set as foreign, and stress the source lacks (RS-20260812b §4.4).
  4. A foreign-word italic is carried by function, and the writing system can do the carrying. Baudelaire keeps the eight Latin and German words still foreign in French and drops every French one; Sasaki keeps none of the twenty-five, because katakana already says "foreign" (RS-20260812b §4.2). Within one hand the same trade shows at span grain (Balmont's 0 of 9 Latin-script against 7 of 10 Cyrillic, table) — a decision about foreignness, not a distaste for italics (RS-20260812g §3–§4). This is a within-hand contrast; the generalisation to "the target language" was withdrawn (§7).
  5. A source may emphasise in a device the target lacks, and the hands then convert. Poe's small capitals (PERVERSENESS, GALLOWS, SILENCE) have no French or Japanese device; Sasaki converts three of six into 傍点 (0.500; the page names two by word), Baudelaire one into italics, Balmont two. A census restricted to italics scores every conversion as an unlicensed addition (RS-20260812b §3; note (bmk)).
  6. A single hand's "licensed" figure can be an average over copy-texts that have nothing to do with each other. Balmont's 0.8723 against 0.1034 splits on the printed volume, a factor of 8.4, with stress retention 1.000 in both and an independent re-keying carrying 27 of 28 body italics — not lost markup. Practice changed or compositors differed; the run cannot say (RS-20260812g §5).
  7. The marked word's lexical status in the target does not predict whether its mark survives. Registered: still-foreign words retained more than naturalised ones; observed 0.300 against 0.444, widening the wrong way at lexeme level (0.300 vs 0.571) (RS-20260812g §2).

3. What this project's own practice found

  1. The demand side, and why the count matters at all (RS-20260811c-source-beliefs, SV→EN, Hansson «Sensitiva amorosa» IX, 925 words, six segments, three arms, three blind non-Anthropic seats answering thirteen statements about the Swedish per segment, scored by script). An arm propositionally identical to a flattened rendering (certified 6 of 6 by a parity seat) but carrying four markedness operators the Swedish licenses nowhere — one exclamation mark, two italicised words, one dash, three sentences begun And — scored 0.1017 on statements those marks bear on (a coin scores 0.5); difference from the flattened arm +0.9537, 6 of 6 segments, exact P = 0.015625; confidence on the wrong answers 5.887 of 6 against 2.500. The lure is graded: exclamation mark and italics moved every seat every time (0.00), the dash two thirds of the time (0.33), the And-run 0.08 — typography lured absolutely; syntax lured less (§5c). The sense-level consequence sentence is in wiki/goodness-senses.md §perceived-source-carriage.
  2. A careful single pass supplies almost nothing deliberately (T-korotenkiy-roman-R06-v1, Garshin «Очень коротенький роман» whole, RU→EN, 1,495 → 2,013 words, expansion 1.3465, 32 paragraphs aligned by construction; RS-20260812 §2). Of 97 source marks, 87 in the English; 58 of 58 exclamation marks, question marks and suspension points retained in the aligned paragraph and none added; dashes 25 → 13, all fifteen omissions in the fourteen paragraphs where the source has direct speech and none in the eighteen narrative ones — the Russian dialogue dash becoming an English quotation mark; semicolons 14 → 16. Licensed share 0.9425, at paragraph grain, so an upper bound. Contamination: not measurable (no English rendering reachable), carried as none for want of another value.
  3. The translator's frozen prediction about his own marks was wrong by the sign on exactly the convention-driven mark (RS-20260812 §3; log D11). Predicted: dashes up on the source's, "perhaps by a third"; found: −48% absolute. The three marks he kept he predicted to 0.00% error. Under a per-1,000-word convention every prediction fails, because he predicted marks without predicting length. No verdict is called: the scoring convention was fixed after the totals were seen.
  4. The self-audit in the flooded channel (T-black-cat-R04-v3, EN→JA, «The Black Cat» whole in one hand; RS-20260812b §5). 17 marks, 17 licensed, 0 added; retention 16 of 19 italic sites and 1 of 2 small-capital sites. The half written before anyone had asked about emphasis (v1, v2) is 13 marks, 13 licensed, 0 added. Sealed prediction: over 70% of italic sites carried — holds (0.842); no added mark — holds; every dropped mark FOREIGN-typed — fails (_ammonia_ is TERM, PERVERSENESS a small capital); total within two of Poe's — holds on italics (17 vs 19), fails on all emphasis (17 vs 21). The log's reason for dropping _barroques_: the italic says foreign, katakana was already saying it. contamination: high on measurement (runs of 16 and 17 characters shared with Sasaki against reference cells of 8 and 9): a self-audit, never an independent hand.
  5. Two hands, opposite conventions, on the same fourteen decisions (T-william-wilson-R06-v1, T-belfry-R06-v1, EN→RU, 1,533 source words, every paragraph carrying a FOREIGN span, logs frozen before any Balmont was read; RS-20260812g §8, process documentation, not evidence). The lead carried 6 of 6 still-foreign words with mark, 0 of 8 naturalised ones, licensed 1.000, no added mark — and on the fourteen shared sites matched Balmont's outcome 4 times. The lead keeps a Latin-script word and its italic and drops the italic on a Cyrillic rendering; Balmont does the reverse. The lead's log reason — "the word arrives still foreign, so the mark still has work" — stands against the only published hand in the corpus, and the lead cannot show which is right. Contamination suspected, near the floor: longest runs 9 and 7 tokens (the 9 is Poe's own French sentence), 7-grams 6 and 1 against reference cells of 3 and 3 tokens with 0 shared 7-grams.
  6. Across both self-audits the translator was right about what he kept and wrong about what he dropped or what moved by convention (RS-20260812b §5).
  7. Two method rules were written on this material (wiki/archive/method-notes-S032-S243.md): (bcg) — a silent extractor produced a Sasaki text with every character and no 傍点, which read as "the translator dropped Poe's emphasis", the opposite of the truth; count the markup in the raw source before reading for form. (bmk) — inventory every device the source uses for emphasis before counting one of them.

4. The options

The evidence sorts a translation's marks into kinds that behave differently:

The instrument is the count — source's marks against your own, by kind, per copy-text — and it diagnoses without prescribing: what to do about a difference it surfaces is not evidenced anywhere in this family (§7.7).

5. Guidance

For a translator

  1. Before translating, inventory every emphasis device the source uses — italics, small capitals, spaced type, capitalisation, 傍点, quotation — and say which of them your copy-text encodes and which your own page can carry. A count that watches one device scores every conversion as an invention. — evidenced (EN→FR, EN→JA, EN→RU), method rule (bmk).
  2. After translating, count the source's expressive marks and your own, by kind, and look at the difference. The difference is where the changes you did not decide are. It diagnoses; nothing evidenced tells you what to do about it. — evidenced (RU→EN; EN→JA and EN→RU as self-audits).
  3. Do not trust your own account of what you did with the marks. In two self-audits the translator was right about what he kept and wrong by the sign about what moved by convention, and wrong about which marks he had dropped. — evidenced (RU→EN, EN→JA).
  4. Where you keep a foreign word in a foreign script, the italic says twice what the script says once (the source page's inference, internal-judgment-only); where you naturalise the word, the italic is the only surviving trace of the author's pointing. Drop in the first case if you choose, keep in the second — and know that the count records the first as a loss and cannot see the second. — evidenced (EN→RU as a within-hand contrast; EN→JA, EN→FR consistent); language-or-hand untested; the lead's own opposite convention stands unresolved.
  5. Where the source emphasises in a device your language lacks, convert into your own device and log it as a conversion. — evidenced as an observed move (EN→FR, EN→JA, EN→RU; 0.167–0.500); the instruction itself untested; logging per (bmk).
  6. Do not add a stress, exclamation mark or italic the source does not license unless you accept that the reader will attribute it to the author, and with more confidence than a licensed mark. — evidenced (SV→EN, three LLM seats; no human reader).
  7. Decide the convention-driven classes once, as a policy, and state it — dialogue dash, suspension points, house-rule italics. Whole classes vanish from the published record with no recoverable decision; a stated policy is the only way a reader could know it was yours. — evidenced as description (RU→EN, RU→DE, EN→FR); the choice itself untested.
  8. State your copy-text, on both sides. The source's rate is an editor's rate as much as an author's, and one hand's licensed rate split by a factor of 8.4 across two volumes. — evidenced (EN→RU; RU copy-texts).

For a pipeline

  1. Enumerate the source's marks from the raw markup, never from a stripping extractor, under a device inventory declared before counting (the census code under the two E-20260812* experiment directories does this). — untested as a pipeline step; executed only inside experiments.
  2. Classify each source mark: STRESS / FOREIGN / TITLE / TERM / SMALLCAP for emphasis (the italic typology blind-coded at 65 of 68; SMALLCAP inventoried by the lead); expressive (! ? …) against convention-bound (dialogue dash, semicolon) for punctuation. — untested.
  3. Set the policy parameters: (a) the target convention for dialogue punctuation and suspension points; (b) the foreign-word channel — script, mark, or both; (c) the conversion device for small capitals; (d) house-rule italics on or off. — untested.
  4. Render, then run the aligned correspondence census — span-level for emphasis, paragraph-aligned with retained = min(src, tgt) for punctuation — reporting licensed share, retention and additions by class, per copy-text, with retention flagged as an upper bound at paragraph grain. — untested.
  5. Check: flag every added expressive mark and every dropped STRESS mark; flag a FOREIGN drop only where the word was naturalised; flag a SMALLCAP conversion as a conversion, not an addition. — untested.

Human entry points. Step 3's four policies are a person's to set — choices about the target's conventions and the reader, not facts the evidence decides. If nobody sets them, default to what practice reached without a rule: keep the source's expressive marks by kind (58 of 58 under R06, uninstructed), follow the target's dialogue convention, drop the foreign-word italic where the script carries foreignness and keep it where the word is naturalised, convert small capitals into the target's device. Step 5's flags need a person: the count diagnoses; nothing here licenses an automatic correction.

Application. Not applied: written at close-out without a translation limb.

6. Not evidenced, and open

7. Sources consumed