Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260727-log-decision-coding/codes.md · rendered 2026-09-09

Page metadata (front matter)
typetypology
idlog-decision-codes
statusdraft
created2026-07-27
updated2026-07-27
linkswiki/goodness-senses.md, wiki/arms/ARM-typology-logs.md, workshop/experiments/E-20260727-log-decision-coding/decisions.tsv, wiki/findings/results/RS-20260727-log-typology.md
internal-judgment-onlytrue

Fourteen classes of translation decision, open-coded from twenty frozen logs

What this is. An open coding of the 361 decisions recorded in the project's twenty frozen translator's logs (decisions.tsv), derived from the material and not from wiki/goodness-senses.md. ARM-typology-logs step 2 requires that ordering: coding the decisions against the nine senses would guarantee they fit, which is how a list survives sixteen sessions without being tested.

How the classes were arrived at. All twenty logs were read end to end in one sitting, before any class existed. Classes were written down as recurring shapes of pressure — what made a decision necessary — and merged or split until every decision in the corpus fell into one. Four passes; the final split was C4 into C4a/C4b, forced by the observation that "the source has a category the target lacks" and "the target obligatorily has a category the source lacks" are opposite conditions with opposite consequences and had been sharing a class.

The coding is on the pressure axis, not the handling axis. A second, orthogonal coding is available and was not done — retain / gloss / substitute / calque / scaffold / flatten / compensate, the strategy set wiki/goodness-senses.md already carries under cultural-mediation. That set describes what the translator did; these classes describe what made a decision necessary. The pressure axis was chosen because the arm's question is whether the senses cover what the logs record, and the senses are dimensions along which a translation is judged — closer to pressures than to handlings.

One coder. Everything below was coded by the lead. E-20260727-log-decision-coding/design.md specifies the independent blind re-coding that tests whether the scheme is shared or idiosyncratic; RS-20260727-log-typology §4 reports what it returned.


The classes

C1 — Lexical gap. No target word covers the source word's sense, or covers it at the right register. Includes false friends (the available cognate means something else), one-source-word-to-several-target-words and the reverse, and coinages with no target equivalent. Examples: Biegsamkeit → "pliancy"; déperdition → "bleeding-away"; удаль → "a kind of dash"; déplorable refused as a false friend; iuvenaliter, for which nothing was found.

C2 — Source-internal repetition. A word, root, phrase or half-line recurs in the source and the recurrence is doing work; the decision is whether to hold one target word across all its sites at a cost in local fluency. Examples: the сид- network across seven sites in «Пари»; うるはし three times in Genji; flet four times in Beowulf; «интересный» four times in «После театра».

C3 — Figure bound to the signifier. Pun, sound-play, rhyme, etymological figure, chiasmus, a word active in two senses at once, a riddle turning on grammatical gender: the effect lives in the form of the word and not in its sense. Examples: Verbrecher/Brecher; шишки (pine-cone / bump); a virgine virgine rapta; giornata as both day and day's wage.

C4a — Category the target lacks. The source marks something grammatically or morphologically that the target has no category for: T/V deixis, honorification, diminutive affection, grammatical gender carrying meaning, middle voice, aspect, formulaic self-abasement. The information must be transcoded into lexis or lost. Examples: the entire honorific layer of Genji; the formal вы in «Роза», where nothing was done and the loss is total; 不佞 → plain "I".

C4b — Determinacy the target compels. The mirror image, and the reason the class was split. The target's grammar obligatorily encodes something the source left open, so the translation states what the source did not: articles, tense, number, subject pronouns, a disambiguation between two live readings. The translator has no option to abstain. Examples: wīf → "woman" at one site and "wife" at another, where Old English says neither; English determiners supplied on nearly every count noun in Beowulf; classical Chinese tense and pronouns supplied "on nearly every sentence"; «лета» read as governing both nouns, closing an ambiguity the Russian keeps open.

C5 — Culture-bound referent. An object, institution, measure, garment, food, place or practice with no target counterpart, or with one that misdescribes it. Examples: торбаса kept italic and unglossed; вершок domesticated to inches; areau → "swing-plough of ancient pattern"; 数珠 → "prayer beads", with "rosary" refused as a Christian object.

C6 — Fixed expression. Proverb, idiom, formula, canonical tag, allusion-by-quotation: meaning is not compositional and the target has a different stock. Examples: the two inverted Makar proverbs; 災棗梨; избрал благую часть rendered with the KJV wording; с пивной котел kept literal with "as they say" intact.

C7 — Register placement. Where in the target's register space the whole text, a passage or a speaker sits — including the decision to archaise or not. Examples: Verga's diglossia answered with plain unliterary English; no archaising in Genji or Beowulf; Yan Fu pitched formal and paratactic; the Bécquer span pitched at 1860s literary English.

C8 — Syntactic architecture. Period length, suspension, apposition, inversion, word order at the sentence and paragraph scale: whether to reproduce the source's shape. Examples: Schleiermacher's 90-word periods kept whole; Turgenev's 60-word suspended apposition broken; the bird simile front-loaded in Ovid; prose chosen over alliterative verse for Beowulf.

C9 — Typographic and convention mismatch. The two writing systems mark different things: dialogue dashes vs quotation marks, guillemets, ellipsis habits, italics, emphasis, list numbering, capitalisation of titles. Examples: Russian dashes → English quotation marks in four separate logs; Nietzsche's punctuation kept at German positions; Yan Fu's seven repeated 一、 numbered 1–7.

C10 — Naming. Proper names, work-titles, epithets, patronymics, place-names: transliterate, calque, substitute, withhold. Examples: Sidonis withheld exactly as Ovid withholds it; three Genji tale-titles handled three ways in one sentence; S—v Lane; La Peña kept uncalqued at both sites.

C11 — Source-text uncertainty. Which source text, and what to do when it is damaged, variant, self-inconsistent or unparsed. Examples: Oft nō seldan hwǣr translated as the edition prints it; the source spelling one place two ways in «Jeli» and the translator regularising; Илюша/Ильюша; four classes of OCR artifact normalised in Schleiermacher; 斧落徽引, a figure the translator could not fully parse and translated anyway.

C12 — Excerpt artefact. The decision exists only because the passage was cut out of a larger text — supplying an antecedent that lies outside the window, rendering a fragment as a fragment, omitting the source's own examples. Examples: celui d'Holbein → "Holbein's ploughman"; seven of ten Ovid windows opening or closing mid-period.

C13 — Prior-rendering pressure. The decision was made in known sight of another translation, or deliberately positioned with respect to one. Examples: Schleiermacher's famous sentence, which "could not be translated in ignorance of the received English"; «институт» kept against a reviewer's stated preference the translator had read; Beowulf pitched plain because a comparison text was known to be archaising; magno stant Danais noticed as obvious English and deliberately not changed.

C14 — Reader-knowledge management. How much the target reader is told: scaffold, gloss in the running text, footnote, or withhold — where the item is not culture-bound and the decision is about the reader's experience rather than about the referent. Examples: Yan Fu's four coinages translated literally and deliberately not identified as "natural selection" and the rest; iam tutus left unglossed; Elend/êlend left in German because the sentence says this is something a German understands.


Mapping onto the nine senses

Applied at class level, not row level: a per-row mapping would be several hundred unverifiable judgment calls, whereas this is fourteen judgments, stated here and disputable here. analyse.py reads this table from its own MAPPING constant; change both together.

class senses verdict why
C1 accuracy, style-correspondence clean The sense list's core case. accuracy covers "correctness at the word level"; where the gap is registral, style-correspondence takes it.
C2 style-correspondence clean Its definition names repetition explicitly as a marked formal feature.
C3 style-correspondence clean Its definition names sound play and "orthographic play".
C4a style-correspondence, cultural-mediation clean, with an overlap worth noting Both senses claim politeness deixis: cultural-mediation lists "honorifics, politeness deixis" among culture-bound items, and style-correspondence's S013 grounding note discusses Russian T/V and diminutive morphology at length as its evidence. Same phenomenon, two senses, neither deferring.
C4b — residue See below.
C5 cultural-mediation clean The sense was written for exactly this.
C6 cultural-mediation, naturalness clean "allusion" is named under cultural-mediation; the target-idiom half is naturalness.
C7 naturalness, purpose-fit clean naturalness is already a "measurable register variable" per OQ-20260723-target-register; a declared purpose re-weights it.
C8 style-correspondence clean "sentence shape" is named.
C9 style-correspondence stretch Its definition covers typographic play — a marked feature a translator may flatten. A convention mismatch is not a marked feature: T-pari-R04-v1 §2g observes that every English rendering must resolve the Russian dash the same way, "which makes it useless as a discriminator between renderings". A sense exists to discriminate.
C10 cultural-mediation clean "names" is named.
C11 — residue See below.
C12 — residue See below.
C13 — residue See below.
C14 cultural-mediation, purpose-fit stretch Clean where the withheld item is culture-bound. Not where it is not: 天擇 is natural selection, a concept English holds perfectly well, and the decision to withhold the English term is about preserving the reader's experience of the source's difficulty. No sense names that.

What the residue has in common

The four residue classes are not four unrelated gaps. Every one of them is a decision that is not about the relation between a determinate source and a freely made target.

C11, C12 and C13 arguably should have no row: a typology of good is not obliged to name conditions of production. That is a finding about what the logs record, not a defect in the list, and it is reported as such.

C4b is different, and it is the one thing here that looks like a missing sense. It is a property of the finished text — that the translation is more determinate than the source it renders — a translator can work at it (Ovid's withheld Sidonis; the duguða biwenede reading chosen to keep the singular focus), and a reader can be worse off for it. TH-20260724-translation-distance-axes C1 already describes the mechanism ("meaning carried by grammar is not translated but transcoded into lexis"); what the project has never had is a sense under which the resulting over-determination is a cost. Proposed as D-20260727-08, which this session may not ratify.