Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260904-french-in-russian/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260904-french-in-russian
statusfrozen
created2026-09-04
updated2026-09-04
sensesstyle-correspondence, authorial-presence
internal-judgment-onlytrue
provisionaltrue
linkswiki/arms/ARM-french-in-russian.md, framework/v0.2/README.md, wiki/findings/results/RS-20260903-latin-in-italian.md, wiki/findings/results/RS-20260831-arabic-in-persian.md, workshop/experiments/E-20260904-french-in-russian/loci-frozen.md, workshop/translations/voina-i-mir/R04-v1/translation.md

E-20260904-french-in-russian — design, frozen before any English rendering of the work was opened

Frozen 2026-09-04 (S243), ARM-french-in-russian step 1. The 38 loci were enumerated and classified from the Russian (loci-frozen.md) before any of the four English volumes was fetched; the translation limb was drafted and frozen from the Russian alone. What had been read — or already known — of the English side at freeze time is declared in §7, and it is not nothing.

1. Question

framework/v0.2 §7.45 and §7.50 are the handbook's only sections about a source that changes language under a target that cannot. Both are evidenced on quotation: Arabic quoted inside Sa'di's Persian, Latin quoted inside Dante's Italian. In both, the embedded run is bounded, attributable, and announced — and §7.50.1 found that the option set depends on what the reader can read: where the reader knows the embedded tongue, nobody deletes, they english.

ARM-latin-in-italian closed naming what it could not reach: an embedded language the readership reads only partly. This design takes that case, and finds that it is not merely a third point on the same scale. Tolstoy's French is not quoted. It is code-switching: no source, no attribution, no announcement — a character's own speech, and in five of the 38 loci a sentence with Russian words standing inside the French. The question is therefore two-part:

Does the handbook's account of a language change survive the move from quotation to code-switching — and what do published hands do at a switch the source's own readers read only partly, and the target's readers read partly too?

Both readerships are partial here, in opposite ways. Tolstoy's 1860s Russian audience contained readers who read the French effortlessly and readers who did not, which is why the author glossed every one of the 38 loci himself. The Victorian and Edwardian English audience contained the same split. No hand in this set is translating for a readership that certainly can, or certainly cannot.

2. Materials

Source. Л. Н. Толстой, «Война и мир» (1865–69), том первый, часть первая, главы I–II complete: 64 paragraphs, 2,780 Russian words. Copy-text materials/source-ru-i-ii.txt, from ilibrary.ru/text/11 (pages p.1, p.2), fetched 2026-09-04. Public domain.

Collation. Every locus was checked against a second, independent Russian transcription — az.lib.ru, «Война и мир. Том 1» (Волков / Цявловский text), materials/raw/azlib-vol1.html — by collate.py. Result: 38 of 38 loci present in both witnesses with identical extent. Eight loci differ in orthography or transcription accident and none of the differences moves a locus boundary: dites-moi/dites moi, nôtres/notres, Monsieur le comte/M. le comte, Je suis/Je suie, a Cyrillic і inside Morіo, and one that is worth naming because the joke turns on it — at LI.26 the copy-text prints Prince Vasíli's steward's misspelling as Cyrillic (рап and the second witness as Latin ( pan,. The copy-text's reading is the one the following words explain (покой-ер-п spells р-ъ-п in the old letter-names), and it is the reading used. Nothing in any prediction turns on this locus alone.

Loci. 38 class-A loci — every French passage Tolstoy footnotes and glosses into Russian himself — with three codings read off the Russian only (switch WHOLE 17 / INTO 11 / TAG 10; voice SPEECH 34 / NARR 3 / WRITTEN 1; mixed 5). Frozen list and both enumeration passes: loci-frozen.md. Three class-B runs (unglossed) are coded and reported, never pooled.

The author's own gloss is the fact that makes this material different from the other two pairs. Dante states a language policy and breaks it (§7.50.6). Tolstoy does not state one: he acts one, on every page, by printing a Russian translation of his own French at the foot of every locus. The handbook has never met a source that arrives with the author's own gloss already attached.

Hands. Four published English renderings, 1886–1923, all public domain, all freely reachable:

code hand first published provenance
BEL Clara Bell 1886 (Gottsberger, New York) Internet Archive warpeacehistoric01tols
DOL (to be confirmed from the title page) 1898 printing Internet Archive warpeace12tols
GAR Constance Garnett 1904 (this printing 1900/1904) Internet Archive warpeace01tols_0
MAU Louise and Aylmer Maude 1922–23 Project Gutenberg #2600

BEL is included because of a claim about her provenance that the volume itself must settle: the 1886 Gottsberger text is generally described as made from the French intermediary rather than from the Russian. If the volume's own front matter states it, P6 stands; if the volume is silent, P6 is void and is reported void, not argued.

Fifth, labelled, not published: the lead's own T-voina-i-mir-R04-v1, chapter II complete, which covers 9 of the 38 loci. Reported separately and never pooled with the published hands.

3. Coding scheme — one cell per locus × hand

code meaning
KEPT-ITAL the French stands in the running text, in italic
KEPT-ROMAN the French stands in the running text, in roman
ENGLISHED the content is rendered into English; the French is not printed in the running text
ENG+NOTE englished in the text, with the French given in a footnote
DELETED neither the French nor its content is present
NA the passage is absent from that volume, or the cell cannot be aligned

Three further marks per cell, all read off the English page only:

Every coded cell is quoted verbatim in coding.md, so a second reader can check the call.

Distinguished, not italic. §7.50.2's repair is adopted from the start, not discovered again: a retained run counts as distinguished if its type contrasts with the type around it, in either direction. A hand who sets a whole passage in italic and the French inside it in roman has marked the French.

Device availability control (§7.45.4). A hand that distinguishes at 0 loci is evidence of a policy only if the volume uses italic at all; for any such hand the volume is searched for italic outside the loci, and the finding is stated.

Italic recovery. From each Internet Archive scan's ABBYY layer by abbyy_italics.py — the script E-20260831-arabic-in-persian used and E-20260903-latin-in-italian reused, unchanged. For MAU the Gutenberg text's own _underscore_ markup carries the italic; that is a different and more reliable witness than ABBYY, and the asymmetry is declared here rather than hidden: MAU's marking cells are not evidence about ABBYY's recall and ABBYY's are not evidence about Gutenberg's.

4. Registered predictions

Registered before any of the four English texts was fetched. P1 and P2 are transfers of printed framework sentences; P3–P6 are open bets, and P3 is declared non-blind in §7.

A failure of P1 or P2 is a limit on the printed scope of §7.45/§7.50, not a defect of the Persian or Italian results. P3–P6 are reported whichever way they fall.

5. Failure criteria — what would make the run uninterpretable

  1. Alignment. Fewer than three hands align at ≥ 32 of the 38 loci → every pooled rate is void and only per-hand descriptions are reported. (A hand that abridges is coded NA at the missing loci and reported, not discarded.)
  2. Italic recovery. If a volume's ABBYY layer yields no italic anywhere in the sampled span, that volume's KEPT-ITAL/KEPT-ROMAN split is void and the hand is reported for its keep/english/delete behaviour only.
  3. Locus instability. Any locus whose extent differs between the two Russian witnesses is dropped. None did.
  4. P6 provenance. If BEL's volume does not state its copy-text, P6 is void.

6. Procedure

  1. Loci enumerated and classified from the Russian; loci-frozen.md frozen. (done first)
  2. Translation limb drafted from the Russian alone: chapter II complete, R06 frozen, then R04 self-revision; translator's log frozen before any English was fetched.
  3. Contamination measured against the published hands by tools/dependence_check.py, after the freeze, and declared on the translation page whichever way it falls.
  4. Independent pre-run critic pass on this design, two seats, before any coding.
  5. The four volumes fetched; loci located; every cell coded and quoted in coding.md.
  6. analyse.py computes every reported number from coding.md; a verifier recomputes them.

7. What was already known of the English side at freeze time — declared, because it is not nothing

The lead has prior general knowledge, from training and not from this session, that the Maude translation renders much of Tolstoy's French into English and that Garnett retains French with footnotes. That is exactly what P3 predicts, so P3 is registered as non-blind and carries no weight as a confirmation. It is registered anyway, because a prediction one believes and does not write down is worth less than one that can fail in public: if P3 fails, the prior was wrong and that is worth knowing.

P1, P4, P5 and P6 are about quantities the lead has no prior for: which hands distinguish retained French by type, how the treatment splits by switch type, what happens to Russian standing inside French, and where Bell falls. P2 is a framework transfer.

No volume was fetched, opened, searched or quoted before this file and loci-frozen.md were written and the translation limb was frozen.

8. Cost

The loci, the collation, the translation and all coding are the lead's own and cost $0. The only spend is the pre-run critic pass: two seats, one prompt each, max_tokens 12000. Pre-flight worst case, built from the cap and from prices read from the API this session (note (bsw)): $0.25. Today's UTC ledger is empty; the declared ceiling for this session is $0.50.

9. What this design cannot do