Repository path: journal/2026-09-04.md · rendered 2026-09-09
2026-09-04 — what translators do when a novelist's characters simply speak another language
This is a long-running study of literary translation. I translate public-domain literature myself under conditions I write down in advance, have outside AI models judge or measure the results blind, and try to distil what survives into a practical handbook.
Where this had got to
Over the last week the handbook grew a pair of sections about a problem every translator meets: what do you do when your author changes language in the middle of the book? I built them on two cases. Sa'di's Persian Gulistan is dotted with Arabic; Dante's Italian Vita Nova is dotted with Latin. In both, I enumerated every foreign passage from the original, then looked at what four published English translators had done at each one — kept it, printed an English version instead, or cut it.
Last week's session, on Dante, closed by naming what it could not reach. Both those books quote another language: someone is being cited, the words belong to somebody else, and the author announces them. Nothing I had measured told me what happens when an author isn't quoting at all.
What I did today
Tolstoy is the obvious case, so I took it. War and Peace opens in French — the first sentence of the novel is French, spoken by a Russian hostess to a Russian prince in a Petersburg drawing room. Nobody is being quoted. The characters simply speak French, because that is what their class did, and Tolstoy's Russian readers of the 1860s were split: some read French effortlessly and some did not. That is why he translated every French passage into Russian himself, in a footnote at the bottom of his own page. No other book in this project has arrived with the author's own translation attached.
I translated the first two chapters whole from the Russian — 2,780 words — before opening a single English version, and froze my working notes. Then I listed all 38 places where Tolstoy switches into French, checked every one against a second, independent Russian text to be sure I wasn't relying on one website's typing, and coded what three published English translators did at each: Clara Bell (1886), Constance Garnett (1904), and Louise and Aylmer Maude (1922–23). That is 96 cells of evidence.
Before coding anything I sent the written plan to two outside AI models with instructions to attack it. Both said redesign, between them raising 31 objections, 11 of them blocking. They were right about most of it, and I rebuilt the plan before opening any English volume. Two of their demands cost me directly, which is how you know they were real ones — see below.
What came out
The handbook's rules turn out to be rules about quotation. Dante's four English translators kept the Latin standing in 83% of cases. Tolstoy's three keep the French in 16% and put the rest into English. That is not a difference in degree; it is a different decision. The readerships are the same people — Victorian English readers of Dante are Victorian English readers of Tolstoy. What differs is that a quotation is somebody else's words, and a code-switch is the speaker's own voice. A translator will leave another man's sentence standing in a foreign language. He will not leave his own character's speech standing in one.
One rule did travel, and it now stands on three languages and three centuries: nobody deletes. Zero of 96 cells drop what was said.
The finding I did not expect, and the one a working translator might actually use: putting it into English is the gloss, moved off the foot of the page and into the text. Tolstoy prints a double text — French above, his own Russian below. All three translators collapse the two into one English text. And then, at the handful of places where they do leave French standing, they print nothing beside it: one glossed retention in ninety-six cells. I had predicted 80% and it fails on every hand. But they have not withheld the service; they have relocated it. So the choice at a code-switch isn't mark it, keep it, or cut it. It is: carry the author's double text, or collapse it — and all three collapse it.
That has a consequence I can put in one line for a practitioner. If you keep the foreign words, you can mark them typographically — italics. If you put them into English, there is nothing left to italicise, so the only device you have left is saying so. And that is exactly what these translators do: Maude writes "written in French", "said she in French", "still in French"; Bell writes "and in French". Four such signals against one gloss. Garnett writes nothing at all — read her version of these chapters and you would not know a second language is being spoken anywhere in the room.
Two things that went against me, reported as they fell
I had predicted that at the five places where Tolstoy plants Russian words inside a French sentence — Anna Pávlovna translating her own French mid-breath, Prince Vasíli reaching out of French for «мой верный раб», "my faithful slave", because French has no word for what he means — every translator would lose the inner switch. I could not test it: no translator keeps the French at any of those places, so there was nothing left to measure. The one moment in 96 cells where anyone registers the inner switch is Maude's, and he does it inside English, with quotation marks — "no longer my 'faithful slave,' as you call yourself" — and then does not repeat the trick eighteen loci later when the same man uses the same phrase again.
And a prediction of mine is left unresolved rather than passed, because of a check the critics forced on me. Their strongest objection was that I was the only coder, and I was also the person who made the predictions and translated one of the chapters. So I paid $0.12 to have a model re-code all 96 passages blind — translator's name hidden, no predictions shown. It agreed with me on 92 of 96. One of the four disagreements is my error: I had written the coding rule myself, handed it to the model verbatim, and then broken it, counting Maude's "Monsieur Pierre" as English while counting Garnett's "M. Pierre" as French. That single cell flips one of my predictions. I have left the cell as I coded it and reported the prediction unresolved rather than quietly fixing it into a pass. It is the cheapest and most useful $0.12 this project has spent.
A third measurement simply could not be made. I wanted to check whether these translators italicise the French they keep. The Project Gutenberg file of the Maude carries no italics at all, and the scanned page-layer of the Garnett contains no italic anywhere in 1,164 pages. So I can say nothing about it — not that they don't mark, but that I couldn't see whether they do.
The translation
Here is Tolstoy's opening as he printed it, with the Russian words he plants inside his own French in bold —
— Eh bien, mon prince. Gênes et Lucques ne sont plus que des apanages, des поместья, de la famille Buonaparte. […] vous n'êtes plus mon ami, vous n'êtes plus мой верный раб, comme vous dites.
— and mine, which keeps the French and lets the Russian become English inside it, so that the relation survives even though the languages have shifted under it:
Eh bien, mon prince. Gênes et Lucques ne sont plus que des apanages, des estates, de la famille Buonaparte. […] je ne vous connais plus, vous n'êtes plus mon ami, vous n'êtes plus my faithful slave, comme vous dites.
The point of doing it that way is that the switch is the scene. She is speaking the language of her drawing room, and twice she drops out of it — once to translate herself for a Russian ear, once because the French word does not exist. All three published translators put the whole thing into English, which reads faster and loses both.
There is one place my own rule broke down, and I have written it up rather than hidden it. Prince Vasíli calls himself Anna Pávlovna's most faithful раб, then misspells it as his own village steward would — рап — and spells the misspelling out in the old Russian letter-names. The devoicing that makes the joke work has no English equivalent; my "slafe" is an invented spelling doing a job that Russian orthography does for free, and I left the letter-names in Russian, which asks a reader to accept a joke it cannot get. Maude, interestingly, solves it better than I did: "your most devoted slave-slafe with an f, as a village elder of mine writes in his reports."
Cost, and what's next
$0.21 of the $5 daily allowance — two critic passes and the blind re-coding. Everything else — the translation, the source collation, all 96 of my own codings — cost nothing.
The handbook has a new section saying that its change-of-language advice is advice about quotation, and that a source whose characters simply speak the other language is a different problem with a different option set. The two older sections now point at it.
One thing is worth flagging as genuinely open, and it is a case I cannot reach without going outside English: nobody here has translated War and Peace into French, where the embedded language and the target language are the same and the switch cannot survive at all. That is the one situation where the handbook has no move to recommend. The 1879 French translation Bell worked from is public domain, so a future session can take it.
Nothing needs your attention.
Later the same day — a fresh look at the whole project, at Tom's request
This is a long-running study of literary translation. I translate public-domain literature myself under conditions written down in advance, have outside AI models judge or measure the results blind, and try to distil what survives into a practical handbook. This entry is not about a translation. Tom asked for a stock-taking: the model that runs these sessions is changing from Claude Opus 5 to Claude Sonnet 5, and he wanted the whole project looked at fresh before the scheduled sessions restart — and any questions about the goals put to him first.
What the record shows after six weeks and 243 sessions
The good news is real. There are 320 filed translations of 174 works in more than twenty source languages, each with a frozen log of the decisions made while translating; five long works have been translated whole; the habit of writing a plan before running anything, sending it to outside critics, keeping every raw output and recomputing every number has held, and it has caught my own mistakes again and again. The daily reports are honest. Money was never the limit.
Four things were not converging on what the charter actually asks for — a practical handbook and a way of judging whether a translation is good.
The handbook could not be used by anyone. It had become one 380-kilobyte file of fifty-two numbered sections, added in the order the sessions happened to run, many of them later withdrawn or narrowed in place. A translator could not open it and find "what do I do when my author's characters speak French." Its own last line reads "Nothing about quality, still."
Quality judgment had been switched off since August 2. Before the AI jury's scores are allowed to count, the jury has to pass an entrance exam: can it detect deliberately damaged translations, on the right criterion? On August 2 it passed every part of that exam except one number at the heaviest damage level — and passed everything at the lighter level. The verdict said a redesign needed one decision, which damage level to treat as the real test, and a rule written on August 1 said that decision was Tom's. Nobody ever asked him. Every daily report since has ended "nothing needs your attention." So for a month every score has been provisional, and the handbook has been allowed to say what translators do but never what is better.
The subject had drifted from prose to sound. Since about August 20 nearly every session has been about rhyme, metre and sound-play in Persian and Arabic verse and rhymed prose — twenty-eight of the fifty-two handbook sections. The charter says prose first. The safeguard built on August 1 against rabbit holes could not see this one, because each little study finished on time and each named the next as its successor: thirty-five in a row, all "resolved."
The hand-off pages had regrown to two megabytes. The pages a session has to read before it can start had reached roughly 600 kilobytes. The diet imposed on August 1 had no enforcement, and a promise to keep a page short has never once held in this project.
What I asked, and what Tom decided
I put four questions to him. He chose: redesign the jury's exam once (if it fails again, that route is closed); finish the current Persian and change-of-language thread first, then return to narrative prose; rebuild the handbook as entries organised by translation problem; and write every entry for a human translator and for an automated pipeline equally.
What changed today
- The handbook has a new shape.
framework/v0.3is an index of fourteen problems — a source that changes language, forms of address the target lacks, realia, register, what a translator tells the reader in notes and prefaces, sound-play, rhymed prose, long works, evaluation, and so on — with a fixed template: the problem; what published translators did, with counts; what my own practice found; the options; guidance in two halves (for a translator, for a pipeline, with the points where a person should step in named); what is not evidenced; and the sources, including what was withdrawn along the way. One entry is finished as the model for the rest — the one on a source that changes language, consolidating the Sa'di, Dante and Tolstoy studies of the last week. The old fifty-two-section file is frozen as the record. - Sessions are now assigned, not chosen. Instead of six rotating tracks that each needed a fresh little study every four sessions, there are three lines of work with finish lines and ordered steps: the handbook; the jury's exam and then the judging of translations; and the practice (four more sessions to finish the current thread, then a long prose work chosen for being the kind of book rarely translated professionally). Each session's hand-off names the next session's assignment outright.
- The pages a session must read are capped by a tool, not a promise, and the old ledgers were moved intact into an archive. A session now starts with about forty kilobytes of reading.
- The session instructions were rewritten for the new model around five kinds of session, with a rule to stop taking on new work at sixty percent of the available context so that verification, this journal, and the merge always complete.
Cost today: nothing beyond this morning's twenty-one cents. Nothing was deleted from the record.
What it means
For the next month the sessions should produce something a translator could read: a handbook entry every third session, each one applied to a fresh passage as it is written; a re-taken entrance exam for the jury within about a week, which — if it passes on even one criterion — finally lets the project say one rendering is better than another with evidence behind it; and, once the current thread is finished, a long piece of narrative prose translated serially with the handbook applied to it as it grows.
What needs you
The scheduled Routine is still switched off. Switch it back on when you are ready; the next
session will pick up the Gulistan span that is overdue. I could not edit the Routine's prompt from
inside this session — it was created from the web interface, which only you can change. The
existing prompt still works, because it tells the session to follow the entry file in the
repository and the old tool name it mentions now runs the new checker. A shorter replacement
prompt, written for the new entry point, is at the end of the assessment page
(wiki/reassessment-2026-09-04.md) if you would like to paste it in.
Later still — the overdue Gulistan chunk, translated
This is a long-running study of literary translation. I translate public-domain literature myself under conditions written down in advance, have outside AI models judge or measure the results blind, and try to distil what survives into a practical handbook. The Routine came back on today and picked up exactly the piece the assessment above flagged as overdue.
What I did
I am working straight through Sa'di's thirteenth-century Persian collection The Rose Garden, one chapter — "On the morals of dervishes," forty-eight short tales — translating it whole, in order, under a fixed set of terminology and style decisions I have been keeping and adding to as I go. Today's piece was the next-to-last stretch: eight short tales, some forty lines, from a bickering argument between a battle standard and a comfortable litter-curtain to the chapter's closing image, an epitaph on a legendary king's tomb. I translated it directly from the Persian, wrote up the decisions before opening any of the four published English versions I keep on the shelf for comparison, and only then checked my own wording against them and against my own earlier work to make sure I had not unconsciously echoed either.
What I found
Five open questions had piled up from the terminology decisions of earlier sessions, and this stretch answered all five. Three are worth passing on:
A repeated word can carry across, even when it is just one word. Persian poetry often ends back-to-back lines on the same word, and I had already found that a repeated phrase survives translation into English cleanly. Today gave the first case of a repeated single word — the imperative "take" — and English carried it too, at some cost to normal word order ("Sa'di — the road to the Kaaba of good pleasure, take; O man of God — the door of God, take"). Small, but it settles a question about exactly what kind of repetition an English sentence can bear.
Two rules I had thought were in tension turned out not to be. Sa'di quotes a Qur'anic verse in this stretch, explicitly introduced ("this agrees with the Qur'an —"), and then follows it with a verse of his own that makes the same point again in Persian. I had one rule saying an explicitly introduced quotation gets marked as foreign in the English, and a different rule (from an earlier Sa'di passage) saying that when he restates a foreign quotation in his own words afterward, the foreign text doesn't need marking. This passage has both features at once, and the answer turns out to be simple: those are two separate questions — does he name his source, and does he then say it again himself — and either can be true without the other. Naming the source is what decides the marking, regardless of what follows.
The best find crossed a session boundary I did not expect it to cross. I keep a private log of words and images Sa'di repeats across different tales in this chapter, because catching them is something only translating the whole book in order lets you do. One recurring pairing — a truly wise person, and a stone that does not disturb him — turned up in today's stretch (tale 46: "if a millstone comes tumbling down from the mountain, he is no true knower who is provoked by the stone"). It is the same pairing, in the same words, that appeared in the tale I translated six days and one working session ago (tale 40, close of last week: "a great sea is not fouled by a single stone; the knower who takes offense is shallow water still"). The two tales are forty-seven lines apart in the book and I translated them a week apart, with no memory of the earlier line in view when I hit the later one — so the echo survived in the English by accident of consistent terminology, not because I was watching for it. That is itself a small lesson for how I keep this kind of running log: it needs to be consulted, not remembered.
I also found one deliberate piece of wordplay that came across for free: Sa'di contrasts a courtier's fitted coat with a dervish's rough cloak using two Persian words that sound almost alike but are unrelated in origin — his point is that the words for the fancy and the humble garment chime against each other while meaning opposite things. English "coat" and "cloak" happen to do exactly the same thing, unrelated words, one sound apart, so the joke crossed over intact.
The translation
Here is the chapter's closing tale, on generosity, in full:
A sage was asked: which is better, generosity or courage? He said: whoever has generosity has no need of courage.
Hatim of Tayy did not remain, but until eternity his exalted name remains, famed for goodness. Pay the tithe on your wealth — for when a gardener prunes back the excess of the vine, it yields more grapes. It is written on the grave of Bahram Gur: the hand of generosity is better than the arm of force.
Hatim of Tayy and Bahram Gur are both pre-Islamic Arabian and Persian figures whose names, by Sa'di's time, meant roughly what "a Good Samaritan" or "a Midas" mean now — a name that carries a whole story in it. The chapter began, forty-eight tales earlier, with dervishes and kings compared; it ends, here, on the same comparison, in a king's own epitaph.
Cost, and what's next
Nothing: the translating itself is free, and nothing was measured or judged today that would call for paying an outside AI model. One short stretch of the same chapter remains — the first ten tales re-translated under the terminology this whole run has now settled, since they were originally done years earlier under a different, narrower approach — plus a closing report on what translating the whole chapter straight through, rather than piecemeal, actually taught. That is the next session's assignment, and it finishes this book.
Nothing needs your attention.
Later still — the handbook's second finished section: who defers to whom, when English can't say it grammatically
This is the same long-running study of literary translation — I translate public-domain fiction myself under conditions written down in advance, have outside AI models judge or measure the results blind, and distil what works into a practical handbook, organized by translation problem.
Where this had got to
The handbook got its first finished section yesterday. Today's session wrote its second: what a translator should do when a source language grammatically marks who outranks whom — a formal versus familiar "you" (French and Russian both have this; English lost it when "thou" died out), a verb ending that changes depending on how much respect the speaker owes the listener (Japanese does this constantly), a suffix glued onto a word purely to signal deference. English has almost none of this machinery. The question the section had to answer, from six weeks of this project's own small experiments across Russian, Japanese, German and Bengali translations: when English can't mark the relationship the same way, is the relationship actually lost on the page, or does it usually come through some other way? And if a translator tries to add it back deliberately, what does that cost?
What I did
I read back through thirteen separate studies this project had already run on this question — scattered across two earlier stages of the handbook that needed reconciling — sorted out which findings held up and which had been corrected or withdrawn along the way, and wrote the whole thing into one entry with practical guidance for a human translator and, separately, for an automated pipeline. Then, to test whether the guidance actually holds up in practice and not just on paper, I translated a short passage fresh — a comic Chekhov sketch nobody in this project had touched before — applying the entry's own advice, and kept a log of which instructions actually decided something as I worked.
What I found
The single biggest surprise, confirmed independently four separate times across four language pairs: where a language marks rank through a single required grammatical slot — above all a pronoun — English essentially never has a same-kind replacement, and translators don't invent one elsewhere either. Russian's formal/familiar "you" survived into a set of English translations at zero out of eleven tested spots. A cluster of Japanese honorific pronouns survived at roughly one site in fifty. This part matched what I would have guessed.
What did not match the guess: that loss usually doesn't cost the reader anything. In one carefully checked case — a scene from Kleist's 1810 German story "Michael Kohlhaas," where a gentleman calls a knight the old-fashioned formal "you" and the knight calls him back the familiar, condescending form — a 1914 published English translation that dropped every trace of that distinction was still judged, by readers shown only the English, to convey which man deferred to which at eight sites out of nine. The same pattern held in a Chekhov story and a Japanese one: what looked like an urgent translation problem mostly wasn't, because the words themselves — who says what to whom, in what tone — were already doing the work.
When a translator does try to put the marking back everywhere, in one measured experiment (the same short scene from Genji, rendered three different ways), it cost about a fifth more words, and, worse, it invented a status the original text never actually stated: readers reading the "maximum effort" version consistently decided that whichever character the sentence happened to be about was the more important one — not because the story said so, but because that was where the translator had piled on the honorific language. A rule I tried that marked "whoever the paragraph's own grammar already treats as senior" produced a smoother-sounding result that was wrong at exactly the three places where the original text itself reverses who's on top.
And a small experiment inside this project found something that undercuts the obvious fallback advice — "just watch the places where the English text goes quiet on this." Two careful readers, given the same seventy lines of Dickens dialogue and asked to say where the English told them nothing about who deferred to whom, agreed with each other on that question at worse than a coin flip. The set of "silent" places isn't even stable between two attentive readers of the identical text. So the workable advice turns out to be the opposite, and duller: assume you are supplying this relationship everywhere, and decide once, deliberately, how the two characters stand to each other for the whole scene — not site by site.
The translation
Chekhov's 1884 sketch "Surgery" turned out to be an unusually good test case: a deacon comes to a rural clinic with a toothache and grovels to the medical assistant treating him — formal address, elaborate blessings, self-abasement — while the assistant, technically not even a qualified doctor, addresses him bluntly. When the tooth-pulling goes wrong, both men's politeness collapses at the same instant into open insult. English has no grammatical way to mark the formal-versus-familiar distinction the Russian uses throughout, so I carried the whole effect through word choice and tone instead, timed to break at the same line the Russian does:
"Our benefactors, that's what you are... We fools would never think of it, but God has given you the light..."
[the tooth won't come out]
"You pulled it, all right!" he says, in a voice both weeping and mocking. "May you get a pull like that in the next world! Much obliged to you, I'm sure! If you don't know how to pull a tooth, don't take it on!"
"And what do you go grabbing with your hands for?" the feldsher snaps. "Fool!"
"Fool yourself!"
I could not find a freely available published English translation of this particular story to check my own rendering against for suspicious overlap with anyone else's wording (a standing precaution against unconsciously reproducing a translation absorbed during training) — I checked the obvious collections and came up empty — so the honest label is that contamination is merely "suspected," on Chekhov's general fame, not measured against anything concrete.
What it means
For a translator: don't hunt for a workaround wherever a source marks rank on a single mandatory grammatical slot like a pronoun — none exists, and inventing an old-fashioned English pronoun ("thou") to fake one imports a historical flavor the original doesn't have. Trust the surrounding words to carry the relationship first; only spend real effort compensating at the rare spot where the marking is the only thing establishing who defers to whom. If you do add something back, prefer one device, not several stacked together, and know in advance that going all-in is mostly just longer prose, not more faithful prose. All of this stays "in my own, unverified judgment" where a human reader's actual response is concerned — the model jury that's meant to check that kind of claim still hasn't passed its qualifying test.
Cost, and what's next
Nothing: the translating itself is free, as always, and nothing was sent to an outside AI model today. Next up is a different kind of session: redesigning that jury's qualifying test itself, which has now failed once and gets one more attempt, by Tom's earlier instruction.
Nothing needs your attention.