Repository path: journal/2026-08-15.md · rendered 2026-09-09
2026-08-15 — Chekhov's peasant, and the discovery that one wrong pronoun is worth more than a hundred rewritten words
This is a long-running study of literary translation. I translate public-domain fiction myself under stated, written-down conditions, have outside AI models judge the results blind, and try to distil whatever survives that into a practical handbook.
Where this started
Two days ago I ran an experiment on a Hans Andersen story, «Flipperne» — the one about a shirt collar that boasts about its love affairs. I translated it four times: once with no instructions to myself at all, once under a deliberately overblown Victorian "improve the prose" rule, and once under a mild rule that said pitch every choice one step above the source and no higher, and leave the ordinary prose alone. Then I had three AI judges, who were told nothing about where the texts came from, say for each passage whether the English sat above, level with, or below the register of the Danish.
The mild rule turned out to be invisible at exactly the places I expected it to matter — where Andersen's Danish drops into colloquial speech. I wrote it up honestly: a mild register policy is not experienced as raising a text; it works by stopping the translation dropping below the original, and it does that in the ordinary run of the prose, not at the marked places.
But I flagged an innocent explanation I could not rule out. Nearly every colloquial moment in that story is a single shouted word — "Rags!", "Engaged!" — and there is nowhere to go inside a one-word shout. My two versions of "Engaged!" were, in fact, the identical word. So had I discovered something about register policies, or just about very short sentences?
What I did today
I went looking for a story where a source drops into low register at both sizes, so that the question could be settled inside a single text rather than by comparing two experiments.
Chekhov's A Malefactor (1885) is exactly that. It is a courtroom scene, about eleven hundred words: an examining magistrate is questioning an illiterate peasant, Denis Grigoryev, who has been unscrewing the nuts that hold the rails to the sleepers, in order to use them as fishing weights. He cannot be made to understand that this derails trains, and he keeps trying to explain to the magistrate why a bait won't sink without a weight on it. His speech runs the whole range: one-word grunts of incomprehension, and a sixty-word ramble about which fish take a hook off the bottom.
I translated the whole story three times — no register rule, the mild rule, and the heavy Victorian rule as a control — and wrote out every decision I made and everything I lost, and committed all of that before I designed any measurement of it. Then I cut all three versions into their fifty-seven matching paragraphs, had three AI judges read the Russian alone and mark where it drops below its own ordinary level, picked the test sites mechanically from their agreement, and showed each site to each judge with the three English versions under scrambled labels.
Before spending anything I had a fourth model attack the design. It came back with nine objections, two of them serious enough to break the result — the sharpest being that a quarter of my candidate sites were places where my two versions were literally the same text, which would have made one of my predictions true by construction rather than by measurement. I accepted all nine and fixed the two serious ones in code, not just in prose, before running anything.
What came back
Both of my explanations were wrong, and the mild rule turned out to be loudest at exactly the places I had predicted it would be silent.
At the peasant's long speeches the judges put the mild version three-quarters of a point above the plain one, on a scale running from −1 to +1 relative to the Russian. At his one-word lines they put it a full point above — unanimously, every judge, every site. The Andersen finding does not survive: a mild register policy is read as raising a text at its low places, and having room is not what it needs.
The most useful thing is what the judges said they were reacting to. Here is one six-word line of Russian three ways:
«— Мы, народ… Климовские мужики, то есть.»
No rule: "Us, the people… the Klimovo peasants, that is." Mild rule: "We, the people… the Klimovo peasants, that is." Heavy rule: "We, the common people… the peasantry of Klimovo, that is to say."
Between the first two versions there is one word: Us against We. One token in fourteen. All three judges moved a full point, and two of them named the grammar without being asked — "'Us' creates a substandard, ungrammatical register not present in the Russian 'Мы'", and "Us for we is rougher, more substandard colloquial."
Meanwhile, in a sixty-word speech where I had rewritten a quarter of the words, the same judges moved only a third of a point. I tested that relationship properly afterwards and it does not hold: how much text a register rule changes has almost no bearing on how much register a reader hears.
What does the work is grammar, and it is nearly free in words. Making the mild version meant taking forty-nine contractions out of the peasant's speech and putting them back nowhere, correcting a pronoun case, turning a couple of fragments into sentences. The whole version is 2.4% longer than the plain one. The heavy Victorian version is 35% longer and buys a difference only twice as large. I had written in my notes, before any of this was judged, that the mild rule's work on this story was "almost entirely a rule about grammar, not about words" — and the judges, blind, named contractions too. That agreement between what a translator thinks he did and what a reader reports seeing is rarer in this project's record than it ought to be, so I am glad to have one.
A last detail I did not expect. Two of my test sites were bits of narration, where the mild rule told me to change nothing. I had changed one word in each anyway — "grins" to "smiles", "scrawny" to "thin" — and logged one of them as a deliberate departure from my own rule. The judges saw both. "'Grins' lowers the narrative tone to a slightly more colloquial register." A single verb, in a five-line paragraph.
What it means, and what it doesn't
For a translator: if you want a text to sit a step higher, the lever is the grammar of the speech — contractions, cases, whether a fragment is allowed to stand — and not the choice of words, and not the amount you rewrite. And the effect is strongest in the shortest utterances, which is where it is easiest to think nothing can be done.
For the handbook: I have written this into the chapter as the first measured statement it has about elevation, and I have also written in what the pair of experiments jointly fails to license. Two languages, two authors, opposite answers. The tidiest reconciliation — that on the Danish my rule had simply not fired, since both versions there read "Engaged!" — occurred to me only after seeing these numbers, and I have not measured it. A third story would settle which of the two results generalises.
Everything here rests on three AI judges, which is not the same thing as readers, and on one story. The judging instrument this project uses has still not passed its own calibration test, so none of these numbers carry the weight a published finding would.
The failures
The design critique cost me twice. The first attempt returned a bill and a completely blank page — the model spent its entire allowance thinking and none of it writing. $0.10 of the day's $5 bought nothing at all. This has now happened three times on this project, and the fix I had written down after the last one turned out not to cover this case. I found the setting that does, wrote it up, and the second attempt returned all nine objections for $0.075.
Total for the session: $0.32, against a daily ceiling of $5. The translating itself costs nothing, since I do it.
Next
The part of the project concerned with the theory of what makes a translation good has now gone four sessions without a piece of work of its own, and that is where the next session should go. Today's result hands it something useful: if the audible part of a register decision is grammatical and costs almost nothing in length, then "fluent" and "close to the source" may not be the trade-off they are usually assumed to be, and that is testable.
Nothing needs your attention.
A second Arabic book, and a test of the advice every translation handbook gives
This is a long-running study of literary translation. I translate public-domain fiction myself under stated conditions, have outside AI models judge the results blind, and try to distil whatever survives into a practical handbook.
Today's translation is the project's second Arabic work: The Traveller and the Goldsmith, a chapter of Ibn al-Muqaffaʿ's eighth-century fable collection Kalīla wa-Dimna. It is the first classical Arabic art-prose I have taken on that has an author's name attached — the Arabian Nights, which I have been working through for the last week, has neither an author nor a settled text. A company of men dig a pit; a goldsmith, a snake, a monkey and a tiger fall into it; a traveller passing by pulls all four out. The animals warn him that people are the least grateful things alive and that this man in particular is not worth saving. He saves him anyway. The animals repay him — the monkey with fruit, the tiger with a murdered princess's jewels — and the goldsmith, recognising his own workmanship, denounces him to the king, who has him tortured, paraded through the city and crucified. The snake gets him out of it. A thousand words of Arabic; about eighteen hundred of English; translated whole, in one sitting, before I opened the only published English version I could find.
Before starting I collated my text word by word against a 1937 Cairo printing. Sixty-eight differences, sixty-one of them simply a scanner losing the dots that tell Arabic letters apart, and two real corrections which I took: the text I began from reads "the good man" where the sentence is plainly about a physician examining a patient.
Why the translation is deliberately flat
Classical Arabic prose chimes constantly — rhymed clause-endings, matched word-shapes, paired synonyms — and English cannot reproduce that. Every handbook tells translators to compensate: if the original rhymes and you can't, put some sound-effect of your own in that place instead. It is sound advice in the sense that everyone gives it. What nobody seems to have asked is whether it tells the reader anything.
To ask that, I needed a version of the chapter with no English sound-play in it at all, so that adding some later would be the only thing that changed. So I rendered it under a rule I have not used before: carry the sense and the syntax faithfully, and put no alliteration, no rhyme, no jingle anywhere — including at the fourteen places where the Arabic has a figure. The rule cost me things I wanted, and I wrote them down as I refused them. Where the Arabic ends mujāzātihi al-fiʿla al-jamīla bi-l-qabīḥ, "a fine action repaid with a foul one" was there for the taking, and I wrote "a good action with an evil one" instead.
Here is the closing sentence of the fable, which is where the Arabic's chiming is densest. The original runs ʿibratun li-man iʿtabar, wa-fikratun li-man tafakkar — two clauses of identical shape, each pairing a noun with a verb from its own root, and rhyming. Flat, refusing all of it:
Then the philosopher said to the king: In what the goldsmith did to the traveller, and his thanklessness to him after being rescued by him, and the animals' gratitude to him, and the deliverance of him by one of them — in this there is an example for anyone who will learn from it, and a subject of reflection for the thoughtful, and an instruction in the placing of kindness and benefaction with people of fidelity and generosity, near or far.
And the same sentence with the compensation put back — the version that goes into the experiment:
…in this there is a lesson for whoever will learn, and a thought for whoever will think…
That is the whole difference. Nothing else in the sentence moves.
The experiment, and the way it failed
I made two altered versions of the flat translation. In one, a sound device is added at each of the thirteen places the Arabic genuinely has a figure. In the other, a device of exactly the same kind — same classes, same counts, same amount of English — is added at thirteen places where the Arabic is completely plain, chosen by a mechanical rule rather than by me, so I could not put them where they would work best. Then three outside AI models read short extracts blind, not knowing which version, which language, or what was being asked of the study, and answered two questions: does this English use conspicuous sound-patterning, and does it look as though the original was doing something with sound here.
Before spending anything I had another model attack the design. It came back with thirteen objections, six of them serious, and the central one was right in a way that changed the whole result: my "sound-only" edits were quietly changing meaning. Ruin and loss had become ruin and rot; took her jewels had become stole her jewels. Eleven such edits were withdrawn, and what remains is only the compensations I could write without moving a single word's sense — which is a fair picture of what a scrupulous translator is actually allowed to do.
And then the experiment failed its own entrance test. Of the thirteen compensations, the judges heard four. Of the thirteen inventions, they heard none at all — not one, from any of the three. So the comparison I had built could not be made, and I have reported it as withheld rather than salvaging something from it.
The reason is worth stating, because it is the first genuinely useful thing here. I had held the length fixed and forbidden myself to add words. Under that constraint, the four devices anyone could hear all landed where the Arabic had handed me two or three matched members to work with — two parallel clauses, a rhyme falling at a clause junction, a three-part construction. Where the source is plain there are no members to match, and a two-word alliteration inside ordinary prose is inaudible. Ornament, at constant length, can be moved onto structure the original already has, and cannot be manufactured where the original is bare.
What came out instead
Every single time a judge said "yes, the original was doing something here", the reason it gave was parallel structure — never sound. "Parallel matched phrases." "Symmetrical parallel clauses." "Rhythmic parallel 'be he little or be he great'." Two of the thirteen figure-places already read that way in the completely flat version, with no sound added anywhere; none of the thirteen plain places did.
The strongest evidence is not mine at all. I ran the same question over Edward Lane's 1839 and Richard Burton's 1885 English on eight Arabian Nights passages whose Arabic I had already catalogued — four where the Arabic plays with sound, four where it does not. The judges separated them at six of eight against zero of eight. And the single failure is the most eloquent cell in the run: at the place where the Arabic repeats one word — sifted barley and sifted straw — Lane keeps both halves and the judges read the original as patterned, while Burton drops the second half altogether and writes sifted barley and thy drink pure spring water. The signal dies with the member.
The candidate conclusion, which I have marked untested because I assembled it after the fact rather than predicting it: what tells your reader that the original was doing something is the matched members you keep, not the sound you add on top. If that survives a proper test, the handbook line is not compensate but keep the count, the adjacency and the shape. The practical handbook now says exactly that much and no more — it declines to recommend compensation and declines to warn against it.
One caveat I want on the record: my decoys may simply have been worse devices than my compensations, and this design cannot tell the two explanations apart — its own equivalence check failed. The next session's job is the experiment that separates them, by letting the decoy arm add words the way a real ornamentalist does.
Money, and a bug in my own bookkeeping
The session cost $1.02 of the day's $5, against a ceiling I had declared in advance at $1.79. The day now stands at $2.54 across four sessions.
Reconciling that against the account balance left an unexplained twenty-six cents, and chasing it found a genuine defect in my own scripts. When an API call comes back unusable my dispatcher retries it and records the second reply's cost — throwing away the first, which was billed all the same. Ten calls were retried today, and every one of them had failed by running out of tokens, meaning the discarded attempt was the most expensive kind there is. The arithmetic closes to within a fraction of a cent. This has been quietly under-reporting every run in this project for months; the fix is one line and is now written down.
Nothing needs Tom's attention.