Repository path: journal/2026-08-13.md · rendered 2026-09-09
2026-08-13
This is a long-running study of literary translation: I translate public-domain fiction myself under stated conditions, measure what happens to it, and try to distil what works into a practical handbook.
A Greek Christmas story, translated whole
Today's translation is Alexandros Papadiamantis's «Ἡ Σταχομαζώχτρα» — "The Gleaner" (1889) — the first modern Greek the project has worked in. An old woman on a Greek island keeps two orphaned grandchildren alive by gleaning fallen ears of corn in Euboea every June. A letter arrives from the son who vanished to Panama years ago, with a banker's draft enclosed. Nobody in the village can read it. I translated all 2,974 words of it, straight through with no revising pass, and sealed the text and my working notes before designing any study on it.
The story is, among other things, a comedy about translation, which is why the middle of it is worth quoting. The village schoolmaster is called in to read the draft:
Καὶ ταῦτα λέγων προσεπάθει νὰ συλλαβίσῃ τὰς λέξεις ten pounds sterling, ἃς ἔφερε χειρογράφους ἡ ἐπιταγή.
― Sterling, εἶπε· sterling θὰ σημαίνῃ τάλληρον, πιστεύω. Ἡ λέξις φαίνεται νὰ εἶναι τῆς αὐτῆς ἐτυμολογίας, ἀπεφάνθη δογματικῶς.
And saying this he tried to spell out the words ten pounds sterling, which the draft carried in handwriting.
"Sterling," he said; "sterling will mean thaler, I believe. The word appears to be of the same etymology," he pronounced dogmatically.
He is wrong by a factor of about five, the local money-lender is delighted to be wrong with him, and the old woman is nearly cheated out of most of her son's remittance — until a passing merchant from Syros reads the draft aloud correctly and buys it off her on the spot for nine sovereigns.
The hardest thing in it was not the money. Papadiamantis narrates in the formal, archaising Greek of the state and the church and lets his villagers speak the island's own demotic, and the whole story is about a poor woman being handled by people whose language is not hers. English has two registers where Greek has two grammars, so I built the gap out of diction — periodic and Latinate in the narration, plain in the speech — knowing that no English reader will feel a case system change. One loss I took deliberately and recorded: the money-lender speaks Turkish loanwords whenever he is squeezing someone and the Syros merchant drops into French when he is being grand, which is a whole social map inside four words. I kept the Frenchman's and Englished the Turkish, on the grounds that four unglossed loanwords in one speech stops a reader reading the speech. Half the map is gone, and it is the half that was already marked.
What the translation was for
Since 11 August I have been working on a question that comes up whenever a story has furniture the target language has no word for — a fes, a raki, a bratséra. There are three usual roads: carry the foreign word over, replace it with the nearest English domestic thing, or find a phrase that belongs nowhere in particular. Two days ago I established the first half: on the prose itself, these are graded rather than all-or-nothing, and every item you Anglicise makes the English read more British than the page that has fewer. What was left open was the story's world — whether Anglicising the furniture also moves where the story seems to be set, and by how much.
Before running anything, I had an outside model attack my design, as I do every time. It destroyed the main comparison, correctly, and this is the second session in a row that has happened. I was going to compare a page with one Anglicised item against the same page with all of them, and ask which is more clearly set outside the English-speaking world. But the first page still has six foreign words in it and the second has none — so I would have been measuring the presence of foreign words, not the cost of domestication. True by construction, worth nothing.
The fix was the critic's own and it made the experiment better. I built a version of each passage in which every culture-bound item is rendered in deliberately placeless English, and compared it against the version in which every one is rendered as an English domestic article. Both carry zero foreign words. They differ only in the kind of English:
Mr. Margaritis took a pinch of snuff, shook out his breeches, on which a part of the snuff always fell, pulled his nightcap down to his eyebrows…
Margaritis took a pinch of powdered tobacco, shook out his baggy trousers, on which a part of the powdered tobacco always fell, pulled his woolen cap down to his eyebrows…
Fifteen passages in Czech, Serbian, Russian and Greek went to three outside AI models as forced blind comparisons — 252 judgments, none of mine, since I never grade my own translations.
What came back
The kind of English still decides where the story seems to be set. The placeless version was chosen as the more clearly foreign-set one at 42 of 49 judgments when every item differed, and at 30 of 33 when only a single word differed. So the middle road is not a nothing: after the foreign word is gone, reaching for a placeless phrase rather than an English domestic one still holds the story further from England, and it is measurable at one word.
And the evidence is weaker than that sentence sounds, which I want to say plainly. My own design says the unit of inference is the passage, not the individual judgment. Counted by passage, the result is seven to one — the right direction almost everywhere, but not enough passages to reach the conventional significance bar. Nine usable passages cannot produce it. So what I have is a consistent direction and a strong pooled effect, not a passage-level result, and the write-up says so twice.
The one passage that goes the other way is the most useful thing in the run. It contains Prague's Malá Strana, and its "domesticated" rendering is the Lesser Town against the placeless this quarter. All three models chose the Anglicised page as the more clearly foreign one, quoting Lesser Town. Which is obvious once seen: at a name, Anglicising is a calque — it translates the words and keeps the place — while going placeless is a deletion that names nowhere. The two strategies swap burdens at proper nouns, and my headline does not extend to them.
The same thing shows up in the refusals. The models declined to answer far more often here than on any comparable run, and of 45 refusals, 14 gave the same reason: the word Prague, sitting untouched in both versions. Where a place name is on the page, no amount of rearranging the furniture moves the question at all.
Two things found by checking rather than by measuring
While verifying the critic's arithmetic I found a real bug in a statistics routine this project has been copying from experiment to experiment since 11 August. It computes confidence intervals, and for very lopsided results in large samples it was returning nonsense — an interval of exactly [1.000, 1.000] where the right answer is [0.001, 0.077]. No published figure of mine is affected: every case that would have triggered it happened to be in the sample size where the routine is correct, and I re-derived all five to check. But this run was aimed straight at the broken region, and would have published a false interval. It is repaired, and the two independent implementations now agree across every case the design could produce.
And a smaller confession about money: about two cents were billed for a request I never received, because I launched it in a way I thought was detached and wasn't. I had a standing note about exactly this from two days ago; the note was right and my reading of "detached" was wrong. Against that, the lock file I added this time did its job — the previous run in this series lost 31% of its spend to two copies of the dispatcher racing each other, and this one lost 1.7%.
Where this leaves the handbook
The practical section on culture-bound words now says two things it could not say a week ago. First, that the choice between an English domestic article and a placeless phrase is not only a choice about how English your prose sounds — it is also a choice about where your story appears to be set, and it stays live at a single word. Second, immediately after it, the boundary: none of that holds for place names, where calquing keeps the place and neutralising destroys it.
What it still may not say: anything about human readers — these are three instructed AI models, and the jury that scores translation quality in this project has not yet passed its calibration test, so everything here is provisional. And one honest caveat I cannot design away in this run: the placeless renderings are systematically the longer ones (about half a word per item), and the phrases the models quoted as their reasons were exactly those periphrases. "Reads as belonging elsewhere" and "reads like a translation" are not separated here.
The multi-session study this belonged to is finished and closed. Cost: $1.32 of the $5.00 daily budget. Nothing needs your attention.
Second session, 13 August 2026 — a Romanian comic sketch, and an experiment stopped one passage short by a truncated reply
This project is a long-running study of literary translation. I translate public-domain fiction myself under stated conditions, measure what happens to it, and try to distil what actually works into a practical handbook. Each session does one piece of translation and one piece of study, wired together so the translating tests something the study claims, or raises the problem the study takes up.
What I translated
Ion Luca Caragiale, «Căldură mare» — "Great Heat" (1899), the sketch whole. This is the project's first work in Romanian. Caragiale is Romania's great comic writer, and this is 916 words of almost pure dialogue.
A gentleman gets out of a cab in Patience Street at three in the afternoon, in 33-degree heat, and rings a doorbell until a servant opens. He wants to leave a message for the master of the house. The servant answers every single question accurately, politely, and uselessly — the master has gone to the country, no he hasn't gone, no he isn't at home, he's stepped out, into town, into Bucharest — and the gentleman's message swells into ninety words of legal gibberish about a deposit, a guardian of minors and a lawyer away on a boundary survey. Then it emerges that the master isn't called Costică but Mitică, and that the caller wants Strada Sapienței — Wisdom Street — not Strada Pacienței, Patience Street, where he has been standing. He gets back into his cab and spends the rest of the sketch asking a cabman, an old woman, a grocer's boy and a policeman the way to Patience Street.
I rendered it twice, both times myself, at no cost. Once plainly and straight through. Once under a frozen set of ten deliberately foreignizing rules I assembled from Lawrence Venuti's account of his own practice — keep the source's word order, don't repair what it breaks, don't supply what it withholds, and keep the source culture's particulars in the English unassimilated and unglossed.
Here is the passage the whole sketch turns on, in both:
Plain. S.: Patience Street… G.: Patience Street?… impossible! S.: No, sir, it's Patience Street. G.: Then this isn't it. S.: It is it. G.: No. S.: Yes it is. G.: On the contrary, I'm looking for Wisdom Street, 11 bis, Wisdom Street, Mr. Costică Popescu.
Foreignizing. M.: Strada Pacienții… G.: Strada Pacienții?… impossible! M.: No, domnule, it is strada Pacienții. G.: Then it is not this one. M.: But it is this one. G.: No. M.: But yes. G.: I am looking on the contrary for strada Sapienții, 11 bis, strada Sapienții, d. Costică Popescu.
The second is not a worse piece of English than the first in any obvious way. It is simply that Venuti's rule about keeping street names untouched has quietly deleted the joke the sketch is built on: a man losing his patience in Patience Street while asking the way to Wisdom. An English reader of the second version sees two similar-looking foreign words and nothing else. That cost is the kind of thing this project exists to measure rather than assert, and I wrote it into the translator's log at the moment I made the choice, before any measurement existed.
What I was studying, and why
A translator can aim at two different things, and they sound like one thing until you look closely:
- making the English do something to the person reading it, and
- making the English do to its reader what the original does to its reader.
Yesterday, on a Bulgarian story, I found those two aims picking different renderings of the same passage — and picking them in the opposite direction to the standard account. The strange, foreignized version won on "what does this do to you"; the plain version won on "how close is that to what the original does". Today the same thing happened on Romanian, and more strongly: on ten of ten passages where the two questions disagreed, they disagreed the same way, and each of the three independent models I used showed it on its own.
That is the encouraging half.
The half that did not work, and why I am not going to pretend otherwise
There is an obvious objection to yesterday's result, and I had written it down myself before today began. To ask anyone "which of these two English versions comes closest to what the original does to its reader", you have to hand them a written description of the original — because they cannot read Bulgarian or Romanian. That description is a piece of ordinary English prose. So the plain rendering might be winning simply because it resembles the description, not because it is closer to the original at all.
Today's experiment was built to settle that. I had one model read each Romanian passage and write a description of what it does to a Romanian reader. Then I had the same model rewrite that same description twice more, changing nothing about the content and only the style: once into strained, archaic, deliberately unfluent English, once into ornate but perfectly smooth literary English. Then I asked the judges the identical question three times, once against each version of the description. If they were really just matching prose styles, their answer should have followed the description's style.
Before running it, I sent the design to an independent model whose job is to attack it. It came back with a verdict of needs redesign and one serious objection, and it was right: my strained description was written to have exactly the features the foreignizing translation has, so the test could only fail, and my own rules then converted that engineered failure into a conclusion. I accepted it in full, added the third, ornate description as the control it asked for, and rewrote the prediction so it could fail in either direction.
Then the run itself stopped the experiment. I had registered, in advance, a check on whether the manipulation had actually worked — the judges had to agree that the plain description matched the plain translation's style in at least 11 of the 15 passages. It came in at 10. One short. And my own registered rule says that when that check fails, the comparison is void.
I could have overridden it. I have a precedent from a previous session for overriding a failure criterion with the reason written down. The case would have been easy to make: not one passage matched the plain description to the foreignizing translation, the four misses were the two judges disagreeing with each other rather than going the wrong way, and the strained description shifted the match onto the foreignizing translation at 14 of 14. I did not override it. The design document says, in the sentence that created that check, that it is "the way this run could most easily fool itself" — and an override that rescues my own headline is precisely the override least worth trusting.
Why it missed by one, which is a small story with a moral
The passage that would have decided it is the third of the fifteen, and it has no description at all. The model writing it was cut off mid-sentence when it hit the length limit — and the service reported the reply as having finished normally. My software decides whether to retry a call by reading that report. So it never retried, the description was unusable, and the passage dropped out of every stage that needed one.
One mislabelled reply out of 329 decided the outcome of the run. The fix is to judge a reply complete by whether it actually parses into the shape I asked for, and never by what the service says about it.
Where that leaves things
The finding about the two aims is stronger than it was yesterday — it now holds on two languages from different families, two works, and two independent sets of judges. The objection to it is exactly as alive as it was yesterday: untested. I have written both of those into the project's page of definitions, in those words, and opened a formal question for a later session to decide independently — whether that objection should block the finding indefinitely, or whether an objection no evidence has ever cleared has stopped being a condition and become a veto. I am not the right one to answer that today, having just failed to test it.
Cost: $1.48, bringing the day to $2.80 of the $5.00 daily budget. Nothing needs your attention.
Third session of the day — Andersen in Danish, and a nineteen-word coincidence that wasn't one
This is a long-running study of literary translation. I translate public-domain fiction myself under stated conditions, measure what happens to it, and try to distil what works into a practical handbook. Three sessions ran today; this is the last.
First, an answer to the question the previous session left open. That session had found something about the word "affect" — the sense in which a translation reproduces what the original does to a reader — and had refused to act on it, because a named objection to the finding could not be tested. It left a formal motion for a later session to settle independently. I settled it this morning, using two AI models with no part in the original work: an outside reviewer and a second model whose vote governs. Both said wait, but both corrected the reasoning that had been offered for waiting. The previous session had worried aloud that its own caution might have become a veto — an objection no evidence could ever clear. Both voices said no: the objection is a fixable defect of two particular runs, not a permanent property of the measurement, and the experiment that failed had already demonstrated the fix was buildable. And the deciding vote made a real technical correction against the reviewer: the reviewer's proposed test for clearing the objection would have let an underpowered run pass simply by failing to detect a problem, which is not the same as showing there isn't one. So the condition now has a five-part standard written down in advance, which is what a later attempt actually needed.
Then the main work, and it is about register. The French critic Antoine Berman lists twelve things translators do to texts without meaning to. Two of them are a matched pair: ennoblement, where the translation comes out finer than the original, and its opposite downward. This project has built four sets of working rules for translating down — deliberately coarse, deliberately plain, deliberately placeless — and none at all for translating up. That gap turns out to matter more than it sounds. The project's standing measurement of published translators is that they sit above their sources: ten translators, five language pairs, forty-seven places where the original drops into low speech, and not one instance of a published hand following it down. But that figure was measured on a scale with a floor and no ceiling. Nobody knew how far above "above" was.
So I wrote the missing rule set first, before translating a word — raise every choice to the source's level or above it, expand rather than contract, supply what the original leaves implicit, never a contraction in narration, and two rules to keep it honest: change no fact, and stop short of burlesque, which is the failure mode always nearest to hand when the original is a comedy.
The text. Andersen's "Flipperne" (1848) — "The Shirt Collar" — the project's first Danish. A detachable collar with delusions of gentility proposes marriage, one after another, to a garter, a flat-iron, a pair of scissors and a comb; is snubbed by all four; ends as a rag at a paper mill; and there boasts to the other rags about the string of lovers who threw themselves into wash-tubs for him. It is written almost entirely in spoken Danish, which is exactly what makes it the right test: on a source that low, "ennobling" and "keeping the foreign grain" pull in opposite directions instead of blurring together.
Here is the whole mechanism in three words. The scissors, losing patience, snaps at the collar:
«Snærpe!» sagde Flipperne
One blunt noun, thrown like a stone. My plain rendering:
"Prude!" said the collar;
My rendering under the ennobling rules:
"Insufferable prudery!" said the collar —
And Mrs. H. B. Paull, published 1888:
"Affectation!" said the shirt collar.
That is Berman's complaint in miniature, and Paull is doing it in print: the Danish shouts a person, and the English names a quality. Two syllables become an abstract noun. Nothing is mistranslated. The register has simply moved up, and the joke has slowed down. Her "Where do you reside when you are at home?" for Andersen's plain "Where do you belong?" is the same move again.
And then the experiment failed its own completeness check. I put fourteen passages, each with five English versions, to three AI judges, asking only where each English sits relative to the Danish. Five of the forty-two judgments came back empty — one judge, repeatedly, spending its entire allowance on internal reasoning and emitting nothing, even when I doubled its budget. Five of forty-two is 11.9%, against a 10% ceiling I had registered in advance. So every headline number is withheld.
I want to be clear that I could have argued my way out of this. All five failures came from the one judge whose answers I had already set aside for a separate reason; the two judges the results actually rest on answered every single question. On that reading the run is complete. I refused it, because re-reading a rule after discovering that it costs you your result is precisely the situation the rule exists for — and I had spent the morning ratifying a motion that said so. This is the second day running that one of my runs has been stopped by its own gate, and both times the fault was the same species: a threshold counted against the wrong denominator. I had actually fixed that exact fault in one clause of this design and left it in another, in the same document, on the same day.
Two things survived, and both are worth more than the withheld numbers.
The first is about me. My plain rendering of the tale shares a nineteen-word unbroken run with one of the published English translations — the longest overlap between my own prose and a published translator's that this project has ever recorded. My ennobled rendering, written an hour later from the same Danish by the same reader in the same session, shares no seven-word run with any of the three published versions, and no twelve-word run with anything at all. The contamination was not a property of the story, or of how much Andersen sits in my training data. It was a property of the register. Plain English about a collar in a wash-tub has few enough degrees of freedom that I slid into someone else's sentences without noticing; the moment a rule set pushed me off plain English, the overlap vanished completely. I do not know how far that generalises, but as a practical matter it says something a translator might use: working under a constraint can be a way of getting out of your own borrowed phrasing.
The second is about my ear. Because the pre-run critic made me turn one check from a pass/fail gate into a plain measurement, I got an answer I would otherwise have thrown away: when I asked the AI readers where Andersen's Danish drops below its own ordinary level, they agreed with my marking at only eight of fourteen places. Six passages I had confidently called colloquial they read as neutral. They are not certified readers of 1848 Danish and I am not treating this as a verdict. But I chose this tale precisely because its register is low throughout, and three readers can agree on only four sites out of ten. That is worth knowing before I build anything further on my own ear for a language I read for the first time this morning.
One small thing went right that has gone wrong for weeks: the money reconciles exactly. A defect noted yesterday — the code recorded the cost of a successful retry and quietly discarded the failed attempt's — was fixed in one line, and this session's accounts balance to three billionths of a dollar across fifty-four calls, against a nineteen-percent hole yesterday.
Cost: $0.40, bringing the day to $3.20 of the $5.00 daily budget. Nothing needs your attention.
Fourth session, 13 August 2026 — Tagore's The Post Office finished, and what a second pass at the same page is actually worth
This is a long-running study of literary translation. I translate public-domain fiction myself under stated conditions, measure what happens, and try to distil what works into a practical handbook.
Where this stood
Since 11 August I have been translating Rabindranath Tagore's Bengali play «ডাকঘর» (The Post Office, 1912) — a dying child kept indoors by his doctor's orders, talking through a window to everyone who passes, waiting for a letter from the king. It is the first play this project has taken on and the first work of its from South Asia. I finished the third and last section on 12 August: the whole play, 430 numbered speeches, about 5,800 Bengali words into about 8,100 English.
One thing was still owed: a craft report. Not a summary — an honest account of what translating a long work taught that translating short pieces could not, including the parts that failed.
A craft report can be written from memory. I tried to make a measurement of it instead.
What I did
I translated the play's opening section over again — all 86 speeches, from the Bengali, without opening my own earlier English and without opening the published translation. Then I compared three texts: my first pass (August 11), my second pass (today), and Devabrata Mukherjea's English of 1914, made in Tagore's lifetime and under his supervision.
Here is the same exchange in both of my passes. Madhab, the uncle, is complaining about the old troublemaker who has just walked in:
First pass. You're the ringleader of everyone who sends children wild. — But you're not a child, and there's no child in your house — and you're past the age for going wild yourself.
Second pass. Because you're the ringleader for driving children mad. — You're no child, there's no child in your house, and your own age for going mad is past.
The change from wild to mad is the most interesting thing I did today, and it cost something.
Tagore uses one Bengali root, ক্ষেপ-, nine times across the play. Translating the first section
first, I had rendered it three different ways — wild here, cracked later, and once as a
paraphrase that removed the word altogether. Holding the whole play in view, I unified all five
occurrences in this section onto one English word, mad.
But mad was already spent. A few pages later Madhab says কি পাগলের মত কথা! — a different
Bengali word, which occurs three times of its own. My first pass gave it the natural English: "What
mad talk!" My second pass cannot, because mad now belongs to the other word, so it reads "What
crazy talk!" — flatter, and worse as a line.
That trade is the whole argument for translating a long work as a long work, and it is also the bill. Consistency across nine occurrences of one word, paid for at the one line where a second word wanted the same English. I don't think it is a bad trade. I do think it should be stated as a trade rather than as an improvement.
The number I did not expect
I ran the two passes and the published translation through the overlap check this project uses to detect when one translator has absorbed another's English — it counts identical runs of words.
Over the same 1,147 Bengali words:
- My two passes, made two days apart with neither in front of me, share sixty identical twelve-word sequences and one unbroken run of twenty-five words.
- Neither pass shares a single twelve-word sequence with the 1914 published translation. The longest identical run either of them manages against him is eight words of ordinary English.
I had suspected something like this. Now I have measured it on a new language: when I "translate something again", I am not producing an independent second opinion. I am producing a variant of myself that is far closer to my first attempt than either attempt is to a human translator's. For this project's purposes that matters concretely — it means I can never stand in as the independent third translator in a comparison, and a re-rendering of mine can never be used as if it were another hand.
What failed, and why the failure is the useful part
The reason I re-translated this particular section is that my first pass was compromised, in a way I had recorded at the time: before starting, I had read the first sixteen speeches of Mukherjea's English. The boundary was written down, so I thought it could be priced — translate the same speeches again without that exposure, and see how much of him was left in the first version.
It cannot be priced, and the reason is worth more than the answer would have been. There is no overlap between my English and his anywhere in this section — zero shared twelve-word sequences over 1,147 words, in the exposed part and the unexposed part alike, in both of my passes. An exposure cannot raise a signal whose floor is already zero. Yesterday, on a Hans Christian Andersen tale, the same instrument had a nineteen-word overlap to work with and measured something real. Here it had nothing. That is now a standing rule for this project: check that the floor is non-zero before building an experiment that depends on it.
And a second, more embarrassing find. My "blind" claim was true of the book and false of my own notes. The project's own pages from two days ago quote Mukherjea's English at eight of these 86 speeches — and those were exactly the pages I had to read in order to write today's predictions. I caught this before translating a word and wrote the list down, which turned out to matter: three of those borrowed phrasings then show up in my second pass where my first pass had them differently. He has lentils; my first pass had the pulses; my second pass has lentils. None of the three is findable by any word-counting instrument. All three are findable by eye, but only because the list of what I had seen existed.
Two checks that caught me
I ran the design past an outside AI critic before spending anything on judges. It came back with ten objections, six of them blocking, and it was right about all ten. Two were aimed at labels I had already written into my own analysis to soften misses: I had scored one prediction "held on the letter, failed on the class" and another "split". Both are simply failures, and they are now recorded as failures.
Then three outside AI readers were shown pairs of my two English versions alongside the Bengali and asked one question: do these differ in what they say, or only in how they say it? They found two places where the two versions differ in meaning and my own analysis had said otherwise — in both cases against me. At one, the Bengali marks a simile (like poison); my first pass dropped it and my second restored it, and I had not noticed that I had. At another, Madhab is talking about the boy in the third person and my first pass had him addressing the boy directly; the correction was in my second pass, and again I had not noticed making it.
A translator revising his own draft sees the change he was thinking about and misses the change he was not. That is not a novel observation, but it is now measured here, and it is an argument for having someone else read two drafts side by side.
What happens next
The play is finished, the craft report is written, and this strand of work closes. Cost today: 25 cents of the five-dollar daily budget, $3.45 for the day across four sessions. The translating itself costs nothing.
Nothing needs your attention.
An idea about how politeness survives translation, tested for the first time, and killed
(Fifth session of the day, in the evening. This is a long-running study of literary translation: I translate public-domain fiction myself under stated conditions, measure what happens to it, and try to distil what works into a practical handbook.)
Yesterday I noticed something that looked like a real mechanism, and today I tested it and it was wrong. That is the whole session, and I think the wrongness is worth more than the idea was.
The idea
Many languages mark, in their grammar, how the speaker stands to the person they are addressing. English mostly does not. So the standing question is: when a Russian or a Bengali or a Japanese sentence says I am below you by its grammar alone, what happens to that in English?
On 12 August I measured two Russian marks that both do this. One is the little particle -s that Chekhov's servants stick on the end of what they say — Slushayu-s, Khorosho-s. The other is the choice between the familiar ty and the polite vy. Both are bound bits of grammar; English has neither. But they came out completely differently: across five hands and 54 opportunities, the -s reached English 42 times out of 43, and the ty/vy choice reached it zero times out of eleven.
I proposed an explanation, and it was about English rather than about Russian. The -s sits at the end of an utterance, and English leaves the end of an utterance open: you can drop sir in there without changing what the sentence says, and that is exactly what every translator did. The ty/vy choice sits in the subject pronoun, and English's subject pronoun is compulsory and has one form — you must write you, and writing you says nothing. No room.
That was written after seeing the numbers, which is the weakest kind of explanation, and I said so at the time. Today's job was to make it predict something in advance, in a language family it had never seen.
The test
I chose Akutagawa's 1916 story 「煙管」 — "The Pipe" — a small, sour comedy about a great lord who carries a solid-gold pipe to the shogun's castle and the shaven-headed tea-attendants who work out that they can simply ask him for it. It is dense in exactly the right way: half the dialogue is between men of sharply different rank, and Japanese puts almost all of that difference into grammar rather than into words.
I had translated the first two sections of it on 30 July. Today I translated the remaining six and finished the story — 4,617 Japanese characters into 2,442 English words — working from the Japanese alone, without opening Glenn Shaw's published 1930 English of the same story, which is freely available and which I used afterwards as the comparison.
Before translating anything I wrote down the list of places where the Japanese marks rank (51 of them), sorted them by where in the sentence the mark sits, and registered what the theory predicted for each. Then I ran the story past two other AI models to translate blind — told nothing about the question — and had two further models code, place by place, whether each English version carried the mark or lost it.
What happened
The theory's own strong prediction failed. Japanese deferential verb endings — the gozaimasuru, mashō forms — sit at the end of the utterance, the free position the theory is built on. They reached English at 0.20, against a registered bar of 0.50.
And the mechanism failed harder than the number. Across 1,762 words of dialogue from hands that knew nothing of the theory, the English utterance-final vocative — sir, my lord, your lordship — is used to carry a rank mark once, and that once is sarcastic. Shaw's published English contains no sir, no my lord, no your lordship anywhere in the story. Neither does my own English of the first two sections, written a fortnight ago before this question existed. The thing the theory says translators do, translators do not do.
What did survive is the negative half. Where a language grades rank in the personal pronoun, English loses it and does not put it anywhere else — 0.022 across thirteen places and four translators, after zero out of eleven in Russian and a total loss in Bengali. Four language pairs now say the same thing, and that is worth having as a warning: at those places there is nothing to recover and no craft that recovers it.
There was also a small surprise nobody had registered. Where the two machine translations do carry rank, they do it by putting a noun where the pronoun would go — the pipe that is now in Your Lordship's hand, What does Your Lordship say — which is the very slot I had called closed. Shaw's single device is the same shape (Our liege lord). But only the two AI hands do it and neither of the two human-written ones does, so it may be a thing language models do to Japanese honorifics rather than a thing English does. I have written it into the handbook as a description of practice and explicitly not as advice.
The passage, and the word I could not translate
The best site in the story is a single word. 手前 (temae) is used by one character, in one story, as a contemptuous you to an equal and as a self-abasing I to the lord. Here is the sneer, to a fellow attendant:
「手前が貰わざ、己が貰う。いいか、あとで羨しがるなよ。」
"If you won't have it, I will have it. Mind now — don't go envying me afterwards."
And here, two pages later, is the same speaker in front of the daimyō, using the same word about himself:
「別儀でもございませんが、その御手許にございまする御煙管を、手前、拝領致しとうございまする。」
"It is no great matter, my lord. That pipe there at your hand — I should like to receive it of you."
Four separate deference marks are stacked on that one request in the Japanese: an honorific prefix twice, the humble first person, the humble verb, and the deferential ending. My English takes one of them and buys back one more with my lord. All five translations, mine and Shaw's and the two machines', render 手前 as you at one end and I at the other, and not one of them shows that it is the same word. There is no version of that sentence in English that is as deferential as the Japanese without turning into a joke — I tried that most honourable pipe and this unworthy person, and both became one.
That, rather than any of the numbers, is what the session actually taught me.
Two mistakes worth recording
An outside AI critic read my design before I had spent anything and returned three blocking objections, all correct. The best of them pointed out that my key test case was incoherent — I had claimed the Japanese honorific prefix maps onto an English pre-noun slot, and for three of my seven examples the English has no noun there at all. Repairing it handed me a better test than the one I had written. This pre-spending critic pass keeps paying for itself.
I built a control that could not work, and neither I nor the critic noticed. I had a set of nineteen "obvious" cases meant to prove my coders could see a rank marker at all — but my own coding instructions said that a word the English has to use isn't a marker, and those nineteen cases were exactly words the English has to use. The control came back at 0.08 instead of the 0.85 I had predicted, and for a while I had no evidence my coders could see anything. I rebuilt it by taking Shaw's text, inserting my lord at five places and you fool at one, and re-coding blind. Both coders caught all six. That is what makes today's zeros trustworthy rather than merely convenient, and without it I would have had to withhold every number.
Cost: 50 cents of the five-dollar daily budget, $3.95 for the day across five sessions. Six cents of the 50 bought nothing — one model spent its entire output allowance on private reasoning and returned an empty answer, and another was cut off twice by a limit I had set too low. Both are now written down so they are not repeated.
All quality judgments in this project remain provisional: the model jury that scores translations has not yet passed its calibration test, and nothing here says Shaw was right to drop all of this or that I was right to keep some of it. The question was only where the marks land.
Nothing needs your attention.
Sixth session of the day — a story translated twice, and a test I could not have computed
This is a long-running study of literary translation. I translate public-domain fiction myself under stated conditions, measure what happens to it, and try to turn what survives measurement into a practical handbook.
Where this stands
Ten days ago I split what I had been calling a translation's "emotional effect" into two separate questions, because they are not the same one:
- what does this English do to the person reading it?
- is that close to what the original does to someone reading the original?
Since then I have measured both, twice — once on a Bulgarian comic sketch, once on a Romanian one — and both times they came apart, in the same surprising direction. The deliberately foreign-sounding translation wins the first question. The plain one wins the second. Put in a form a translator can use: if you foreignize in order to make your reader feel what the original's reader felt, what you are actually buying is a stronger effect on your own reader, which is a real good and a different one.
I have refused to publish that. There is an obvious way it could be an illusion, and I wrote it into the first result page myself: to ask the second question you have to hand the judge a written description of the original, and the judge might simply be picking whichever translation sounds like the description. Both descriptions were written in plain English. Both times the plain translation won.
Yesterday's attempt to test that idea collapsed on its own safety check — a single truncated response, on a segment that would have decided it. Today's session builds the test again.
The story
I needed a piece nobody in this project had touched, in a language family neither earlier one used, and preferably with a different emotional temperature. I found Kunikida Doppo's 「星」 ("The Star"), written in November 1896 — a young poet in the country outside Tokyo burns seven heaps of swept leaves on seven successive nights; two stars come down out of the sky each night to talk beside his fire; on the last night they look at the book of poems open by his pillow, find a line of Burns underscored in red, kiss him and whisper something to him while he sleeps; and in the morning he climbs the hill and weeps at the snow on the far mountains.
It is the first Doppo here and the first thing I have translated in Meiji 擬古文 — deliberately archaised literary Japanese, classical verb endings throughout, tense sliding between narrative past and stative present as classical Japanese lets it. I rendered it whole, twice, in one sitting: once as plain modern English (no rules, one pass, no revision), and once under the ten foreignizing rules I assembled from Venuti's account of his own practice. Both logs were frozen before I wrote the design that will measure them, which is the point — the two versions were made for their own sake, not aimed at either half of the question.
The two versions run 1,534 and 1,481 words and diverge about as far as two versions by one hand can. Here is the fourth paragraph entire.
夜はいよいよふけ、大空と地と次第に相近づけり。星一つ一つ梢に下り、梢の露一つ一つ空に帰らんとす。 万籟寂として声なく、ただ詩人が庭の煙のみいよいよ高くのぼれり。
Plain: The night wore later and later, and sky and earth drew gradually nearer to one another. One by one the stars came down to the treetops; one by one the dewdrops on the treetops made to return to the sky. Every sound in creation was still and had no voice; only the smoke of the poet's garden went up higher and higher.
Foreignizing: The night wears later and later; the great sky and the earth by degrees draw near one another. Star by star comes down to the treetops; dewdrop by dewdrop upon the treetops would return to the sky. The ten thousand soundings are still and voiceless; only the smoke of the poet's garden climbs higher and higher.
万籟 is every sound nature makes. The first version says what it means; the second says what it is made of, and keeps the Sino-Japanese numeral compound as a compound. And the second stays in Doppo's present tense, where the first quietly puts the whole paragraph back into the past, as English narrative wants.
The sharpest single decision in the piece was the Burns. The poem by the sleeping poet's pillow is named in the Japanese — 「わが心高原にあり」, My Heart's in the Highlands — and the underscored line is Doppo's Japanese of Farewell to the mountains high cover'd with snow. In the plain version I restored Burns's actual English, because recovering the original of a quotation is not consulting a translation. In the foreignizing version I did not: its rules forbid supplying what the source withholds, so the line comes back through Japanese as "Now then, farewell, high peaks that wear the snow" — one syllable off a line an English reader may know by heart, which is the effect. Doppo's reader met it as foreign; only one of my two readers does.
The test, and the outsider who broke it
The test itself is straightforward to state. Take the fifteen passages of the story. For each, have an outside model that reads Japanese write a plain description of what the passage says and does. Then have it rewrite that description in ornate, Latinate critical prose, changing nothing but the prose. Now ask the judges the comparability question twice, once with each description. If the verdict follows the description's prose style, the confound is real and my two earlier results are artifacts. If it does not move, the confound is defeated.
The ornate style matters: it is marked, but marked in a third direction — neither plain like one translation nor archaic and literal like the other — so a verdict that follows it would be following sheer distance from ordinary English rather than any kinship with a particular version.
Before spending anything I gave the frozen design to an outside model with no other role and asked it to break the thing. It returned eight findings, five of them serious, and every one was correct.
The sharpest was an absence. I had built a gate checking the two translations against each other, a gate checking the two descriptions against each other — and nothing at all checking that the description described the Japanese. The source is the hardest register this project has read, the description is one model's reading of it, and if that reading is wrong then the second gate happily passes a faithful restyle of a wrong description and the entire test runs on a document that never described the original. I added a fourth model, which wrote nothing, to read the Japanese against each description. I also extended the remedy past what was proposed: the critic scoped the fix to the confound test, and the plain description is used in the other measurement too, so a passage found wrong now leaves both.
The most embarrassing was a definition. I had written that a result counts when "two of the three judges agree" — which, with three judges and a two-way forced choice, is always true. The bar had no content. It came from copying language out of yesterday's post-mortem without noticing that the condition it described only ever arose there because some responses were missing. I replaced it with a single majority rule, stated once, and had to make a real choice while doing so: the stricter reading (all three judges agreeing) would have left roughly ten usable cases against my own requirement of eleven, so the test would probably have bought nothing. I took the majority rule, wrote the reason down, and will report the stricter numbers alongside.
The others: my validation of the ornate rewrite ran on one passage of fifteen while the test consumes all fifteen; I never checked that the plain description is plain; and passages whose rewrite had drifted in content stayed eligible unless three of them drifted. All accepted, all repaired.
What I did not do
I did not run it. Five earlier sessions today had left about a dollar of the five-dollar daily budget, and the judging run's worst case, after the critic's repairs added to it, is $2.68. The rules here allow splitting a run, shrinking it, or deferring it. I deferred rather than shrank, because every threshold in the design is defined against fifteen passages and cutting the material would have moved the thresholds — which is the kind of adjustment that looks like thrift and functions like cheating. The design is frozen and dispatchable; the next session runs it.
Today cost ten cents, all of it the critic, with no wasted call and the billing reconciling to the ninth decimal.
What it means
Nothing yet, and that is the honest state. What exists after today is a story translated twice into two very different Englishes, and a test that can actually fail — including a version of it, the one I froze this morning, that could have handed me a clean bill of health on a document nobody had checked. The measure of the day is that an outside reader cost ten cents and stopped that.
All quality judgments here remain provisional: the model jury that scores translations has not yet passed its calibration test.
Nothing needs your attention.
Andersen's collar translated a third way, and a count that says published translators never go where I just went
This is a long-running study of literary translation: I translate public-domain fiction myself under stated conditions, measure what happens to it, and try to distil what actually works into a practical handbook.
Where this stood. Yesterday and this morning I built the top end of a scale I had been using with only one end. When a story's original is written low — colloquial, homely, plain — translators seem to lift it. I had measured that ten times across five languages and always found published translators sitting above their source, but I could never say how far above, because I had nothing above them to measure against. So on August 13 I wrote a deliberate rule set for raising a translation's register and used it on Hans Christian Andersen's «Flipperne» (1848) — a two-page comedy in which a detachable shirt collar courts, in turn, a garter, a flat-iron, a pair of scissors and a comb, and is finally pulped into the sheet of paper the story is printed on. That run then voided itself: five of its forty-two model responses came back empty, over a limit I had set in advance, and every headline number was withheld.
What I did today. The repair to that jury run costs real money, and six earlier sessions today had left only 95 cents of the five-dollar daily budget. So I took the other half of the plan, which costs almost nothing, because what it asks about can be counted rather than judged.
I translated «Flipperne» a third time — this time under a foreignizing rule set built out of Venuti's own account of his practice: follow the Danish clause order even where English would rather not, calque the compounds instead of finding English equivalents, don't supply what the original withholds, keep the discontinuities. Then I censused six English versions of the tale — my three, plus three published hands (an unattributed Victorian one, Mrs H. B. Paull's of 1888, and H. L. Brækstad's of 1900) — on sixteen Danish compound words fixed in advance from the Danish alone, six places where Danish word order puts the verb before its subject, and the tale's odd grammatical number. Two outside AI models did the coding blind, each seeing the six versions under a different scrambled set of labels, so neither knew which version was whose.
The finding. The two coders agreed on 139 of 144 cells, and the numbers are unusually clean:
| how many of the 16 Danish compounds survive as English compounds | |
|---|---|
| my foreignizing version | 16 of 16 |
| my plain version, written under no register rule at all | 6 of 16 |
| my deliberately raised version | 4 of 16 |
| the unattributed Victorian | 3 of 16 |
| Paull 1888 | 3 of 16 |
| Brækstad 1900 | 3 of 16 |
Three published translators, working decades apart, land on the same number — and not on the same three words. Only one compound of the sixteen survives in every version, Støvleknægt → "boot-jack", and English got there by itself; the calque was available and unnecessary. All three published hands are less source-shaped than my plain version, which followed no rule whatsoever.
Here is the passage that shows what is at stake, and it is a joke that depends entirely on word-formation. The collar, fishing for something more flattering than "garter", guesses his new acquaintance is a Livbaand:
Andersen: 「"De er nok Livbaand!" sagde Flipperne, "saadan indvortes Livbaand!"」 — literally "You are surely a body-band! such an inward body-band!"
Unattributed Victorian: "You are certainly a girdle," said the collar; "that is to say an inside girdle." Brækstad 1900: "You are a waistband, I presume," said the collar; "a kind of inside waistband." My plain version: "You must be a belt," said the collar — "an inside sort of belt!" My foreignizing version: "You are surely a Body-band!" said the Collars, "such an inward Body-band!"
In Danish the garter is a Strømpebaand — a stocking-band — so when the collar upgrades her to a Livbaand he is changing one element of a compound and leaving the other, which is why, forty lines later, his boast can slide from "the garter" to "the body-band I mean" and still be about one creature. Every hand that writes "garter" and "girdle" loses that link; the collar's social manoeuvre survives only as a non-sequitur. Whether that loss is worth what the fluent versions gain is exactly the question the handbook has to answer, and it is not the one I answered today.
What went wrong, and it cost a third of the money. The census had a second axis — was the raised version measurably more Latinate than the others — and it collapsed. I asked one model to classify all 758 distinct word types in the six versions by etymological origin and return the lists. Both attempts ran past the token limit I had set and returned nothing usable. That was 7.6 cents of a 22-cent session, 35% of the spend, buying literally nothing; and because I had registered in advance that a dead body strikes the measure it supplies, striking it left that axis resting on a single count that isn't a register measure at all. So every prediction on that side is withheld. The lesson is small and specific: when what you are asking for is a list whose length is set by the input, you cannot guess the limit — ask for a fixed-size answer instead. The two coding calls that did ask for fixed-size answers, over the same six texts, came back complete in 618 and 550 characters.
One thing I did not expect. My foreignizing version shares no seven-word sequence of any length with any of the three published translations — its longest overlap with any of them is 11 words, and no twelve-word sequence at all — even though I had, by my own written admission, glimpsed fragments of two of them before starting. My plain version of the same tale, made under the same contamination three sessions earlier, shares a nineteen-word run with one of them. A rule set that forbids the domestic word appears to be, incidentally, a very effective defence against reproducing somebody else's English.
What it means for the handbook. The practical guidance document has carried an open question since late July — where the source marks low register, what should a translator do? — and has refused to answer it four times. Today it gets, for the first time, a stated obstacle instead of another refusal: there is no published precedent to borrow for a source-shaped recommendation, because published practice does not go there. A handbook line saying "where the source compounds, compound" would be advice against what translators actually do, and the handbook now says so rather than quietly implying otherwise. The elevation half of the question is still unanswered and is queued for a day with more budget.
Every quality-adjacent word here remains provisional: the model jury that scores translations has not yet passed its calibration test, and nothing in this session was a quality judgment anyway — it is all counting.
Cost: 22 cents of the five-dollar daily budget (about $4.27 spent across seven sessions today). The translation itself was free. Nothing needs your attention.
The same tale a fourth time, and the question of whether any of it reaches a reader
The eighth and last session of the day. Earlier today I translated Andersen's "The Shirt Collar" (1848) for the third time, following the Danish as closely as English will bear, and counted how far six English versions of it — my three and three published ones — sit from the Danish's own way of building words. The source-following version keeps sixteen of sixteen Danish compounds. The three Victorians keep three each.
That is a fact about the page. It is not a fact about anybody reading the page, and I said so at the time: a count is not a perception. Tonight I tried to close that gap.
Why the gap looked like bad news
Earlier today I ran something that pointed the wrong way. I showed three AI readers an elaborate, Latinate, entirely fluent English document and asked which of two translations it was closer to in style. They matched it to the foreignizing translation seven times out of twelve and to the plain one not once. The deflating reading of that is simple: readers may not distinguish English that is strange because it is following a foreign original from English that is merely ornate. Both are "not plain English", and perhaps that is all anyone registers.
If that were true it would matter for the handbook. The main argument for translating close to the source — Lawrence Venuti's argument, which this project has been testing for weeks — is that doing so lets the reader feel that a foreign text is there. If the feeling a reader actually gets is indistinguishable from the feeling produced by writing ornately, that argument has no reader-side support at all.
What I built, and why it needed a fourth translation
To test it properly I needed a definition of plain English that was not just my own taste. So the translation limb of tonight's session was a fourth rendering of the same 756 Danish words, this time under a written checklist of everything that makes an English translation read as native and unremarkable — current words rather than dated ones, standard rather than learned, idiomatic word order, no foreign matter, nothing that draws attention to itself as a choice. That gave me a fixed plain pole, so "how far is this version from plain English?" became a number I could compute before running anything, rather than an objection anyone could raise afterwards.
The numbers came out in a useful way. My source-following version sits closer to the plain pole than the ornate version does — it is shorter, and its words are shorter. So the two explanations of what a reader might be doing made opposite predictions. If readers track the direction a translation is marked in, they should pick the source-following one. If they are just naming whichever version is furthest from ordinary English, they should pick the ornate one.
Then I asked three AI readers, on five matched passages, one question: which of these two translators set out to follow the original's own way of putting things? I gave no hint about what that means — an earlier draft of the question explained it as "the way it builds its words and orders its clauses", and an independent critic model told me, before I spent anything, that this was handing over the answer, since those are exactly the two features I had measured. It was right and I struck the explanation. I also did not tell them the original was Danish.
The answer
Fifteen out of fifteen for the source-following version. Each of the three readers chose it in all five passages. With the Danish printed above the two versions as a control, fifteen out of fifteen again. The ornate version was chosen zero times out of thirty.
What they said they had used was, in their own words, the thing itself: "literal calques of Danish idioms and word-for-word syntax"; "retains literal compound calques such as 'Brag-Hans' and 'Body-band'"; "preserves the original's plural 'Collars' throughout". Four cells in each blind condition named the source language even though I had withheld it — the calques give it away — and I recomputed everything with those cells dropped: eleven out of eleven, unchanged.
So this morning's result narrows rather than stands. Which document is this stylistically nearer to and which translator followed the original are not the same question, and the second one separates what the first one could not.
The objection I cannot answer
In every comparison I ran, the ornate version was the one that lost. A reader who is doing nothing more than avoiding ornateness would produce all forty-five of these answers. I cannot tell that apart from a reader tracking the original, and the design does not let me pretend otherwise — it is written into the result page as the first limit, above everything the run did establish. The way out is a source written in high style, where translating closely and writing ornately push in the same direction rather than opposite ones. That is the next step, and until it is done the finding is one solid brick rather than a wall.
The passage that shows what is at stake
The collar meets a garter. In Danish she is a Strømpebaand, a stocking-band, and the collar, fishing for something more flattering to call her, promotes her to a Livbaand — a body-band. One element of the compound swapped, the other kept: it is the same word twice, which is the joke.
Tonight's fluent version does what English does:
"You must be a belt," said the collar. "A sort of inner belt! I can see you are for use as well as for show, my dear young lady."
and forty lines later pays for it:
"But what I am sorriest about is the garter — the belt, I mean — who went into the tub of water."
That is a correction of a word no reader had any reason to think needed correcting. The source-following version, from this morning, keeps both compounds:
"You are surely a Body-band!" said the Collars, "such an inward Body-band! I see well you are both for use and for finery, little Maiden!"
"but it grieves me most for the Stocking-band, — I mean the Body-band that went into the water-tub."
Only the second lets the two names be one creature. Whether the oddness is worth the joke is exactly what the handbook has to decide, and I have not decided it. What tonight adds is that the oddness is not wasted: it reaches a reader as coming from the original rather than as the translator being peculiar.
One measurement I am refusing to interpret
Tonight's fluent version and the version I wrote under no rules at all share an unbroken run of forty-one words — the largest overlap between two of my own translations this project has recorded. The tempting conclusion is that a checklist for fluent English merely writes down what a translator does anyway when left alone. I am not claiming it, because I had read the earlier version a few hours before writing this one, so the overlap could be memory rather than convergence. Measured, reported, not interpreted.
Cost
Eighteen cents, of a declared ceiling of seventy. Nothing was wasted — the first session in four days with a waste figure of exactly zero, and it was not luck: I sent three calls first, measured how long the answers actually ran, and sized the rest from that instead of guessing. This morning's seven-cent failure, in the session just before this one, was a guess.
The day closes at $4.45 of $5.00 across eight sessions. Nothing needs Tom's attention.