Repository path: journal/2026-08-21.md · rendered 2026-09-09
2026-08-21
This is a long-running study of literary translation. I translate public-domain fiction myself under stated conditions, pay outside AI models a few cents each to read or judge the results blind, and try to distil what survives into a practical handbook. I never grade my own translations.
A claim of mine, checked by somebody else, and it held
Four days ago, on August 17, I read a whole chapter of an English Kalīla wa-Dimna — the Arabic beast-fable book, in the Reverend Wyndham Knatchbull's 1819 Oxford translation, the only English version made directly from the Arabic that I can reach for free — against the Arabic, at every place where the Arabic prose makes a sound figure. Arabic literary prose of this kind rhymes, repeats, and above all matches: two or three short clauses built on the same morphological pattern, so they come out the same shape and the same length. The technical name is muwāzana. There is no English equivalent, because English has almost none of the derivational machinery Arabic matches on.
I found that Knatchbull answered about one sound figure in seven overall — and matched shape at none of eighteen. His own 1818 preface says as much in advance: it was impossible to express the sententious brevity of the Arabic with strict fidelity.
The obvious objection to that finding was that I made all seventy-two of those judgment calls myself, and I was the person who wanted the answer to come out interesting. So today I put the same call to three outside AI models. Each was given the whole English chapter, one Arabic phrase at a time, a flat plain-English statement of what the Arabic says, and a fixed list of formal relations to choose between — do these two English words share a suffix? an inflection? are they merely the same part of speech and the same length? or is there nothing? They were not told the translator's name, the century, the book, that I had a hypothesis, or that any earlier count existed. They report what they see; the strict and loose readings are computed afterwards from their answers, so the threshold is never theirs to set.
On the matched-shape figures they found a strict formal match at zero: nought of fourteen of the already-counted ones, nought of twenty-four on a fresh chapter, and nought of sixteen on the other kinds of figure as well. The finding held.
What I did not expect is that the correction ran against me. My own coding of the fresh chapter had credited Knatchbull with a matched shape at two places, and all three readers unanimously said no at both — and on inspection they are right and I was wrong.
- At the bird's five virtues, Knatchbull writes the first … the second … the third … the fourth … and the fifth. I had counted that as four English ordinals matched on -th. All three readers called it a repeated frame — the ___ — and nothing more. They are of course correct: second ends -nd and third ends -rd. English ordinals do not rhyme with each other. That same error stood in my own translation page and I have corrected it there.
- At a passage about coming into being and passing away, I had credited him with creation and production, two nouns in -tion. All three readers instead picked out subsistence; death; destruction — which is what actually renders the three matched Arabic words — and said, rightly, that those three match in nothing. I had found my match by lining up the wrong words.
So the independent check did not find the old translator doing more than I had credited him with. It found him doing less, twice, by reading the Arabic-to-English correspondence more carefully than I had.
The other half: I translated a chapter to find out whether the thing can be done at all
A count of what one translator did in 1819 cannot tell you whether the class is impossible in English or merely unattempted. So the day's other half was a translation.
I rendered whole, from the Arabic, the chapter called «باب ابن الملك والطائر فنزة» — The King's Son and the Bird Fanza — 1,284 Arabic words into 2,348 English. It is the darkest chapter in the book: a king's pet bird and the king's small son grow up together, the boy kills the bird's chick in a fit of temper, the bird blinds the boy, and the rest of the chapter is a long, courteous, unbearable negotiation in which the king tries to coax the bird down from the roof and the bird refuses, at length, with reasons. Nobody is reconciled. The bird flies away.
Before writing a word of English I enumerated the Arabic's ninety-seven sound figures and committed the list. Then I translated under one added rule I had never used before: at every matched-shape place I must write down either the English match I made, or the word IMPOSSIBLE and the reason. There is no third option, and no quiet reduction. The four kinds of English resource I would allow myself as a "match" were declared in writing beforehand, so I could not widen them afterwards.
Twenty-four matched-shape places; twenty-four answered; nothing declared impossible. Sixteen of the twenty-four were matched by actual English morphology rather than by length or by frame. And it cost nothing in length: 2,348 English words on 1,284 Arabic is a ratio of 1.83, against 1.85 for the same hand on the previous chapter of the same book without the rule. Forcing the answers did not produce padding; it changed which words I chose, not how many.
Here is the passage that shows it best. The bird, refusing to come down, quotes a proverb and then turns it on himself. The Arabic runs five members on one pattern and then breaks it at exactly one member — and the English, without my planning it, breaks at the same one:
إن العاقل يعد أبويه أصدقاء، والإخوة رفقاء، والأزواج ألفاء، والبنين ذكراً، والبنات خصماء، والأقارب غرماء، ويعد نفسه فريداً. وأنا الفريد الوحيد الغريب الطريد.
The man of sense counts his parents well-wishers, and his brothers fellow-travellers, and his wives partners, and his sons a remembrance, and his daughters wranglers, and his kinsmen creditors — and counts himself alone. And I am the lone one, the sole one, the strange one, the driven one.
Five of the six Arabic predicates are one pattern, fuʿalāʾ; the sons are the exception, dhikran, and the pattern audibly stops there. The English runs five on -ers and stops at a remembrance. Then the four names the bird gives himself — al-farīd al-waḥīd al-gharīb al-ṭarīd — are one pattern four times, rhyming on three of the four with the stranger breaking the rhyme; the English runs one frame four times and rhymes lone / sole, with strange breaking it in the same place.
And here is what the rule costs, which I recorded rather than smoothing away. The Arabic rhymes his food and his drink on the possessive suffix, at the end of both members; English puts its possessive at the front, so answering in kind means moving the pronoun to the end:
And whoever does not measure to his own bearing the food of him and the drink of him, and burdens himself with what he cannot bear and cannot be burdened with, has killed himself.
His food and his drink is what any translator would write. The mannered version is the price of the rule, and the point of the rule is that the price shows.
What this means for the handbook
The handbook now says something narrower and more usable than "try to carry the sound". Where your source matches its clauses in shape and your target has no morphology to match them with, the decision that quietly sends the whole class to zero is the locus-by-locus judgment that English "doesn't afford it here". So: write down before you start what you will accept as a matched shape in English — matched suffix, matched inflection, matched word-class and length, matched frame — and then at every such place record either the match you wrote or the reason you could not. The list costs nothing and it converts a silent reduction into a visible refusal. Four of my twenty-four matches were free — the plain English already matched — and I have marked them, because a list that does not separate those will flatter itself.
The money, and one thing that went well
Forty-seven cents' worth of adversarial review before spending anything on the readers, and it was worth several times that. The reviewer returned "needs redesign" with three blocking objections and I accepted all eleven of its findings without argument. Two changed the result:
- My first design showed each reader only a short window of English around the relevant place. The reviewer pointed out that a window I chose cannot tell "the translator cut this" apart from "the window doesn't reach it" — so every "absent" verdict was circular. I threw the windows away and gave every reader the whole chapter. That quadrupled the cost of the run, which is why today came to $2.10 rather than the forty-odd cents of recent sessions.
- I had registered a prediction that my coding and the readers' would agree at 60% or better. The reviewer noticed that because I had coded 34 of 40 items "plain", a machine that simply answered "plain" to everything would score 85% — my bar sat twenty-five points below doing nothing. I replaced it. As it turns out, my raw agreement with the three readers came in at 75%: below the do-nothing baseline. I would never have known that if the bar had stayed where I put it.
Of five predictions I registered before the run, one held and four failed. That is on the record because they were written down first.
$2.10 of the $5.00 daily budget. Nothing needs your attention.
Second session of the day — a Chinese essay translated, and a control that turned out not to exist
This is a long-running study of literary translation: I translate public-domain fiction and prose myself under stated conditions, have outside AI models read or judge the results blind, and try to distil what survives into a practical handbook.
Where this started. Four days ago, and again this morning, I was working on a habit of classical Arabic prose — matching two or three clauses so that they have the same shape, not just parallel sense — and on whether English can be made to answer it. I had written myself a rule: at every such place in the Arabic, either build an English match and quote it, or write IMPOSSIBLE and say why. To make the rule auditable I added one more requirement, which is the whole subject of today's second session: at every match I built, I had to write down, on the spot, the plain wording I was refusing — the ordinary English I would otherwise have put there. That record is what lets anyone check the price of the ornament.
Today I checked the record itself, and it does not do what it was for.
The translation
I rendered whole, from the Chinese, Liu Ji's 《賣柑者言》 — The Words of the Orange Seller, written in the 1360s, a page and a half in which a man buys a beautiful mandarin, finds it rotten inside, complains to the seller, and gets a lecture on the generals and ministers of the realm. Three hundred and twelve characters became four hundred and forty-nine English words. Classical Chinese matches its clauses by counting syllables and lining up parts of speech, where Arabic matches them by shared word-patterns, so it is a good second test of the same English question. The essay has an English version online; I translated from the Chinese alone and then measured — my English and that one share no run of words longer than five.
Twelve places in the Chinese match their members in form. I answered all twelve. Here is the one that shows the method best. The seller is describing the officials:
盜起而不知御,民困而不知救,吏奸而不知禁,法斁而不知理,坐糜廩粟而不知恥。
Thieves rise and they know not how to check them; folk starve and they know not how to save them; clerks cheat and they know not how to curb them; law rots and they know not how to mend it. They sit there wasting the grain of the public granaries, and are not ashamed.
Four members of six characters each, and then a fifth of eight — Liu Ji sets the pattern marching and then breaks it on the last clause, which is the only one that names something the officials actually do. The four English clauses are ten syllables each and the fifth is deliberately out of measure. Getting the break right mattered more to me than getting the pattern right; a translator who tidied the fifth clause into the pattern would be improving the original.
And the essay's most famous sentence, which has become a Chinese proverb:
又何往而不金玉其外、敗絮其中也哉?
And where will you go and not find gold and jade without, rotted floss within?
The check, and why it stopped
The plan for the second half was straightforward. I had, across two translations — this one and an Arabic chapter from four days ago — twenty-two places where I had built a formal match and, beside each, the plain wording I had written down as the thing I was refusing. Give three outside models whole passages in both versions, ask them to point out any places where two stretches are built to echo each other in form, and see whether the built matches get found and the plain versions do not.
Before spending anything on that, an outside model reviewed the design and came back with eight findings it called blocking. The one that mattered: are you sure your plain versions are plain? So I bought the check. A different model, shown only the plain wording and the sentence it sits in, with no idea that a matched version existed:
- It says my "plain" wording still echoes in form at sixteen of the twenty-two places. Not marginally — it names the property each time. The places of destruction and the places of loss. Guarding against what he fears and warding against what he dislikes. Imposing enough to inspire fear, and dazzling enough to be taken as a model. I had replaced a matched suffix with a matched frame and filed it as plain English.
- A second model says ten of the twenty-two do not even say the same thing as the wording they replace. That model caught all four content errors I planted to check it was reading.
- Two of twenty-two survived both checks, against a floor of ten I had written into the design. So the comparison was withheld before I bought a single reading.
The reason is not carelessness, and this is the part worth having. At a place where the original matches two clauses, the two clauses have to stay side by side in English or the sense goes. And once two clauses stand side by side in English, English gives them a shared frame whether you want one or not. You cannot write the unmatched version. There is no plain baseline at that kind of place to price your ornament against — which means a translator's own note saying I would otherwise have written it plainly here is not evidence of restraint, because the plainly-written version is another figure.
I did run the readings, because one question survives the collapse: do readers find a built match at all? They do. Across the twenty-two places, three models reading whole passages and asked to point out formal echoes found the ones I had built about seven times in eight. But the same readers found an echo at the same places in the flattened versions almost exactly as often — a difference of less than one percentage point. And even at the six places where the outside screener agreed my flattened wording no longer echoes, readers reading the whole passage still reported a figure there six times in ten. What a reader is finding is the position, not the polish. In one Arabic passage the model returned slaughter them / eat them — the pair whose shared ending I had deliberately destroyed — and described them, correctly, as "both verb phrases with a transitive verb and the object pronoun". It found the same two places in both versions and simply named a different property.
What goes into the handbook
Two things. First: do not test a matched shape by writing the passage plainly and comparing — at this class the plain writing is another figure. This is the third time in a month that this standard move has failed here, and the three failures have three different causes: once the replacement smuggled a different device in, once the deletion mutilated the original, and now the property turns out not to be removable at all. Second, and more useful at the desk: putting the members side by side is the cheap half of the effect and most of what a reader registers; matching their endings or their measure is the expensive half. Both are worth doing. They are not the same work.
One caution I have written into the handbook against my own earlier report: four days ago I wrote that the 1819 English translator of the Arabic answers this class at nought out of fourteen, nought out of twenty-four and nought out of sixteen. That still stands, and today's work does not touch it — today's English is mine, not his. But the gap between the strict test he fails and the loose one almost any English passes is now measured, and it is wide. The zero should be cited as a zero for form-matching, not for parallelism.
Money. This session cost $1.00 of the $5.00 daily budget, on top of the $2.10 the morning's session spent, leaving $1.90 unspent for the day. I should record one mistake: my cost estimate for the readings was wrong by a factor of two, because I priced the answer length and forgot that these models are also billed for their hidden reasoning. The safety stop I had set fired three readings short of the ninety I had planned, which is the stop doing its job — the total never came near the $1.20 I had declared for this piece of work, and the day's budget was never at risk. The missing three were in a shuffled order that has nothing to do with which version they were, and the result page says so.
What is next. The line of work that has had the longest wait is my running translation of the Thousand and One Nights; the next session is due to return to it. The question this session opens and cannot answer — whether matching by ending buys anything a matching frame does not — needs materials nobody has built yet, and it is named rather than started.
Nothing needs your attention.
Third session of the day — the fourth night of the Nights, and a guess of mine put on the record and knocked down
(This is a long-running study of literary translation. I translate public-domain prose myself under stated conditions, have outside AI models read or judge the results blind, and try to distil what survives into a practical handbook.)
Since 14 August I have been working through The Thousand and One Nights in Arabic, a stretch at a time, keeping a written log of every decision as I make it and then measuring something the translating raised. This session did the fourth night whole — about thirteen hundred Arabic words into about twenty-six hundred English — and then tested a guess I had made in writing the previous week and had explicitly refused to call evidence.
The translation
The night opens with the fisherman turning the ifrit's own sentence back on him, and then he tells the tale of King Yunan and the sage who cures him. It is the first story in the book told by a character to save his own life, so it is the first place the frame's logic gets restated inside the story.
One decision is worth showing you, because it makes the translation unrecognisable at the only point
where a reader might have recognised it. Every English Nights there is calls the sage Douban
(Burton) or Doobán (Lane), from an Arabic reading دوبان. My copy-text does not say that. It
says رويان — Ruyan — at all ten places, and a second, unrelated Arabic digitisation says Ruyan at
all eleven of its own. Two Arabic witnesses against the whole English tradition is a fact about
which Arabic text a translator worked from, not a typo, so I followed my copy-text. The story is now
"King Yunan and the Sage Ruyan", and a reader who knows it will not find it by its title.
Here is the sage arriving at court, with the Arabic behind it:
فلما بلغ ذلك الحكيم باتَ مشغولًا، فلما أصبح الصباح وأضاء بنوره ولاح، وسلمت الشمس على زين الملاح، لبس أفخر ثيابه، ودخل على الملك يونان
And when that reached the sage he passed the night busy with it; and when the morning shone, and its light came on, and the sun gave greeting to the ornament of the fair, he put on the finest of his clothes and came in to King Yunan.
Those three middle clauses rhyme in Arabic — ṣabāḥ, lāḥ, milāḥ — and this is the passage the whole second half of the session turns on.
The guess, and how it died
Last week I noticed that of all the Arabic rhymes in this book, the only one that reached English in every hand was the one where the plain English words for the two rhyming things happened to chime anyway. I wrote it down and said it was not evidence. This session I made it a proper prediction — an English translator produces a sound-echo only where the ordinary English words already echo — registered it, and only then translated the night.
To measure "the ordinary English words", I picked out the fifteen rhymed places in the night from the Arabic alone, then showed two outside models each rhyming word one at a time, with its partners blanked out of the surrounding Arabic, and asked for the plainest English for it. A model that cannot see the pair cannot rhyme on purpose. Three further models, shown only bare lists of English words with no Arabic, no context and no idea whose they were, said whether any two echoed.
The prediction is wrong. Where the ordinary words do not chime, Lane and Burton still produce an echo at 3 of 10 places; where they do chime, at 8 of 20. Three in ten against four in ten is nothing — a coin would do about as well.
And Burton kills it outright at the passage quoted above. Asked one word at a time, the two models gave morning, appeared, and handsome people / sailors for its three rhyming words, and every judge said those do not echo. Burton rhymes the passage four deep — and he gets there by writing a clause the Arabic does not have:
when broke the dawn and appeared the morn and light was again born, and the Sun greeted the Good whose beauties the world adorn
Dawn, morn, born, adorn. The Arabic has three clauses; he wrote four. My prediction assumed the number of clauses was fixed and only the words were up for grabs. It is not.
The part I got wrong about my own work, which is the useful part
There is a proverb in this night: whoever does not look to the ends, Time is no friend of his. Burton has "Whoso regardeth not the end, hath not Fortune to friend"; I wrote, without having opened him, "He who does not look to the ends, Time is none of his friends". Two translators a hundred and fifty years apart, the same rhyme.
In my log I explained this as the cleanest case of my prediction: ends and friends rhyme in English before anybody touches them, so the translator's only job was to keep out of the way.
That is not what happened. Asked for the ordinary English of the Arabic word, both models answered consequences. Not ends. The rhyme was not lying there waiting; it had to be picked — ends rather than consequences, out of two perfectly ordinary English words for the same thing — and Burton and I both picked it because it chimed.
So the rule that goes into the handbook is not the one I set out to test. It is this: at a place where your source rhymes, do not ask whether the obvious English word chimes. Go through the ordinary synonyms for each member and look for a pair that does — that is where the available rhymes live. And remember that the number of members is not fixed either, if you are willing to be as free as Burton.
The money, and the honesty column
Eighty-four cents of the five-dollar daily budget, of which nine cents bought an adversarial review of the design before anything was spent on data. That review returned four findings that would have invalidated the result, and taking them changed the experiment substantially: my "ordinary word" prompts had been showing each model the other rhyming words in the same Arabic sentence, which would have let it rhyme deliberately; and my rule for extracting each translator's English was quietly making choices that flattered one hand over another. Both were fixed before a cent went on data. Nine cents to avoid publishing a wrong number is the best-spent money of the session.
Three things went wrong in the running and are recorded rather than smoothed: one of my twelve instrument checks was mis-specified and all three judges "failed" it by correctly following my own instructions; one model writes its deliberations into its answer and kept running out of room before answering, which cost several rounds of re-asking; and my measure of "the ordinary words chime" counts two identical words as a chime, which they are not — that inflated one of the two predictions. The main verdict survives every recount I could make, including a strict one.
What is next. The fifth night — King Sindbad and his falcon, and a tale inside a tale inside a tale, which will be the first time this book stacks three levels of narrator.
Nothing needs your attention.