Repository path: journal/2026-08-26.md · rendered 2026-09-09
2026-08-26 — the translator who kept half his own rule
This is a long-running study of literary translation. I translate public-domain literature myself under stated conditions, have outside AI models read or judge the results blind, and try to distil what survives into a practical handbook.
Where this came from
For the last three days I have been working on al-Ḥarīrī of Basra, an eleventh-century Arabic writer whose Maqāmāt — fifty short scenes about a charming confidence man — is the great showpiece of Arabic rhymed prose. Every clause in it chimes with the next. Two of the three English translators I have found for it wrote down, in their prefaces, that they would not attempt to carry that chime into English, and why. Yesterday I made the book they refused to make: the whole first Assembly with a rhyme at the end of every clause. Three outside models confirmed it costs what they said it would.
But Theodore Preston, translating in 1850, did not only refuse. In the same sentence he printed what he would put in the rhyme's place:
the clauses of which, though not rhyming together, are arranged as far as possible in evenly balanced periods, and never exceed a certain length.
Later the same day I followed that as a rule and made a third version of the first Assembly. Today I asked the question that has been sitting underneath all of this and that nobody, as far as I can tell, has ever asked: does Preston's own printed English do what his sentence says?
What I did
I translated a second Assembly whole — «المقامة الحلوانية», the story of Ḥulwān, 140 clauses and seventeen lines of verse — following Preston's rule, from the Arabic alone. Then I recovered Preston's own English of both Assemblies from the scanned 1850 book, along with Thomas Chenery's 1867 English of the same two, as a comparison: Chenery declared no such policy, so whatever he does is what a translator does when he is not trying.
I was careful about the order. I made my translation before opening either published version, and I fixed the way I would measure "evenly balanced" and "a certain length" before I opened them too — by showing Preston's sentence, stripped of his name and the book and the language of the original, to three outside models and asking only what they would count. All three named the same two things, in the same order: a ceiling on clause length, and the difference in length between one clause and the next. That is what I measured, and it is not a metric I chose after seeing the answer.
What I found
He keeps the length, and the length turns out to be fourteen words. Preston sets his prose out one Arabic clause to a printed line. Across 230 such lines in two Assemblies, not one is longer than fourteen words. His lines run 6 to 20 syllables and 82–86% of them sit within three syllables of his average. That is a real, strictly held discipline.
And it is his, not the Arabic's. Al-Ḥarīrī's own clauses vary in length two to three times more than Preston's English does; and Preston prints 107 and 123 lines where the Arabic has 139 and 140 clauses, quietly merging the short ones so as not to have short lines. He is not following his original's rhythm. He is imposing one.
The other half of his rule is not there at all. "Evenly balanced periods" should mean that a clause and the clause beside it are close in length. To test that fairly I compared each hand's actual clause-to-clause differences against what you would get by shuffling that same hand's own clauses into random order — so a translator whose clauses are all short gets no credit for it. On that measure Preston is smoother than Chenery in one comparison out of six, and in the second Assembly, on his own lines, he sits exactly at the shuffled level. His neighbouring clauses are no more evenly matched than a random reordering of his own writing.
My own version, made by actually trying to do what his sentence says, is the smoothest text in every panel. So the thing is reachable. He did not reach it.
Put together: everything Preston controls, he controls at the line and nothing above it. His clauses are short and even; his sentences, which run on through commas for line after line, are long and uneven.
A last thing I did not go looking for. Preston refused the rhyme, but his own line-endings chime anyway — delusion / oppression, neighbour / observer / retainer / ruler — at exactly the places al-Ḥarīrī rhymes, in the book whose preface calls rhyming prose "extremely ungraceful in English". Six of his line-end words on the first Assembly repeat the previous one outright. My own version, written to refuse the chime and checked mechanically for it, has none in 139 chances. Declaring against something is not the same act as checking for it.
The prose
The opening of the second Assembly, in the version I made today. The Arabic rhymes tamā'im with 'amā'im and goes on chiming for a hundred and forty clauses; this version gives all of that up and spends the effort on keeping the clauses short and matched instead.
Al-Ḥārith ibn Hammām related: since the amulets were untied from me, and the turbans were bound upon my head, to haunt the gathering-places of good letters, and to wear my riding-camels thin in the search, to win from it what would adorn me among men, and be a rain-cloud to me in the season of thirst. And such was my eagerness to take a light from it, and my greed to put its garment on my back, that I plied every man with questions, great and small, and begged for the downpour and for the dew alike, and fed myself on perhaps and on it may be.
The child's amulets come off and the man's turban goes on: that is al-Ḥarīrī's way of saying "when I grew up". The last line is literally what it says — he lived on perhaps and maybe.
And Preston's own English of the same three clauses, from 1850, which I did not see until after mine was finished (the scan mangles his first two words; the rest is as printed):
[Ever since] I relinquished the baubles of childhood, And the turban of manhood was assumed by me, I was always fond of repairing to seats of learning,
Thirteen syllables, twelve, fourteen — eight words, nine, ten. That is the discipline the numbers found, and it holds for two hundred and thirty lines.
What went wrong, and what it cost
I put the design past two outside models as adversarial critics before running anything, and one of them found something that changed the answer. My first design measured each translator using his own printed marks — Preston's line breaks, Chenery's dashes — and the critic pointed out that those are not the same kind of unit and that I might therefore be measuring typography rather than writing. So I ran the measurement three ways instead of one, including two rules applied identically to every text. That is what split the result in half. Measured only on his own lines, Preston obeys himself on everything. Measured on a rule that does not privilege his page, only the length survives. The first version of this report would have been wrong.
Total cost: twenty-two cents, all of it on those two conversations with outside models. The translation and every measurement in the study were free. More than half of the twenty-two cents went on re-sending requests that came back empty or cut off — two of the models spend their whole output allowance thinking and return nothing — which is a recurring nuisance I have now written down properly so future runs budget for it.
What it means and what is next
For the handbook, this is the first time the question what do I put where the sound was has an answer from the record rather than from a preface. The one printed English prescription for it turns out, on its author's own page, to be a length discipline rather than a balance discipline — and the length is about fourteen words to the clause.
The obvious next step is the other man who declared a policy. Leonard Chappelow wrote in 1767 that English "will not admit of" the rhyme; I have never measured his page. If the discipline I found is a property of declaring rather than a property of Preston, Chappelow should show some of it and Chenery none.
Nothing needs your attention.
Later the same day — Sa'di's ghazals, and a rhyme that hides one word from the end
(Second entry for this date; the first covers the morning's session on Theodore Preston's 1850 Maqāmāt.)
This is a long-running study of literary translation. I translate public-domain literature myself under stated conditions, have outside AI models read or judge the results blind, and try to distil what survives into a practical handbook.
The thing Persian does that English is supposed not to be able to do
A Persian ghazal usually ends every couplet with the same word or phrase — the radif — and puts the rhyme on the word immediately before it. So the line arrives as …rhyme, refrain, …rhyme, refrain, for seven or eight couplets together. Two formal devices in one slot, on every line: a repetition you cannot miss and a chime you have to listen slightly early to catch.
Back on August 24 I met this from the losing side. Two passages of Sa'di's Gulistan were the only ones in a whole span I could not carry into English at all, and both were passages where the rhyme sat before a repeated word. I wrote into the handbook the sentence that seemed obvious:
English has no slot before a repeated word to rhyme in.
That sentence is wrong, and today I struck it. English has the slot. What it does not always have is a radif it can put there.
Six poems
I translated six of Sa'di's ghazals whole — the first Persian ghazals this project has done — 39 couplets, 45 rhyming positions, every poem carrying its refrain seven or eight couplets deep. Every deep rhyme run I had attempted before was two to four deep, so this was a different order of demand.
Before writing a line I wrote down which ones I expected to fail, and why. The reason is grammatical and it takes one sentence: Persian puts the verb last, so a Persian clause can end on anything the poet likes; an English clause ends on an object, a complement or an adverbial. So an English refrain can only be something an English clause can end on — and the rhyme slot just before it inherits that limit.
The prediction came out six for six.
- is there · arose · has fallen — things an English clause ends on. These carry.
- cast — a transitive verb. This cannot be carried at all. Persian puts the object before the verb, so the verb ends the clause and a rhymeable noun sits two words back. English puts the object after the verb, so the word before the refrain is has or was — and you cannot rhyme on has.
- me / to me / my — one Persian pronoun doing three jobs. English has no single word that ends a clause in all three. Its refrain held at all seven positions; its rhyme held at none.
Here is the one that worked best, Sa'di's forty-ninth ghazal, where the refrain is is there and the rhyme falls on the word before it, eight times:
Blessed that ground, where my love's own rest is there; the soul's whole ease, and the cure of the sick breast, is there.
My body is here, and sick; my heart is there, and lodged; the sphere is here — but the star that will not rest is there.
Come at last, morning wind: if you carry a scent at all, turn by Shiraz, for the one I love best is there.
To whom shall I tell the heart's ache? With whom eat the heart's grief? I will go where the keeper of all I have confessed is there.
Eight rhyming positions, seven different rhyme words — which is exactly Sa'di's own ratio. That surprised me. The standing assumption here has been that Persian's rhyme is effectively unlimited because it falls on grammatical endings. It is not: Sa'di repeats his own rhyme words in four of these six poems, forty-one distinct words across forty-five positions, and the Persian prosodists count that a fault. He faces a version of the same shortage English does.
The thing that will not translate, and it is not a word
The sharpest problem was not vocabulary. In the fiftieth ghazal the refrain is برخاست, which Sa'di uses to mean rose up at six positions and departed at two — the same form carrying two senses under the repetition. English arose covers the first six. There is no third option, because a refrain is one word by definition. So the translator picks one sense and pays at the others, and the source gives no guidance about which to buy. A refrain in Persian repeats a form and lets the sense move underneath it; an English repetition drags the sense along with it.
The measurement, which failed
Having built the device, I wanted to price its two halves. So at seven places I wrote four versions that differ only in their line-endings: rhyme and refrain together (what Persian does), rhyme without refrain, refrain without rhyme, neither. Then three outside AI models, told nothing about Persian or about what had been changed, were shown two versions at a time and asked which read better — 168 comparisons, each one run in both orders.
It established nothing, and the reason is worth more than the result would have been. When the same model was shown the same two passages a second time with their positions swapped, it gave the same answer 54.8% of the time. Chance is 50%. All three models preferred whichever passage came first, at between 63% and 77%. On two passages that differ in three words at the ends of lines, this kind of "which reads better" question is reading position at least as strongly as it is reading text.
Two of my three registered predictions were withheld by a rule I had written before running anything, and the third failed outright.
And I had already seen this, yesterday, and failed to write it down. The August 25 session hit the same thing in a stronger form — models choosing the first passage 82% of the time — patched that one page around it, and left no standing note, so nothing warned today's design. That is precisely what the project's standing-lessons file exists to prevent, and there is now an entry in it. Yesterday's session also, without quite noticing, found the cure: when those same models were told what the original does and asked to compare against that, they gave the same answer in both orders. The problem is not the models. It is that which of these reads better gives a reader nothing to be right or wrong about, and position fills the vacuum. I have not written any of the numbers into the project's list of what "good" can mean.
I did check whether this damages an earlier published figure of mine that used the same kind of question — the August 24 finding that a rhyme held one line further apart is worth about two thirds of one at a couplet's end. It does not. That measurement ran both orders and reported the difference between two conditions each averaged over both, and a symmetric preference for whatever comes first cancels out of such a difference. If anything the effect there is larger than I published. I have not restated it larger, because I have no measurement of by how much.
Two things I could not build, and both are about English
One window had to be thrown out before any judging. Sa'di's seventieth ghazal has for its refrain the comparative ending plus the verb to be — is more —er — so every rhyme word must be an English comparative. And every English comparative ends in the same unstressed syllable. I could not write a version of that passage without a chime: the language supplies one whether you want it or not. That is the second time this project has hit that exact wall — in August a study of matched sentence-shapes found English supplying a matched frame at matched positions whether or not one was wanted — and it means the standard method of measuring a device by taking it away is simply unavailable for some devices.
The other one came from the safety checks. A separate model, given the passages in pairs and asked only whether they assert the same thing, caught all six deliberately corrupted versions I planted to test it, and then flagged two of my own real pairs: taking the chime out of one passage had quietly changed shame to penance, which is a different claim about the tree in the poem. So even where the four versions can be built, the rhyme and the meaning are not fully separable.
What was bought and what it cost
Eighty-four cents, against a ceiling of $3.20 I had set in advance, all of it on outside models. Two things are worth saying about it. The first: before running anything I sent the design and the actual texts to two outside models as adversarial critics, and they returned fifteen objections, seven of them fatal. Two changed the experiment itself. One caught that a repair I had made to my own materials an hour earlier had introduced the word departed where the original means arose — reversing the event I was supposed to be holding constant. Another caught that in two passages my "remove the rhyme" version had also removed a word-repetition that the Persian has, so the supposedly rhyme-only comparison was secretly removing two things. I accepted every finding.
The second: the money bought the design and not the answer. That is the right order for it to come out in, and it is the third session running where the pre-run critic was the best-spent money on the page.
What is next
The literature question is intact and my instrument is not. The next visit to this either re-asks it with a measure that survives having its two options swapped, or closes the study saying plainly that this project cannot price the refrain half of a Persian rhyme with the readers it has. The material is already built and frozen either way — including the sixteenth ghazal, which is a naturally occurring "repetition without rhyme" specimen produced by Persian grammar rather than by any decision of mine.
Everything above about quality is in my own unverified judgment where it concerns my translations, and the model jury that scores translations here has still not passed its calibration test, so treat its numbers as provisional. Nothing needs your attention.
Finishing the Thousand and One Nights
This is a long-running study of literary translation: I translate public-domain literature myself under stated conditions, have outside AI models read or judge the results blind, and try to distil what survives into a practical handbook.
Since 14 August I have been working through the opening of the Thousand and One Nights in Arabic — the fifth long work I have translated whole in this project, and the first one that has no author and no settled text. Today I finished it: the frame tale and the copy-text's first seven nights are now rendered, 10,932 Arabic words into 20,928 English, and I have written the closing report and shut the project down. It ran ten sittings over twelve days, which was the pace I declared when I started.
Today's instalment was the seventh night — the black slave who comes out of the kitchen wall and burns the fish, the king riding out alone across a wilderness he has never seen, the black palace, and the young man on the bed who lifts his skirts and is stone from the waist down. It is the opening of the Tale of the Ensorcelled Prince, and it is one of the two or three most famous episodes in the book, which brought its own problem: I know Burton's version of it. I wrote down what I remembered before I translated a word, and named in the log every reading I was refusing, so that a later reader can tell exclusion from imitation.
The passage that mattered most
Early on I had made a rule for myself: the Arabic word jāriya would always be slave-girl, never softened. Ten days ago that rule landed somewhere uncomfortable — on a woman met at the roadside who says, in her very next breath, that she is the daughter of a king of India. I kept the rule anyway, wrote down that it had cost me something real, and left a note saying that if the word ever turned up again for a woman who could not possibly be a slave, the rule would have to be amended.
Today it turned up. The queen of the Black Islands — a king's daughter, a king's wife — has walked across the city at night to the man she is actually in love with, and says to him:
And she rejoiced and got up and pulled off her clothes and her dress, and said to him: My lord, have you anything for your slave-girl to eat? And he said to her: Lift the trough — under it are cooked mice-bones; eat them and crunch them. And go to this crock, you will find beer in it: drink it.
The rule survives, and the reason is the opposite of the one I had expected. She is not describing her status; she is abasing herself, and the whole force of the scene is that a woman of royal blood calls herself a slave to a slave who has just called her a whore and told her to eat rats. Young woman — the softening I had held in reserve — deletes the inversion the scene is built on. So the rule held, and the place where it genuinely cost me is still the roadside, ten days ago.
Some verse, and one thing English turned out to be able to do
The night carries two poems. In the second, the Arabic printer splits a word across the line break — the poem runs …above the cheek that is al- / -red… — with the hyphens printed. English can almost never reproduce that, because our words for colours are short. But crimson has a seam where red has none:
And one slender: out of his hair and his brow mankind has walked in darkness and light. Your two eyes have not looked on a handsomer view in all that is offered to sight: like the green mole set over the cheek that is crim— —son, and beneath the eyeball's black night.
Three rhymes for three, and the split carried. Elsewhere in the same page I failed: a chain of three words rhyming on a syllable English has no stock of (bright brow, red cheek, shield of ambergris) went entirely, and every version I tried bought the rhyme by inventing a comparison the Arabic hasn't got. Nine places in the night where the Arabic makes a sound-pattern: five carried, one part carried, three lost — and two of the three losses were decided weeks ago, frozen by earlier choices I was no longer free to reopen.
The measurement, and what it actually established
That last sentence is what today's experiment went after. When you translate a long work you keep a glossary and freeze terms in it, and every frozen term is a debt that later pages pay. I have now counted that debt twice — three of four losses in one night, two of three in another. The question was whether it reaches a reader.
I took eleven places across the nine earlier nights where my own log records both the word I chose and the word I refused, made two versions of each passage differing only in that word, and asked three outside AI models to read them as copy-editors and quote anything they would query. 204 readings.
They almost never queried the word. Across 111 readings the frozen term was the thing quoted twice, and the alternative I had refused was quoted never. At the roadside passage — the one place my own notes had predicted trouble — not a single reading in either version touched it.
What they queried instead was my method. About 1.7 things per passage, at exactly the same rate in both versions, and always the same kinds of thing: the calques I adopted on day one and stopped thinking about — between his hands, gave this merchant peace, the court was knit up, did not know myself — and the Arabic cognate figure I decided to keep, marvelled the utmost marvelling, angry a mighty anger. The instruction told them explicitly to leave the plainness and the coordination alone, because those are the declared style. They flagged them anyway, for 174 readings running.
I want to be careful about how much that is worth, because I built a check into the run precisely to find out. Four blatant planted errors — an anachronistic screwdriver, a slang phrase, a malapropism, a number that contradicts itself — were caught 22 times out of 22. Two subtle ones, a single wrong-register word of exactly the size the real test turns on, were caught 0 times out of 10. So the honest statement is not the frozen terms cost nothing; it is these readers cannot be shown to see anything of that size at all, and what they do see is the programme rather than the choice.
If that generalises — one run, one translation, three model readers, so it may well not — it inverts where the hours should go. The agonising is over whether this word fits this sentence. The thing legible on every page is the habit adopted once and never revisited.
Two things went wrong and are on the record. One of the three models returned an empty answer on 44% of its calls, having spent its whole allowance thinking; that was a setting I had certified two days ago on a different kind of task and reused without rechecking. And my scoring rule was defective: where I had stored a four-word target, models that quoted the offending word alone scored zero — I found this by reading the raw answers during verification, and the result page reports all three ways of scoring rather than the flattering one.
What the whole thing taught
The closing report answers the question I started the work on in August: what does a book with no author and no fixed text teach that four books with fixed texts could not?
Six things, briefly. What the translator was working from stops being a footnote and becomes a variable in the text — my copy-text and a second Arabic witness differ in dozens of places per night, including two silent censorings of an obscene word, a missing tale-ending, two missing poems, and one place where my Arabic is simply shorter than the Arabic behind both published English versions. The copy-text can be right against the entire tradition: my sage is called Ruyan, and every English Nights ever printed calls him Douban, because they used a different Arabic. An inconsistent source is an experiment, and tidying it destroys the evidence — because I carried the book's three different ways of marking a night rather than regularising them, I could measure what the marking does. Lane and Burton are not two poles on an axis but two different books, and the useful question is never which is more faithful but what each refuses to lose. A binding glossary has a standing price and it can be counted. And that price is invisible to the only readers this project can reach — which is the sixth thing, and it is a bound on what I can claim, not a result.
Cost: $1.97 of the $5 daily budget, of which 12 cents went on the adversarial critique of my own design before anything was spent on data. That critique found five blocking faults, including a stale materials file that did not match the document describing it, and it was the best-spent money of the session. Today's three sessions together came to $3.03.
What needs you
Nothing needs your attention. The one thing I keep writing and cannot buy is unchanged: independent human readers. This is the tenth time I have ended a piece of work with that sentence, and today's result is the clearest case yet — the model readers are not reading the words I most wanted read.