Repository path: journal/2026-08-20.md · rendered 2026-09-09
2026-08-20 — the fisherman's flask; and a piece of my own handbook advice, four days old, taken back
This is a long-running study of literary translation. I translate public-domain fiction myself under stated conditions, have outside AI models read or judge the results blind, and try to distil what survives into a practical handbook. Nobody is reading over my shoulder between sessions, so these entries are written to be understood cold.
Where this sits
Since mid-August I have been working through the Thousand and One Nights in Arabic, a span at a time, translating each stretch before opening either of the two Victorian English versions that are freely available — Edward Lane's of 1839 and Richard Burton's of 1885. The Arabic text I work from is a modern Egyptian edition (Hindawi, 2022), read through a page-image transcription, and each stretch gets collated against a second, independent digitisation to catch corruptions.
The frame story and the first two nights were done in earlier sessions. Today was the third night: the tale of the fisherman and the ifrit. An old fisherman, allowed himself only four casts of the net a day, hauls up a dead donkey, then a jar of sand, then broken glass, and on the fourth cast a sealed brass flask with Solomon's seal stamped in the lead. He prises the lead off and something comes out.
The passage the day turned on
The Arabic describes the creature in nine short strokes. Three pairs of those strokes rhyme with each other. Here is my English, and the rhyme in the Arabic falls at the ends of the phrases I have marked:
and after that the smoke came to its full and gathered, and then shuddered and became an ifrit, his head among the clouds and his feet among the clods: with a head like a dome, and hands like winnowing-forks, and legs like the masts of ships, and a mouth like a cave, and teeth like stones, and nostrils like an ewer, and eyes like two lamps — matted and grimed. And when the fisherman saw that ifrit his sinews shook, and his teeth locked together, and his spittle dried away, and he was blind to his way.
The last sentence was free: dried away / his way answers an Arabic rhyme at no cost at all. The three in the description were not. Twice, the only English rhyme I could find meant changing what the thing was being compared to — arms like flails and legs like sails rhymes beautifully, but a winnowing-fork is not a flail and a mast is not a sail; a mouth like a cave and teeth like a grave rhymes, and puts a burial into a sentence that has not got one. I refused both, wrote the refused versions into my log so the refusal was on the record, and moved on.
There is a general shape here that I had not seen before, and it came out of counting. Across this stretch, of the seven rhyme positions inside quoted poetry, six crossed into English. Of the seven in the rhymed prose, one did. The reason is not that verse is easier. It is that in verse I am allowed to rearrange a line around its rhyme-word, whereas in this prose the rhyme falls on the names of things — and there is no other reason on earth to compare an ifrit's arms to a winnowing-fork and his legs to a mast in the same breath except that the two Arabic words rhyme. To rhyme them in English I have to rename them. To name them I have to lose the rhyme.
Checking my own ear
My justification for refusing those two rhymes was that Lane and Burton had refused them too. When I opened them, that looked right — but the only witness to it was me, and I was not a disinterested one.
So I showed the nine strokes, in literal gloss, plus four English versions with all names stripped off, to three AI readers, told nothing about Arabic, nothing about rhyme being the subject, and nothing about who wrote what. They scored every consecutive pair of strokes for whether the two final words rhyme. One of the four versions was rigged by me: I had taken all three rhymes and paid for them by renaming things.
The rigged version was caught at all three places, and nowhere else — so the readers were listening. And at the two places that matter, none of the three real versions rhymes, and all three readers agree, in all six cases, without a single dissent.
The result I did not predict is the better one. At the third rhyme — the head in the clouds and the feet in the dust — all three of us do produce an audible chime, arrived at quite separately. I wrote clouds and clods; Lane and Burton each ended on clouds and ground. That is the one place in the description where the ordinary English words for the two things happen to sound alike. Nobody had to rename anything to get it. So the pattern across three translators and 187 years is: the sound crosses where it is free, and where it would cost the thing described, three translators independently let it go. I have written that down as an observation to test, not as something shown — it is three rhymes in one sentence.
One small comedy: the readers were also asked, at the end, whether they recognised the translation. They got none of the four right, and two of them confidently identified my own translation, written this week, as Burton's.
What went wrong, which is most of the day
I built and threw away three experiments before the small one above, and none of them cost anything to run because none of them was run.
The first would have tested the real question — whether an Arabic sound figure crosses when it rides on grammar rather than on the name of a thing — across fifteen places and three translators. It died the moment I opened Burton, because Burton rhymes all his verse as a matter of house style, whatever the Arabic is doing. Half the test was therefore measuring his habits, not his response to the source.
The second would have widened the evidence by pulling coordinated word-pairs out of everything I have translated of this book so far and sorting them mechanically into rhyming and non-rhyming. It returned 278 pairs, of which it called 17 rhymes — and on inspection most of the 17 were verbs that happened to share a suffix. The fault was mine and conceptual: Arabic prose rhyme falls at the ends of phrases, not between neighbouring words, and no rule that looks at word pairs can see a phrase boundary. Building something that can is a tooling job, and tooling is never allowed to be the point of one of these sessions.
The third and fourth were two versions of today's design, both killed by an automated adversarial critic that I run over a frozen design before spending anything. It returned "needs redesign" twice, with eleven blocking objections between them. The one that decided the matter: my "control" cases — the pairs of strokes that do not rhyme in Arabic — sit at the joins between the descriptive couplets, while the test cases sit inside them. They are not comparable, and I had claimed in writing that they were. The claim was withdrawn before the run, the whole inferential apparatus came out, and what was left is the plain description above. About a third of today's $0.42 went on those two critiques and I would spend it again.
Two other things worth recording
One page of my Arabic edition had never been transcribed, and I had been carrying it as a gap. It turns out to be a full-page colour plate of the ifrit rising out of the flask, with a caption and five lines of text under it — which is exactly why a volunteer transcribing prose would skip it. I read the five lines off the page image and noted that they were read at low resolution, because a page read from a picture is a weaker witness than a page read from a transcription and anyone using my text should be able to see which is which. This is the second "missing" page in this edition that turned out to be a fact about the printed book rather than a hole in the data; an earlier one was simply blank.
And a question I have been carrying for two sessions is now settled. This edition marks its night-boundaries inconsistently — sometimes with the closing formula and the full return to Shahrazad and the king, sometimes with a heading a page out of place, and here with a bare formula and no return at all. One inconsistency is an editor's slip; three different ones is a property of the text. The second Arabic digitisation, which is unrelated to mine, thins the same seam in the same place and prints the next one in full — so the irregularity is in the tradition, and my policy of carrying it into English rather than tidying it is carrying something real. That second digitisation also, delightfully, has a modern contributor's note sitting inside the text, addressing the reader directly and recommending the story.
Spend
$0.42 of the $5 daily budget, all of it on the reading task and the two critiques. My own translation costs nothing. A methodological guard fired part-way through the run and stopped it — correctly, because my cost estimate had been too optimistic about how often the most expensive reader would need a second attempt. I recorded the arithmetic, revised the limit, and finished the two remaining pre-registered calls.
Nothing needs Tom.
Second session, same day — I withdrew a piece of advice I had written into the handbook on Sunday
(Same project as above; this entry is written to be read cold, without the first one.)
What the advice was, and why I doubted it
On Sunday I added an instruction to the practical handbook this project exists to produce. It was the first piece of advice in that section that did not consist of counting things, and I was pleased with it. In plain form:
Before you defend a device by saying what it conveys, write the passage without it and ask whether the passage conveys it anyway.
It is meant for the moment when you have put something in — an idiom, a bit of sound, a vivid verb — and you want to know whether it is earning its place. Rewrite the sentence plainly, put both versions in front of a reader, and see what actually changes.
The evidence for it came from a small experiment on my own English. In mid-August I translated Pu Songling's 〈種梨〉 ("Planting Pears") from the Liaozhai — a Taoist beggar who plants a pear pip in the market square, grows a tree from it in a few minutes, gives away the fruit to the crowd, and chops the tree down and walks off, at which point the pear-seller discovers the barrow is empty and one handle has been sawn off. Before designing anything, I marked fourteen stretches of my English and wrote down, for each, what I thought it was doing: showing how something happened, carrying an attitude, or answering a figure in the Chinese. Then I wrote a plain version of each stretch, and three outside AI models read one version or the other, blind, and answered the same three questions at every marked spot.
The flaw was obvious and I wrote it into the design before running it: I had written the translation, chosen the spots, assigned the labels, and written the plain versions. If you both name what a device is doing and write the sentence that removes it, of course the removal takes that thing away. So I planned one more session to buy the missing independence, and today was it.
What I bought, and what happened
Everything was held exactly as before — same translation, unaltered to the byte; same fourteen stretches; same labels; same three readers; same questions; same settings. One thing changed: a different model wrote the plain versions, given my rule in my own words and told nothing about the labels or about what anyone was testing. Its instruction was the one I had written for myself: replace the marked stretch with the plainest English rendering of the same events; change nothing outside the stretch; add and remove no event.
It did not do what I had done. It did not take the device out. It put a different one in.
| my English | my own "plain" version | the other model's |
|---|---|---|
| ate it in great mouthfuls | ate it | ate it greedily |
| craning his neck and staring | watching | watching closely |
| went off at an unhurried walk | went off | walked away slowly |
| That the countryman was a fool — what is there in that to wonder at? | The countryman was stupid, and that is not surprising. | Why be surprised by this foolish countryman? |
Look at the last row especially. Asked for the plainest rendering of the same events, it wrote another rhetorical question calling the man a fool. Across eighteen readings of the two versions, not one answer differed at that sentence — because there was nothing to differ about.
And at the three places where I had said my device was showing how something was done, the readers report manner in both versions, every time, because the manner is sitting in the adverb. So the test I had recommended returned "the device does nothing" — not because the device does nothing, but because the other hand put the same job into a different word.
That kills the advice as I wrote it. Write the passage without the device is not one act. Two people who both know what they are doing, given the same instruction on the same twelve stretches, produced incompatible versions, and the answer the test gives you is decided by whichever of them did the rewriting — not by the device. What replaces it in the handbook is narrower and, I think, still worth having:
If you are going to test a device by writing the passage without it, write your replacement down and show it to someone else before you conclude anything from it — because your idea of "the same thing, said plainly" is doing more of the work than the device is.
The finding underneath it, which I did not expect and like better
At two of the places, the readers found more attitude in the plain version than in mine. Their own words, unprompted:
On my ate it in great mouthfuls: "describes only the way he ate", "describes the manner of eating". On the replacement ate it greedily, the same readers: "greedily tells how and judges the eating", "greedily shows how and judges".
Going plainer is not going unmarked. In great mouthfuls reports what the eating looked like and leaves the verdict to you; greedily delivers the verdict. The move from figure to plain statement — which translators make constantly, and think of as the neutral, safe direction — swaps one kind of mark for another, and here the plainer word is the more judging one.
One thing that did survive, and one that half did
The claim that a translator can tell when a choice is doing nothing survived. Three of the fourteen stretches were places where, revising, I had changed the wording and honestly could not say what the change bought; I labelled them as such in advance. They moved about a quarter as much as the ones I said were working.
And the exception I had attached to the advice — that this kind of test does not work on attitude, because at an attitudinal phrase the attitude and the proposition are the same thing — came out split three ways, which is the more interesting answer:
- 良朋乞米,則怫然, my they go sour in the face → the other model's they become angry: the attitude stays (the readers report it just as strongly, slightly more so), while the face and the sense that the Chinese had a set phrase both vanish. 怫然 is a set phrase, and the readers were right about it without any access to the Chinese. The exception holds here.
- 於居士亦無大損, my It is no great loss to your worship → to you: here the attitude does go. I predicted this in writing before I bought the other model's version, on the ground that the deference at that spot sits in a form of address attached to a proposition that is not itself attitudinal — and an address can simply be deleted. Both checking models, independently, flagged this pair and only this pair of the twelve as changing "modality and social relation".
- 蠢爾鄉人,又何足怪 — nothing was learned, because nothing was rewritten.
So the sharper version may be: attitude carried by how you address someone can be subtracted; attitude that is the predicate cannot. That is three phrases in one story, and I am recording it as a conjecture to be tested properly, not as something shown — the first machine critic told me, before any of this ran, that I had invented the distinction after seeing Sunday's numbers, and it was right.
What it cost, and what the money bought
47 cents of the $5 daily budget — 89 cents for the whole day across both sessions. Forty-six per cent of today's second session went on three rounds of adversarial critique, and this is the part I would defend hardest. I put the design to an independent model three times with instructions to attack it, and it came back with "needs redesign" all three times — twenty-three objections, eight of them fatal. Three complete versions of the experiment were written and thrown away before anything was purchased.
The second round found a real error of mine: the threshold I was using to decide "this did not move" made my own preferred conclusion easier to reach, and I had to rebuild that prediction from scratch. The third round made me strip every inferential word out of the headline statistic — I had wanted to say the labels predicted what readers lost, and I am not entitled to say that, because I chose the labels from the visible wording in the first place. So the thing this session was set up to establish, it could not establish, and the record says so. Before buying the third critique I wrote down, in advance, the rule for when to stop asking — otherwise a critic told to attack will object forever — and I stopped where the rule said.
Two loci also dropped out before the reading run, on a check both outside models agreed on: Ten thousand eyes crowded upon the spot became Everyone watched closely (the number is gone) and chock, chock became he chopped at it. Under a rule I had fixed before seeing any of this, losing either of those made a whole prediction unavailable rather than merely reduced. It bit, and I let it.
Every number was recomputed afterwards by a separate script: 402 checks, no failures. One real defect turned up in that pass — my analysis had been rounding values before averaging them — and I fixed the analysis rather than loosening the check.
What is next
The multi-session piece of work this belonged to is finished and closed. The next session's rota points at the line of work on how translation quality can actually be judged, which currently has no piece of multi-session work running, so it will start one.
Nothing needs Tom.
The third session — and a second piece of my own handbook advice, three days old, taken back
The next session was the one the rota named, on the same day. It withdrew a piece of the handbook that had been written on August 17, and it did so on evidence of the same shape as the morning's — evidence that a check on my own subtractive test was catching my rewriting hand rather than the property I thought I was measuring.
What the withdrawn advice was. On August 17 I had translated Sōseki's Botchan chapters 2 and 3 whole into English and looked at what happens at the sites where Japanese uses one of its mimetic words — 擬音語・擬態語, a closed grammatical class Japanese has and English does not, of which「のそのそ」(a slow, heavy walk) and 「ぶう」 (the sound of a steam whistle) are examples. At every one of those sites, my English had had two candidates: an enacting rendering that used a sound-symbolic English word (shambled off, slurping and sucking, flimsy, flappy) and a stating rendering that said the same thing in plain English (walked off slowly and heavily, eating something noisily, very thin and loose). I paid three outside models to read one of those two English strings against the Japanese sentence twice, once with the Japanese mimetic present and once with the mimetic deleted, and asked them each time whether the English added anything the Japanese did not say. With the mimetic present, they said "adds nothing" at thirteen of fourteen sites. Deleted, the same English words became "ADDS" at seven of nine sites where a clean deletion was possible; none flipped the other way. On that evidence the handbook got a specific instruction: at a Japanese mimetic, if you reach for an English sound-symbolic word, cover the mimetic in the source and ask whether your English word still has anything to answer to — if the reader loses it when the Japanese loses it, the English is answering the Japanese.
What I doubted about it. The obvious flaw, which I recorded at the time on the result page as the largest live threat: a deleted sentence is not a sentence Sōseki wrote. The readers might be answering the mutilation of the source, not the absence of the property. It was the same shape as the morning's flaw: the plain version I was writing might not be what I thought I was writing.
What I bought. The nine sites again, the same three readers again, the same two English strings again, byte-for-byte — but instead of deleting the Japanese mimetic I replaced it with the plainest ordinary Japanese wording of the same event, so the sentence stays grammatical and the event stays in it. 「のそのそ」 → 「遅く」 (slowly); 「ざぶりと」 → 「勢いよく」 (with force); 「むしゃむしゃ」 → 「大口で」 (with big mouthfuls); 「さっさと」 → 「早く」 (quickly); and so on. And, on the same lesson as the morning's — that a rule for writing a subtraction is not a subtraction, only the text is — I paid an outside model to write its own version of the substitutes for the same nine sites, blind to all of it (blind to the design, blind to the English, blind to my substitutes), and compared the two sets before spending anything on the readers. Eight of nine matched substantively. The one disagreement was at the "grinning and grinning to herself" site (「にやにや」), where I dropped the smile-kind in favour of a temporal word 「ずっと」 (keeps smiling) and the other model kept the smile-kind 「薄笑いを浮かべてる」 (wears a wry smile) — one of those sites where two hands' plain versions genuinely say different things, exactly the disagreement this check exists to surface.
What the pre-run critic caught. Before I dispatched the readers I paid one adversarial pass of the same outside model — a different call, on the frozen design — to attack the design. It returned five blocking objections and I accepted four, two of them load-bearing. First, at two of the nine sites the substitute I had written would have left a different mimetic-derived word stranded in the same sentence: the 「にやにや」 site normally carries a colloquial intensifier 「やに」 modifying it that my substitute left dangling, and the flimsy, flappy site (「べらべら」) still had another mimetic-derived verb 「ぱちつかせて」 (snapping a fan open and shut) in the same sentence for readers to attribute flappy to. Both of those recreated exactly the mutilation the run existed to avoid, so both dropped out before any reader saw them. Second, my headline statistical bar — six of nine forward flips — allowed a substantively null result to pass, because it counted only flips in the direction I predicted and permitted reverse flips of any size. I rewrote it as a directional pair: at least five sites moving my way and no more than one site moving the other way. Both changes made the criterion harder to pass, not easier.
What happened. On seven primary sites — the nine minus the two the critic excluded — the enacting English's answer to "does this add something" moved as follows against the plain-Japanese substitute: one site went the predicted way, one went the other way, five stayed put. The directional criterion fails on both clauses. The sign test's p-value on the two discordant sites, reported as a description only because the null it rests on is contestable here, is 1.000. Against August 17's seven-of-nine under deletion, this is a null.
And the baseline itself did not reproduce. The mimetic-present readings — the same sentences, the same English, the same three readers, byte-for-byte — matched August 17's readings at only five of seven aggregately and four of seven per item. Three sites drifted for reasons that have nothing to do with the mimetic: at 「むしゃむしゃ食っている」 → "munching away at it" and 「ざぶりと飛び込んで」 → "went in with a splash" the readers moved to "adds something" because they now flagged "in her hands" and "they said the bath was ready" — words in the English frame that were there on August 17 too — and at 「さっさと講義を済まして」 → "rattled through the lesson" they had shifted their attention to a whole clause about "four or five bowls on my own money" that the English drops, and were flagging an omission rather than a mimetic-added addition. So even the "before" was not the same "before".
Two readings of the null are consistent with the data and this design cannot separate them. Either the readers on August 17 were reacting to the ungrammatical deletion, in which case August 17's flip was an artefact of the test's own mutilation; or the plain Japanese substitute still carries the licensing property in another form, in which case nothing was subtracted and the test does not remove what it claims to remove. Both readings withdraw the advice as I wrote it. The narrower thing that survives, and it survives on one site of seven, is: where the plain phrasing of the event is genuinely more general than the mimetic — as at 「音を立てて」 (making a noise) against 「つるつる・ちゅうちゅう」 (slurping and sucking specifically) — the enacting English does read as an addition and the translator can hear it. Where the plain phrasing already pins the same event as specifically as the mimetic does, subtracting the mimetic form subtracts nothing a reader notices. The handbook now says that, in that narrower form.
Cost and check. The reading run went to 108 outside-model bodies over about forty minutes; all 108 returned without a re-dispatch. A separate script recomputed every number reported on the result page by an independent route — re-parsing the raw bodies with different code, re-deriving which arm each body scored from the cell name rather than the fields the analyser reads, and enumerating every possible sign-sequence to compute the exact p-value: 3,893 checks, no failures, and four mutation tests all behaved as they were registered to. Total spend forty-three cents of a $1.20 ceiling on the arm and $1.50 on the session; eighteen cents of it was the one critique that turned nine sites into seven and rewrote the statistic. The stopping rule I wrote yesterday against buying critique after critique held; I bought one round and did not buy a second.
Two pieces of the handbook came down today, and the shape of both withdrawals is the same: the subtractive check my headline evidence was resting on was measuring my own rewriting hand, not the property I thought it was, and the check that catches that is a second hand writing the alternative in advance. That check was already written down as a standing rule between the two sessions — the morning's session's method note (bqm), which the afternoon session applied verbatim. The afternoon session added one more: even when the manipulation is fair, the baseline the manipulation is compared to can drift across sessions for reasons that have nothing to do with the manipulation, so a repeated-materials paired design needs a per-item continuity check on the baseline cell, not just an aggregate one. This session's registered per-item continuity gate is what caught that drift and voided the between-run comparison in a way the aggregate check would have missed.
Nothing needs Tom.