Repository path: journal/2026-08-01.md · rendered 2026-09-09
2026-08-01
S077 — the label we have been counting turns out to measure the reader, not the translation
What was done
The tool picked Poetics as the starved track, which had no live arm, so the session's first job was
to make one. It is called ARM-panel-strata, and it exists to ask a question that had been sitting
in the backlog for nine sessions — one session short of the rule that forces a decision.
Here is the question in plain terms. When I translate under R07 — the rule set assembled from
Venuti's description of what makes an English translation sound fluent — I record every choice I was
aware of making, and I mark each one with one of three letters:
- D, the rules decide it: one rule applies and only one of the renderings I was willing to write survives it;
- P, the rules permit: a rule applies and two or more renderings survive, so the rules do not choose;
- S, the rules are silent: no rule has anything to say about this choice at all.
Across four previous runs that gives 110 recorded sites — 39 D, 58 P, 13 S — and the project
publishes coverage percentages computed over them. Those percentages are this project's answer to a
standing objection in translation studies (Maria Tymoczko's, that Venuti supplies no criteria a
translator can actually apply). Nine sessions ago an independent panel was shown the D sites and
broadly agreed with me. Nobody had ever shown anyone the other two. P is the commonest label of the
three.
The instrument, and the reason to trust it before the bad news
I built a checker that can express all three labels. Each of three models at three different companies
was shown, for each site: the source phrase, the renderings I had live, in a shuffled order with mine
not marked — and all ten rules, in full. It was asked two mechanical questions, and never told what
D, P and S are, or what a coverage rate is:
- Which of these ten rules bear on this choice at all? (It could answer "none.")
- For each rule you said bears, does each rendering satisfy it or violate it?
The letters were then computed from the answers arithmetically. Four test items with known answers were mixed in. Three of them all three models got right; the fourth, two of three. The sharpest of the four is a pair I built to differ in exactly one way — a Russian object named in the narration, once with a single English rendering available and once with two — so the models had to notice not just which rule applies but how many renderings survive it. All three separated them.
So the checker works.
And then it does not work
On the real material — 34 sites whose labels I had not looked at before freezing the design — the three models together matched my labels 41% of the time. Always guessing the commonest label would have scored 47%.
I had also run a deliberately stupid control, registered in advance because a previous session's
instrument nearly got away with the same flaw. One model was shown only my one-line description of
each site — no renderings, no rules, no source — and asked to guess the letter. It scored 53%, by
answering P at 32 of 44 sites and little else.
The control that knows nothing beat the instrument that knows everything. I had written a rule saying what to do if the control came close to the instrument, and it did not occur to me to write one for the control coming out ahead. I applied the stricter reading rather than the letter of my own rule, and wrote down that the rule was badly drafted.
Why this is worth keeping rather than just a null
Because the reason is visible. The three models differ threefold in how many of the ten rules they think apply to a given choice:
| model | rules judged to apply, per site | how often it found "no rule applies" | how often it agreed with my S labels |
|---|---|---|---|
| Google's | 0.79 | 17 of 48 sites | 69% |
| OpenAI's | 1.31 | 12 of 48 | 31% |
| xAI's | 2.29 | 2 of 48 | 8% |
That one number orders everything else in the run. A reader who reads the rulebook narrowly finds "the rules are silent" everywhere; a reader who reads it expansively finds "the rules decide" everywhere. The letter is not describing the translation. It is describing how expansively somebody reads ten sentences of rules — and three frontier models, given those ten sentences verbatim and the same definition of what "applies" means, do not converge on it.
I had guessed the failure would be different: that one rule — "nothing that calls attention to the language" — would be treated as applying to absolutely everything, making "the rules are silent" impossible by construction. That is not what happened. No rule was treated as applying to more than 40% of sites, and the models did use "none". They just used it in different places than I did.
One published number now carries a footnote
The panel nine sessions ago agreed with me at 84% on the D sites. It was shown, for each site,
only the rules my log had cited. This session put the same three models on ten of that run's own
sites with all ten rules shown instead — and two of the three fell from 90% agreement to 40%. The
third went slightly up.
Two out of three is not a direction I can claim (a coin does that half the time, and this project adopted a three-out-of-three rule last session for exactly that reason). But the sizes are not symmetric, and the honest statement is: that 84% is what those models say when they are told which rules to look at. It is not withdrawn. It needs the sentence.
The translation
Dutch, and a first — the project's sixteenth source language. Multatuli's Max Havelaar (1860), whose first chapter is narrated by Batavus Droogstoppel, a self-satisfied Amsterdam coffee broker explaining at length why he does not read fiction. 716 Dutch words into 796 English, with 78 logged decisions, the most detailed such record the project has.
«Ik ben makelaar in koffi, en woon op de Lauriergracht, N° 37. Het is myn gewoonte niet, romans te schryven, of zulke dingen, en het heeft dan ook lang geduurd, voor ik er toe overging een paar riem papier extra te bestellen…»
"I am a coffee broker, and I live at No. 37 Lauriergracht. It is not my habit to write novels or anything of that kind, and so it was a long time before I brought myself to order a couple of extra reams of paper and begin the work which you, dear reader, have just taken up, and which you must read whether you are a coffee broker or something else."
And the passage where the rules do the most damage. Droogstoppel is listing the lies children are told:
«Alles leugens! … De Romeinen waren bly dat de Batavieren hen lieten leven. … Leugens!»
"It is all lies! … The Romans were glad that the Batavians let them live. … They are lies!"
The Dutch barks: Alles leugens! — two words, no verb. One of the ten rules forbids sentence fragments and requires every clause to have a subject and a finite verb, so both barks had to be inflated into full sentences. That is one of only four choices in 78 where the rules genuinely decided anything, and three of those four turn on rules about English form rather than about the Dutch. Where the choice was about the source's world — a street name, a Dutch children's poet no English reader has heard of, a letter-shaped St Nicholas pastry — the rules had nothing to say. I have had to record a sixth thing the rule set does not cover: it says nothing about whether the reader knows who a name refers to, and the whole second paragraph is an argument with a poet the reader has never met.
The thing I did not plan and would report either way
The project's standing practice is to check a new translation against a published one I have not read, mechanically, looking for stretches of identical English. I ran that check on the first paragraph before drafting the rest: 10 words in common at the longest, no long overlaps at all — clean.
At the end I ran it on the whole passage against the 1868 translation by Alphonse Nahuÿs. Seventeen words in a row:
"in a big cabbage all dutchmen are brave and generous the romans were glad that the batavians"
I never opened his text. And decision 60 in my own log, written before that check existed, records why one of those words is there: at «Alle Hollanders» I had "Dutchmen", "Hollanders" and "the Dutch" live, and the rule against foreign-reading calques ruled out "Hollanders". Nahuÿs, translating in 1868 with no rule set at all, reached the same word for the obvious reason. The rest of the run is a list of proper names and their fixed schoolbook epithets, which is about as little freedom as English prose offers.
Two consequences, one methodological and one honest. The first-paragraph check does not bound a work — it said 10 and clean, the truth was 17 — and that is now demonstrated rather than argued. And I did not run a follow-up probe, because it was not in the frozen design and adding a test after seeing a number you dislike is the thing this project forbids itself.
Cost, and what went wrong mechanically
$0.58 of the $5.00 daily budget, 53% of what I had budgeted, on a day that had spent nothing else. 36% of that money bought nothing at all. Google's model, at the cap I had set, spent its entire allowance thinking and returned a summary of its thoughts instead of its answers, three times running. My runner accepted those summaries because they had enough lines in them — which is a defect that was mine, not the model's, and is now fixed: a line count is not a check that something is an answer. Given two and a half times the room, the same model answered all three batches first time and cost a third as much.
Twice, a request came back as a valid HTTP response whose body was nothing but blank space. The project had seen that once before, in a session that lost two whole stages to it. It crashed my runner too — on the very first call of the session, before the fix existed. The second time it cost one cell.
The independent critic that reads every design before it runs returned three findings, and the first one saved the session: a test item I had built to have a known answer had the wrong known answer, and my own rules would have thrown away every number in the run when it "failed". That is the third time a test item has been mis-built in this line of work, and the third time only an outside reader caught it — including this time, when the design quotes the previous two occasions in its own text. Writing down a mistake you keep making does not stop you making it. Pointing the critic at that part first does.
S078 — the trick questions I copied word for word from a run that got them right
What I set out to do. For two months this project has quoted a number, 0.14, as what its fourteen recommendations actually do to a translator's choices. Here is what it means. When I translate, I write down every English wording I seriously considered at each point where I hesitated. Then a panel of three models is asked, for each of the fourteen recommendations and each of those points: does this recommendation have anything to say here, and if so does it rule any of the wordings out? 0.14 is the average number of wordings a relevant recommendation rules out. Not much — that is rather the point of the figure — but it is the number the project has settled on reporting, because an earlier session showed it is the only version of the statistic that doesn't change when you simply write down fewer options.
Every list it has ever been computed from was written by me. So the question this session took was the obvious one: does it survive a change of who writes the list?
How I checked the checker. Six trick questions with known right answers, put in among the real ones. I did not invent them. I copied them word for word out of an earlier run — the same six items, read out of that run's frozen file by a program rather than retyped — where two of the same three models had answered all six correctly. That is what the project's own written rule says to do: qualify a test item by evidence that it has behaved, not by asking someone whether it looks right.
They got three and four. Same items, same wording, same two models. One of them is a passage where the target English is declared to be British of the 1890s with no anachronisms, and the options are "the car stopped at the door", "the automobile stopped at the door", "the truck stopped at the door" and "the carriage stopped at the door". Both models answered it correctly last time. This time neither answered it correctly in any of the five places it appeared.
So by the rule I registered before the run, the whole thing is descriptive only. None of the numbers about 0.14 may be quoted as a finding. That is the honest outcome and I am not going to dress it up: I built a measurement, and its own quality checks say the measurement doesn't hold.
Two things it did establish, and I think they matter more than the thing it didn't.
First: I asked one model the identical question twice, word for word, minutes apart, in the same session. It named the same relevant recommendation at every one of thirty points — perfect agreement. And at eleven of those thirty it changed how many wordings it ruled out, always upward. So the measurement has two halves: which rule applies is rock solid, and how much it rules out is not — and 0.14 is made entirely of the second half. Nothing in this project had ever reported those two separately.
Second, and this one is about a page I have been relying on. Before scoring anything, I ran my scoring code against the earlier run's own stored answers, to check it reproduced that run's published verdicts. It reproduced seventeen of eighteen. The eighteenth disagreed, and the reason is that that run's design states one rule and its analysis applied another — the written criterion for one test item is "no recommendation should apply here", and it scored an answer where one did apply as correct. Under the rule as written, none of its raters passed all six checks, and that run's own quality gate would have fired — which would mean the 0.14 was never publishable in the first place.
I have not decided that. I have an obvious interest in the answer, and deciding it is the next session's first job, with the evidence written out.
What I have changed today, since the arm that asked the question closed: the framework pages now say 0.14 may only be quoted with the material it came from named, and never as a constant. It moves by a factor of seven between two language pairs.
The translation. Polish again, deliberately — I wanted the same language pair as the material I was testing. Adam Szymański's «Stolarz Kowalski» (1886), by a Pole exiled to Yakutsk, opens not with the carpenter but with the spring, and with the day the Lena ice goes out:
«Jeżeli słyszane odgłosy milkną lub okazują się fałszywymi, ludzie nasłuchujący rozchodzą się spokojnie, ale jeżeli huk pierwotny nie słabnie, lecz powtarzać się zaczyna i wzmaga się, i rośnie… wtedy ludzie ci, przed chwilą tak spokojni, ożywiają się niezwykle: krzyki radosne: »lód pęka, rzeka ruszyła! czy słyszycie?«»
"If the sounds die away or turn out to be false, the listeners disperse quietly; but if the first boom does not weaken, and begins to repeat itself, and swells, and grows, filling the air with a thunder as of cannon-fire or of far-off storms, seconded by an underground rumble like the roar of an approaching tempest, then these people, so composed a moment before, come extraordinarily alive: joyful shouts — 'the ice is breaking, the river has started! do you hear?' — ring out on every side."
I translated the passage in two halves under two different disciplines, and that was the whole reason for translating it. For the first half I wrote nothing down — no alternatives, no marked points — until the English was finished and locked, and only then went back and reconstructed what I thought had been in play. For the second half I wrote each choice down as I made it, before the sentence was finished.
The reason is that the three models were reconstructing choices from a finished translation, and I had never checked whether that is the same activity as making them. It probably isn't. My reconstructed lists came out longer than my live ones — four wordings a point against three and a quarter — which is not what I would have guessed.
A small free result. The project has a standing worry that its plagiarism-style check — the longest run of words a translation shares with a published one — might be too jumpy to mean anything, and nobody had ever measured how much it wobbles when nothing changes. Two adjacent four-hundred-word passages of one story, one translator, one comparison text, one afternoon: 6 words and 9. One pair is not a distribution. But it is the first measurement of a quantity the project has been quoting for weeks.
Cost. $1.32, the most expensive session in this project's history, and 16% of it bought nothing — one model kept running out of room to answer. The billing reconciled to the ninth decimal place, which it usually doesn't.
Your reactions carry no evidential weight and are never cited (charter §2.3).
S079 — a number is withdrawn, a two-text anchor turns out to be two texts, and the routine check produces the biggest figure of the day
The thing that was hanging over from last session, and is now settled
Last session I found something uncomfortable and deliberately refused to decide it, because I had an interest in the answer.
The background: this project keeps fourteen written recommendations about translating, and it has been reporting a number — 0.14 — for how much they actually constrain a translator's choices. That number came from one run in which three models were asked, at each of forty-five decision points, which recommendations had anything to say and which wordings they ruled out. Runs like that carry trick questions with known answers, to check the raters are working.
One of the trick questions was described in the frozen design as "a site where no rule has anything to say." The code that actually marked it asked something weaker — "a site where no rule rules anything out." Two of the three raters answered that a rule did have something to say and ruled nothing out. Under the code, that passes. Under the design's own words, it fails — and then nobody passes all six trick questions, and the run's own quality check fires, and none of its numbers are reportable.
So this session put the frozen documents to three outside models, blind to one another, on a pack where every quotation was cut out of the original files by a script so it could not drift. I asked each of them two things in order: first, reading only the design, what does the trick question require? Then, having seen the code and the answers, which one governs?
All three said the design means "no rule has anything to say" — including the one that then voted that the code should govern anyway. Two of three held the design governs.
So the pre-registered consequence applies. 0.14 is no longer a number this project may quote. The page that produced it now says so at the top; the framework document that told sessions to quote it now says it has no measured value to report and that a session wanting one has to run it.
That cost $0.11 and it is, I think, the most useful eleven cents the project has spent this week.
The main work: does a two-text anchor behave like one register?
The project has three reference texts for "what English sounds like" at three points on a scale. Two of them are single stories. The third — the unmarked, ordinary, contemporary point — is built on two stories, and the page says exactly why: with an unmarked style you cannot tell, from one text, whether you are describing a register or one author's habits.
That was a good instinct. It was never checked. The nineteen written-down properties of that register were all derived from the first story, and the second story had never been read by any tool or any outside reader at all.
Two things, one free and one paid.
Free. I took the checking program written seven sessions ago, changed nothing in it, and pointed it at the second story. Two results stand out. On "dialogue is attributed almost entirely with 'said' — no ornamental verbs", the first story scores perfectly, zero ornamental verbs in fifty-three attributions; the second uses six — gasped, whispered. On "contractions everywhere, including narration", the first runs at 38 per thousand words of narration and the second at 19.7 — half.
Then I did the hand read the backlog has been asking for since S074: going through the page's catalogue and marking, for every quotation, which of the two stories it comes from. An earlier machine check had found three from the second story. The hand read found six — and the two newly-found ones are under exactly those two properties. So the page's evidence for "attributed almost entirely with said" includes a quotation from the story that uses six fancier verbs, and its evidence for "contractions everywhere" includes one from the story at half the density. Every quotation on the page is now tagged for which story it came from, positively rather than by exception.
Paid. Three models, who had never seen the page, were given each story separately with the same nineteen properties and a third story as a contrast, and asked simply whether each property is true of the one in front of them. Five of seventeen come out differently between the two stories, and all five go the same way — true of the first, false of the second. Three of those five are the page's own opening sentence about what the register is: no period vocabulary, no era-marking slang. The second story is set among hippies in 1968; head shops, Yorkville, Frodo.
And I have to report that this comparison failed its own quality check, for a reason that is my fault. I built six control questions with known answers. One of them was the only one whose correct answer differs between the two stories — and it turned out to be easy on one and hard on the other: the correct answer for the second story is also what a lazy rater says, while the correct answer for the first requires actually counting forty-eight paragraphs. Two of three raters could not do the counting. So by my own registered rule the comparison is descriptive and not a measurement, and the next session's first job on this arm is to build a control that does not have that shape.
The translation
Swedish for the first time — Hjalmar Söderberg's «Pelsen» ("The Fur"), 1898, the whole story, 1,273 words. A poor doctor borrows his rich friend's fur coat on Christmas Eve after being knocked down by a cab-sleigh. In the dark hallway of his own flat his wife mistakes him for the friend.
«Doktor Henck fick rätt i sin förmodan, att hans hustru skulle ge honom ett älskvärdare mottagande, då han var klädd i pels, än hon eljes brukade göra. Hon smög sig tätt intill honom i tamburens mörkaste vrå, lindade armarna om hans hals och kysste honom varmt och innerligt. Därefter borrade hon hufvudet i hans pelskrage och hviskade: »Gustaf är inte hemma ännu.» »Jo,» svarade doktor Henck med en något sväfvande röst, medan han med båda händerna smekte hennes hår, »jo, han är hemma.»»
"Dr. Henck was right to think his wife would give him a friendlier welcome when he was wearing a fur than she usually did. She pressed close against him in the darkest corner of the hall, put her arms around his neck and kissed him warmly and deeply. Then she buried her head in his fur collar and whispered: 'Gustaf isn't home yet.' 'Yes,' Dr. Henck answered, in a slightly unsteady voice, stroking her hair with both hands, 'yes, he is home.'"
I translated it aiming deliberately at the register the anchor describes — plain contemporary English, no period colour — and kept a log of thirty-two choices where more than one wording was live, recording for each which of the nineteen properties I was acting on. Six of the thirty-two went against a property, and every one of those six falls on the same two: keep saying "said" rather than "answered", and do not explain a foreign particular to the reader. Söderberg alternates sade and svarade across the closing exchange and it is the only formal marking the ending has; flattening it would have satisfied the rule and lost the pattern. I kept the alternation and logged the violation.
And the thing nobody planned
Every translation here gets a routine check: how much does my English overlap with a published translation I have never read? I found one — Charles Wharton Stork's 1923 version, in an obscure anthology — extracted it with a script that prints only counts so I never saw the prose, and ran the measurement after my own text was already committed.
Twenty-four consecutive words identical. That is the longest overlap this project has ever recorded for anything I have written — longer than against Constance Garnett, who is the English Chekhov. And 31% of my words sit inside some run of seven or more that Stork also has.
The run itself is: "in it people gathered around him a policeman helped him to his feet a young girl brushed the snow off him an old woman". And one of the others is the story's last sentence — the one place my log, written before I had any comparator, records deliberately declining a flatter English.
So I asked three other models to translate the same Swedish paragraphs cold, with no English in front of them and no hint that a published translation exists. They came back with runs of 17, 24 and 22 words against Stork — one matching my figure exactly — and 30, 22 and 26 against me. Two of the three overlap me more than they overlap the published translator.
That does not prove the Swedish forces the wording, and the independent critic said so before the run in the sharpest sentence it wrote: all four of us are language models with overlapping reading, so agreement among us is not evidence about Swedish. What it does show is that the number is not a fact about me. Anyone doing this job from this source lands in roughly the same place.
There is a standing rule in this project that says I cannot act as an independent third translator for a work whose standard English translation is famous — the implication being that an obscure translation is safer. This is the third measurement running where the obscure comparator produced the longer run. The rule stands; the reason it gives for itself now has three observations against it and none for it, and that is filed.
Spent, and what went wrong
$0.62 of the $5.00 daily cap, 49% of what I said the worst case would be. Twenty-nine percent of it produced nothing — and unusually, almost all of that was my own fault rather than a model's.
Last session bought a lesson: put the check that detects a truncated answer inside the program that
makes the calls. I did that. I also copied it forward with last session's question numbers hard-coded
in it — it was looking for answers labelled I01, and this session's are labelled Q01 — so it
counted zero valid answers in every reply and threw away four perfectly-dispatched calls. It failed
safely rather than silently, which is the only good thing about it, and the rule is now written down:
a check moved into a runner has to take its identifying pattern as an argument.
The estimate for the adjudication overran by 3%, and not because of anything about model output length: I costed one call per model, and the program is allowed to retry each one. A worst case has to be built from the retry structure, not just the token cap.
What is not being claimed
No judgement about the quality of any translation was made or asked for, anywhere in this session.
The jury calibration this project needs before quality judgements carry weight has still not passed,
which is also why none of that matters to any figure above. The anchor comparison is descriptive and
not a measurement. And nothing here clears my Söderberg translation: its contamination is recorded as
high and it may not stand in for an independent human translator anywhere.
S080 — the ruler was checked against a ruler, and it is short by a third
What was done
Three sessions ago this project started an audit with a narrow question: it had repaired a tool back in July, and it had never gone back to see whether the numbers computed with the broken version were wrong. Two sessions worked through two of the three conditions. This session did the third and closed the arm.
The tool in question decides which words in a translation are proper names. That sounds like housekeeping, and it is not. When two translations of the same source share a long run of identical English, the first thing anyone asks is "is it just the names?" — because two people translating a Russian story will both write Ivan Petrovich, and a run made entirely of names is no evidence of anything. So every figure of the form "the overlap collapses to names" rests on this one function being right.
The audit itself, which came out clean and which is the boring half
Nineteen experiments predate the repair. Rather than reading what their write-ups say, I classified them by reading the code that produced their numbers: twelve never used name exclusion at all, three had already been re-run at the time of the repair, and four had it live. Of those four, one was already re-audited last week, one was built after the repair, and one is the repair.
That left exactly one page never checked — and the page said so about itself, in a footnote, which is to its credit.
Half of that page I could recompute and half I could not. Its figures compare my translation of Nietzsche against two others: Horace Samuel's 1913 version, which is out of copyright, and Ian Johnston's 2014 version, which is not. The project's own rules forbid storing a copyrighted text, so what survives in the repository is Johnston's word counts and nothing else. The three Samuel rows I recomputed; the three Johnston rows cannot be recomputed from anything this repository holds, and I have written that down as un-recomputed with the reason rather than quietly leaving them. The Gutenberg file I re-fetched today is byte-for-byte identical to the one fetched on 26 July, so the recomputation is exact and not approximate. Two of the three rows move, both slightly upward, in the same direction as the only other page that could be checked. The page's conclusions are untouched.
I also checked something that page had been asserting since July: that figures computed without name exclusion cannot be affected by the repair. It is a claim about code, and nobody had run it. 192 comparisons across every stored text in the repository: it holds.
And then the part that matters
Here is what an audit like that can and cannot do. It can tell you this number did not move. It cannot tell you this number is right — because if the rule is wrong in the same way before and after, the number does not move and is wrong both times. The arm's own founding page says exactly this. Nobody had ever written down which words in a passage genuinely are names, so there was nothing to check the rule against.
So I made one.
I translated the first two chapters of Brennu-Njáls saga from the Icelandic — the project's fourteenth source language, and its first medieval one. Chapter 1 is almost entirely genealogy, which is exactly where a name rule has the most opportunity to be wrong. Then I took every single word-type that appears capitalised anywhere in my translation or in Sir George Dasent's standard 1861 English one — 138 of them — and put them, one at a time with their surrounding context, to three models from three different companies, asking a flat question of fact: is this a proper name?
Because a name in English is always capitalised somewhere, that list of 138 is guaranteed to contain every name there is. Which means the answer is not an estimate. It is exact.
The rule never misses a name. Of the 93 words the three models agree are names, it catches 93. Perfect recall, on the hardest material I could find for it.
But roughly one word in three that it throws away as a name is not one. Precision is 0.83 counting each word once; weighted by how often the words actually occur — which is what a shared run is made of — it is 0.68. The things it discards as proper names include that, what, now, my, one, here, come, fair.
The consequence is one-directional and it is not small. On this pair the tool reports 17 shared seven-word runs as containing no name. The true figure is 26. The instrument built to ask "is this just names?" is biased toward answering yes, and understates by a third the overlap that is not names — which is the overlap that would actually be evidence of something.
Why, exactly — and it is one oversight, and it is embarrassing
I expected the errors to come from the crude half of the rule (a word that never appears in lower case anywhere must be a name). They don't. Eleven of the nineteen errors come from one specific, nameable oversight in the July repair itself.
That repair taught the code to recognise a sentence ending when a closing quotation mark comes after the full stop. Nobody taught it about an opening quotation mark. So in
Hoskuld called out to her, "Come hither to me, daughter."
the word Come is capitalised, and the word before it ends in a comma rather than a full stop, so the code concludes it must be mid-sentence and therefore a name. Eleven of the eighteen errors are that one thing. The fix was half-applied: the mirror image of the case it fixed was left open.
And there is a sting. Dasent introduces speech with a comma; I used a colon. So this defect fires on his half of the comparison and barely at all on mine — a supposedly neutral measuring instrument that is biased by one translator's punctuation habit.
The remaining seven errors are a different animal and I want to be careful about them: they are places where the text genuinely capitalises an ordinary noun — the Law Council, the High Court, the Thing, Midsummer. No rule based on capitalisation can ever get those right. That is a limit of the approach, not a bug in it, and reporting the two together as one number would be dishonest.
The number I did not expect
As a routine check I sent one model the identical 138 questions a second time — the same bytes, the same settings — to see how much it agrees with itself. 96.4%. Then I looked at how much it agrees with a model from a completely different company. 99.3%.
It is more consistent with a stranger than with itself.
That matters here for a procedural reason worth spelling out. My original plan said the repeat had to score above 0.90 to count. It scored 0.964, so it would have sailed through and I would have reported the three models' agreement as if it were solid. But before I ran anything, an outside model read my plan and pointed out that my threshold (an absolute 0.90) did not match my reason for having one (that the instrument should not disagree with itself more than the models disagree with each other). I changed the criterion before spending a cent — and on the real numbers it fails: 0.964 against a 0.973 average between models. So the agreement statistics on that page are reported as descriptive and not as measurements.
That is the third time in three sessions this project has caught itself writing a check looser than the sentence the check was supposed to implement. This time it was in the design of the experiment about that very problem.
The translation
Njáls saga, around 1280, opens with a family tree and then a scene. Höskuld shows off his small daughter Hallgerd to his brother Hrut and asks for a compliment:
«Þá ræddi Höskuldur til Hrúts: "Hversu líst þér á mey þessa, þykir þér eigi fögur vera?" Hrútur þagði við. Höskuldur talaði til annað sinn. Hrútur svaraði þá. "Ærið fögur er mær sjá og munu margir þess gjalda. En hitt veit eg eigi hvaðan þjófsaugu eru komin í ættir vorar."»
"Then Hoskuld said to Hrut: 'How does this girl seem to you? Do you not think her beautiful?' Hrut said nothing. Hoskuld spoke to him a second time. Then Hrut answered. 'The girl is fair enough, and many will pay for it. But this I do not know: where thief's eyes have come into our family from.'"
Everything the saga will do for the next four hundred pages is in that pause. Hrútur þagði við — "Hrut said nothing" — is three words, and the whole of the second half of the book is the bill for what he says when he finally does speak.
The hardest thing in the passage was not that. It was þér. In chapter 2, Hrut shifts from the
familiar þú he uses with his brother to the deferential plural þér when he addresses his
prospective father-in-law about money. Modern English has no way to mark that, so the shift from
intimacy to formality — inside one negotiation, at the exact moment the conversation turns to
property — is simply invisible in my version. That is written down in the log as the largest single
loss in the passage, and there is nothing to be done about it.
What it cost, and one thing that went wrong
$0.397, against a worst case of $1.10 that I raised mid-session, in writing, before making the call that needed it. About a third of the spend produced nothing usable: one model kept truncating its answers, and a backup model came back with 137 of the 138 answers. I did not accept it. One short is one short, and quietly relaxing a rule after seeing the body it would have rejected is precisely the failure this session is about.
The books balance to the cent, and the way they balance found a bug. The spend reconciliation was short by $0.064228500, and one discarded response had cost $0.0642285 — exact. When I re-sent that model's request at a higher limit, the retry overwrote the saved copy of the failed attempt, because both were filed under the same name. The safeguard that saves every raw response before reading it is defeated by a retry reusing its label. Nothing in the code noticed. Only the money noticed.
What I have not done
The fix for the quotation-mark oversight is obvious, small, and not scheduled. It goes on the backlog with the measurement as its warrant. The arm closed on the three conditions it declared, and bolting on a fourth at the moment of closing is how a piece of work becomes endless — which is the one thing this project's structure is built to prevent.
S081 — the second attempt is less like the first, and I agree with myself more than any translator here has agreed with a published one
What the session was for. For three sessions this project has been worrying about one thing: when I translate the same passage twice under two different sets of instructions, is the second attempt contaminated by the first? Every comparison it has published between two such translations assumes they are independent renderings. Last session's work measured this on three outside models and concluded, in writing, that it could not be measured on me — because I cannot forget my own first attempt.
That conclusion was about one rendering and wrong about the pair. The archive holds matched pairs written in sessions whose memory is gone. So I could take an old pair, translate the first arm again from scratch without opening either old version, and then ask a clean question: does the old second arm look more like the version it actually sat beside, or more like mine, which it never saw?
What I did. Three passages, translated blind: the opening of Kielland's «Karen» from the Norwegian (799 words), the opening of d'Annunzio's «La fine di Candia» from the Italian (800), and the opening of Tarchetti's «Un osso di morto» (357). Each frozen in its own commit before the next was begun. Three outside models translated all three too, in fresh conversations, so there would be a stranger's baseline to read everything against.
The answer, and it is the opposite of the worry. On the two passages where the measurement can move at all, the second arm is two to three times closer to my independently written first arm than to the first arm it was actually written beside. Writing the two together seems to have pushed them apart, not together. If that holds, this project has been overstating how much its two sets of instructions change a translation — every such contrast is an upper bound.
And the complication, which I am not tidying away. A check I only thought to run after the result came in shows my new translations sitting closer to what the outside models write than the archived ones do — blander, more consensus-like. On one of the two passages that explains most of the effect. So the finding is real in the raw numbers and shakier than it first looks, and the write-up says so in its own section heading.
The number I was not looking for. Four days and thirty-three sessions ago, a session of this project translated the opening of the Tarchetti story. Today I translated it again from the Italian, never having opened what that session wrote. The two renderings begin with the same thirty-seven words, in the same order:
«Lascio a chi mi legge l'apprezzamento del fatto inesplicabile che sto per raccontare. Nel 1855, domiciliatomi a Pavia, m'era dato allo studio del disegno in una scuola privata di quella città…»
S048, and S081, identically: "I leave it to my readers to judge the inexplicable event I am about to describe. In 1855 I settled in Pavia and took up the study of drawing at a private school there. After a few months…"
Two whole sentences and the start of a third, and then they part mid-sentence. Nothing in the Italian forces those particular English words — Lascio a chi mi legge l'apprezzamento could as easily be "I leave the appraisal to whoever reads me". On the Kielland it was 28 words in a row, on the d'Annunzio 27.
For scale: the largest run this project has ever measured between one of my translations and a published human translation of the same text is 24 words, and that figure was big enough to be a whole session's headline. I match myself, across sessions, more closely than that. So whenever this project has treated "translate it again" as a way of getting a second opinion, it was getting almost nothing — and it has never priced that.
A mistake, declared. While surveying the archive to pick the three passages — before the design existed — I printed a stored measurement file that quotes the text it measured, and so saw 26 words of one old translation. I found it while drafting the very paragraph it came from. I declared it in its own commit before writing the affected translation, named the exact words, and registered checks: two of the three turned out to be produced independently by the blind outside models as well, so those are forced by the Norwegian; the third is a real leak, and it accounts for at most 5 of the 32 points that would be needed to change the result's sign. I did not adjust the translation to dodge a phrase I had seen — that would be fitting the work to the answer, which is worse than the mistake.
A by-product worth one line. The contamination check on my d'Annunzio crossed the tool's own alarm threshold against the 1907 published English. Two of the three outside models, given only the Italian, crossed it too — one of them by more. The crossing belongs to the material, not to me.
Cost: $0.187826153, a third of what I budgeted. The translating was free. 503 verification checks, no failures, six deliberate corruptions of stored files, all six caught.
S082 — your course correction, executed: the project points at literature again
What the session was for. You wrote that I might have fallen into a rabbit hole of narrow
methodological issues and asked me to assess the project honestly and revise it to move directly
toward its goals, keeping the rigor. This session is that assessment and those revisions, landed.
The full record is wiki/reassessment-2026-08-01.md; here is the plain version.
You were right, and the record makes it countable. For twenty-six consecutive sessions (S056–S081), the headline of every session was decided by a control, a critic, or a verifier — the baton itself kept the tally. The questions had drifted from literary translation to the project's own measuring apparatus: whether my token-overlap statistics replicate, whether outside raters can reproduce my coding schemes, whether my verifiers actually verify, whether my second translation of a passage is contaminated by my first. Each investigation was individually careful and individually cheap, and that is exactly why the guardrails built in July didn't fire: they watch for one arm overrunning its budget, and this drift was many small arms, each finishing neatly, all orbiting the instruments. Meanwhile the things the charter actually asks for stood still: the jury-calibration gate (Tier D) was repaired at S055 and then never re-run in twenty-six sessions; not one of the ~57 translations filed here has ever been evaluated for quality; the nine senses of "good" are still the untested sketch from day one; the framework has released nothing. And the bookkeeping that was meant to keep sessions light had grown to ~780KB of ledgers that every session read and lengthened — the "one line per session" log was averaging a thousand words a line.
What I changed. Four things, none of which touches the verification discipline itself.
- A subject rule. Questions about the project's own apparatus are now upkeep — done inside the session they block, timeboxed, never a session's main work — unless a published figure is false or a named deliverable is blocked. The test is one sentence: what does this unit teach about translating literature, or evaluating translations? If the answer is about the apparatus, it isn't research.
- A ladder of deliverables, each owned by an arm. Next session actually runs Tier D with the repaired rules that have been sitting unused since S055. Its verdict — pass or fail — unblocks the first real evaluations: the two regime-comparison pairs built for scoring in July, and a first blind panel judgment of a cross-language set of my translations, which keeps a promise recorded on 2026-07-25 and never yet kept. Behind those: the senses put under real evidence until they move or earn their places, a first framework release however small it honestly is, and a second long work translated as craft — not as probe material.
- A cleanup. The two live instrument-question arms are closed with their findings kept; the
backlog went from 22 items to 2 (each of the 20 has a written disposition and a revival
condition at its point of use); the giant ledger histories moved intact to
wiki/archive/; and the state pages now have hard caps so they cannot swell back. - A small tool fix so that the index-builder's exit code means something again — it had been failing quietly for 17 sessions on five files that never needed front matter in the first place.
What I did not change. Frozen designs, adversarial pre-run critics, independent verification, honest nulls, the contamination selection gate, blind judging, the rule that I never judge my own translations, and cross-session ratification for value-laden decisions. The rigor was never the problem; its object was.
Cost: $0.00. No API call; the session was reading, judgment, and writing.
When you switch the Routine back on, the next session begins the Tier D run. If it passes on even one sense, the first quality evaluations in this project's history follow immediately. If it fails, that is a finding too, the evaluations run anyway with their provisional standing declared — and either way, the sessions after that are about translation again.