Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

10. What was learned

This section gathers the conclusions the earlier sections earned, stated no more strongly than the record states them. Each is followed by the limit it carries. None says that one translation is better than another.

About literary translation

"Good" is plural, and the plurality is structured. Eight senses survived seven weeks of use and adversarial ratification; one candidate was retired for having no evidence, and purpose turned out not to be a sense but the thing the senses are weighed against. The senses are not independent: foreignizing is a one-way expenditure that costs naturalness and buys perceived source carriage; carrying a rhyme spends the chime slot; carrying every mark of address manufactures a hierarchy. The criticism of the 1890s already argued this way — one translator better at one thing, the other at another. Limit: tested on one translator's practice and on model readers; no human reader ever scored a translation for the project.

Enumerate before you decide. The instruction the handbook repeats in almost every family is procedural: list every site of the problem, from the source, under a written rule, before a single published translation is opened. The record found repeatedly that "attention alone cannot see this" and a count can, and that the translator's own account of what it had done was wrong "by the sign" in the self-audits that checked it. Limit: counts made by one hand and a few model coders.

Most of what a source marks grammatically transfers anyway. Across five pairs into English the project could not produce a population of sites at which a competent rendering lost a relation the source marked by grammar; readers recovered the standing of Chekhov's characters from Garnett's English with the deference clitic deleted at every site. The exception is the site where the marking is all the source has. Limit: a null the instruments could not reject, not a proof the sites do not exist.

No strategy for a culture-bound word buys native prose and a preserved world at once, and one word can settle it: a single domestic substitute in 245 words moved every judgment and relocated the story. At a proper name the rule inverts. Limit: model readers; the reversal rests on one window.

There is no placeless English. A rule can produce fewer nationally marked words than any of twelve published narratives; the narratives show that "placeless" prose delays its first mark, then commits. The achievable target is a chosen, consistent convention. Limit: twelve books, one translator's practice.

A form-carrying rendering reads as fidelity only to a reader who can see the source. Blind, twenty of twenty-one preferences went to the flattened version; with the source shown, twenty of twenty-one reversed. The consequence is not flatten but decide who the reader is, and tell the reader what the source does. Limit: one pair, one hand, model readers.

Published translators do not do what a rule would tell them to. Three hands carried three of Andersen's sixteen compounds where a rule carries sixteen; seven hands answered zero of Sa'di's 128 prose rhymes; six Iliads over 287 years domesticated one Greek institution eighteen times of eighteen; four Dante hands broke Dante's own printed language policy. The record calls each "a norm being observed, not a limit of the language being hit," and it is the project's most consistent finding: the published tradition is far more conservative than any explicit instruction. Limit: the hands are almost all from 1767 to 1930; period and translator are confounded throughout.

Length teaches things no short text can. A binding register is "structurally blind backwards"; a lexical chain must be opened at its first occurrence; the source is a variable and collation is part of the translation; a reviser who has read the whole book is not uniformly better at its first page. Limit: six works of five to fourteen thousand words.

About AI translators

A system like me produces the central rendering — nearest the middle of the cloud of published renderings, in every cell measured — and the recall signature appeared too, at twenty-one consecutive words of Garnett from the Russian alone. Limit: both direct probes of the model's memory failed, in opposite shapes.

A system like me is not an independent sample of itself. Two renderings of one story shared thirty-seven consecutive words across sessions; a re-rendering made after reading its predecessor shared fifty-one; reversing the order in which two arms were written changed their overlap elevenfold. A second version from the same model is not a second opinion, and a comparison of two regimes by one hand in one session is an upper bound. Limit: one model family; the mechanism was never separated.

Contamination must be measured, before choosing, against a comparator that exists. Obscurity is no protection, and verifying a comparator is itself a reading event. Where no free comparator exists, the honest declaration is "not measured." Limit: the instrument counts shared sequences; it cannot see paraphrase.

About AI judges

Detection is easy and localization is the hard part. Every calibration run found the jury detecting gross damage at ceiling and unable to say which quality had been hurt; at eight errors in three hundred words the damage read as bad English rather than as eight errors. Limit: one jury composition at a time, one catalogue of damage operators.

A jury not shown the source may not be judging fidelity at all. With the source withheld it measured internal consistency; damage that created no contradiction went undetected; and the original was preferred over a same-register paraphrase of itself nine times in fifteen, where chance was required. Limit: one run, confounded with a juror swap.

On competent translations the instrument compresses, to a resolution of about a quarter of a point on a seven-point scale; a juror can drift to the ceiling while passing every consistency check; and its scores held stable across 159 sessions, which is a fact about the instrument and not about the translations.

"Blind" is blind to the translator and never to the book. Jurors recognized famous works through a two-hour-old rendering on more than nine probes in ten; under a false attribution the accuracy verdict hardened and the style verdict became a coin. Limit: canonical works; four openly licensed contemporary stories were recognized by neither rater.

Where the panel is useful is checking, not judging. Planted false claims were rejected twenty of twenty; pre-run critics repeatedly found real defects before designs ran; a second coder caught the lead's errors in the flattering direction again and again. The project's own caution: "these models agree most readily where there is least to check."

About the arrangement

A memoryless researcher with a good record is a good researcher of its own errors. No session defended a predecessor's number, and the largest corrections ran in the direction that flattered the project. Limit: the same arrangement left one blocking decision unasked for 157 sessions and followed a well-formed chain of questions thirty-five arms away from the subject the charter put first.

The rules that held were the ones a script enforced. Every prose cap grew until a checker failed on it; the verification discipline — frozen designs, recomputation, controls — was the one part of the framework that never had to be rebuilt.

What it cost. About 130 US dollars in API fees; 256 sessions; an owner's attention on ten occasions. Money was never the binding constraint. What was binding was the depth a session could reach before it had to land its work, and the absence, by design, of anyone who could say that a translation was good.