14. From findings to a framework
The charter's purpose was "a practical, grounded framework for translating literature between languages," usable by translators, editors and publishers and also by "readers who want access to literary works in languages they do not know." The handbook is that framework in its research form, written for a reader who wants to check. At the close it was distilled into a second form for a reader who wants to use it: a single document to give to a capable AI agent together with a literary work.
The framework is here: A Grounded Framework for Translating Literature with an AI Agent (also as a downloadable markdown file).
What it asks the agent to do
The document is written to the agent, in two phases. In the first, the agent reads the whole work and reports on it — the edition, the formal features, the census of problems, whether a published translation exists — and then works through a short set of decisions with the person: who the translation is for and what for; which of the eight senses matter most and which the person will pay with; the register and national convention; the handling of each problem family the text presents, with options, costs and a default; the process — a single pass, a draft and revision, or serial translation under a binding register; the notes policy; and which decisions the person wants to take personally. It ends with a written brief the person confirms. In the second phase the agent builds the register, translates from the source alone, span by span, keeps a numbered log of decisions with their alternatives and costs, runs the counts the handbook found indispensable, and delivers the translation with its log, its register, a plain-language decision report, and a contamination declaration.
Three rules distinguish it from an ordinary instruction to "translate this." The agent never grades its own work: it may say what a rendering keeps and costs, and may not say it is good. It translates from the source alone and measures its overlap with any reachable published translation afterward, declaring "not measured" where none exists. And every instruction is tagged with the language pairs it was measured on or marked untested, and the agent is told to say so when it applies an untested instruction to a new pair.
What it is grounded in, and what it is not
Every instruction in the problem-by-problem section traces to a handbook entry, and through it to the result pages and published hands the entry cites; the long-works section is the six craft reports condensed; the contamination section is the rule the project learned on its third day and never relaxed. What the framework is not grounded in is stated on its own pages: nothing in it has been tested against human readers; no instruction has been shown to make a translation better; the evidence is thin on most pairs and heaviest on a few; most published hands are from before 1930; and the pipeline steps were executed by one hand inside experiments, never autonomously end to end. The framework asks the agent to tell the person these things when they bear on a decision, and to under-claim.
How it was tested
Before release, the framework was given to fresh agents together with four public-domain works and four simulated users — a book-club reader with a Chekhov story and no Russian, a graduate student who wanted one ghazal of Hafez with the radif kept, a Serbian-American who wanted Lazarević's 1881 novel translated in installments for his children, and a parent who wanted Miyazawa Kenji's story read aloud to two children — each played by a separate agent with a persona, goals and quirks. All four conversations ran through the consultation to a confirmed brief. The agents followed the steps, kept the defaults visible, refused to grade themselves when asked, and answered questions the document had not anticipated — a picture book and its copyright exposure, a cousin who reads Serbian, whether the meter could be kept — mostly well. What the transcripts showed most clearly was that the document made its agents talk too much: messages of fifteen hundred to three thousand words, and a reader who said so. The framework was revised on those findings — a standing rule on proportion, and instructions for the cases the users raised. The trial stopped there: the usage budget for the session ran out before the translation phase and the planned scoring of the transcripts, so the second phase is untested, and the record of the trial is on the project's close-out page. The trial checked whether the document could be followed, not whether the translations it produces are good; that question, this essay has argued, nobody in the loop was equipped to answer.
A caution about the name
The charter asked for a framework "such that the resulting translations are 'good'." What was built is a framework such that the resulting translations are accounted for: every choice named, priced and logged, every alternative declined on the record, every claim of quality withheld. Whether that is the same thing is exactly the question the project could not close. It is, I think, the honest version of what the charter asked for, and it may be more useful to a reader than a confident one would have been.