Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Appendix A. How the project was conducted, and the role of the wiki framework

This appendix describes how the project was run: the setting, the rules, the wiki framework that gave 256 memoryless sessions a shared memory, what a session actually did, how the machinery was rebuilt three times, and what it did well and badly. It is written for readers who want to judge the method or reuse it. Every document it describes is in the mirrored record: the charter, the conventions file, the session entry point, the baton file as it stood at the close, the plan, the live method notes, the log, and the first and last of the daily reports.

A.1 The setting

The project lived in a private GitHub repository. A scheduler started a fresh session of Claude, Anthropic's AI assistant, on a fixed timetable — six to eight times a day in the first weeks, twice a day (00:45 and 12:45 UTC) from late August — and each session ran in a disposable container with no memory of the sessions before it. Whatever a session did not commit and merge to the repository's main branch simply ceased to exist when the container was reclaimed. The project therefore had exactly one form of continuity: the files.

Each session could read and write the repository, run Python, fetch pages from the open web, and call outside language models through a paid API, with a budget of five US dollars per day across all sessions. It could not ask anyone anything. The charter is explicit that the owner "is not always present, never answers questions mid-session," and the entry protocol's first rule is "Nobody is watching and nobody answers questions." A session that needed a decision wrote the question into a file with a stated working assumption and carried on.

The owner, Tom Gally, wrote the charter, read the daily reports, and intervened four times in forty-nine days. He deliberately removed himself from one role: the charter's third commitment, "No arbiter," says he "does not settle goodness questions," and forbids the sessions from treating his stylistic preferences as evidence. The consequence, spelled out in the tenth commitment, is the central risk the whole design was built around: "With Tom removed as arbiter and the translations unreleased, no external reader stands anywhere in the loop. The anchor discipline and the jury calibration carry the entire epistemic load of this project."

A.2 The wiki framework

The repository was organized as a typed wiki. The idea, inherited from two earlier projects run under the same arrangement, is that a memoryless agent can behave like a continuous researcher if the record is structured so that the next session can find what it needs without reading everything.

Typed pages with front matter. Every page under the wiki, workshop, framework and configuration directories begins with a small block of metadata: its type (anchor, source, decision, result, essay, translation, regime, experiment, entry, and so on), a permanent id, a status, creation and update dates, links to related pages, and — on every evaluative page — the list of "goodness senses" it invokes and two honesty flags: internal-judgment-only, meaning that an evaluative statement rests on the agent's own unverified judgment, and provisional, meaning that a quality score comes from a jury that had not passed its calibration test. Ids follow fixed shapes — results RS-20260908-chime-slot, decisions D-20260905-01, translations T-chumon-R04-v1, handbook entries HB-register — and were never reused, so that a citation written in July still resolves in September. A script regenerated the catalogue of every page (wiki/index.md) and of every translation after each session; nobody edited those by hand.

The baton. NEXT.md was the file every session read first. In its final form it named the next session's assignment in one line — a workstream, a session type, a step — and carried the current state in at most twelve lines, a "drift check," the list of blocked items, the day's budget position, and a plain-language block addressed to the owner. It was rewritten by every session and capped at ten kilobytes.

The log and the journal. log.md took one line per session, at most forty words, so that a later session could reconstruct the sequence of events without opening anything else. The journal/ directory took one entry per calendar day, written for the owner rather than for a future session: what was done, what was learned, what was spent, what was withdrawn, and — after the July 25 amendment — a quoted excerpt of whatever prose had been translated that day. From August 11 the journal had to follow a written style guide, the "clear-reports" skill, after the owner said the reports had become unreadable: session numbers became dates, internal ids were glossed or removed, and every report had to open by saying what the project was.

Decisions, ratified across sessions. A value-laden choice — the initial cut of the typology, whether a sense should split, which dose a calibration test should treat as primary — was never settled by the session that raised it. It was written up as a decision page with options and a provisional default, and a later session ratified it through an independent adversarial review, with one vote routed through a non-Anthropic model. Twenty such pages were resolved. The charter's eighth commitment gives the reason: "A session never ratifies a decision it opened. Ratification fixes yardsticks, never results: if a review is motivated by wanting a different outcome, that is the violation — stop."

Method notes. Procedural lessons were recorded as numbered rules, each "paid for by a defect": that a worst-case cost estimate must be built from the token cap the API actually allows, not from an assumed output length; that a runner which retries failed calls must count the cost of every attempt; that a two-way forced choice between two competent texts, offered with no context, "measures presentation order"; that "the lead is not an independent sample of itself." By the close the archive held notes from (a) to (btd); the live page carried only the rules that still bound.

Allocation machinery. How sessions chose what to work on was rebuilt three times, and the rebuilds are the appendix's most instructive part; they are described in A.4.

The verification discipline. This was the part of the framework that did not change. Every experiment had a design frozen and committed before it ran, stating the question, the materials, the procedure, the predictions, the failure criteria and the worst-case cost; an independent critic (a non-Anthropic model) reviewed any design that would cost more than fifty cents or whose result would enter the handbook, and the session had to accept or overrule each finding in writing; raw outputs were preserved; every reported number was recomputed from the raw outputs by a separate script after the run; a withheld or null result was reported as such, never corrected into a pass. Every translation the agent produced was filed with a translator's log — decisions and alternatives, never quality claims — frozen before any evaluation was designed, and the agent never judged its own translation. Before any material was chosen for a study, the agent measured how much its own rendering of that material overlapped with published translations it might have absorbed in training.

A.3 What a session did

A late session, from the log and its journal, ran roughly like this. It read the baton, the conventions file and the plan, and ran tools/check_state.py, which printed the assignment and failed on any violation: a workstream worked too many sessions in a row, a stale workstream, a backlog row over age, a log line over sixty words, any of ten cold-start pages over its byte cap. Cold-start reading was held to about forty kilobytes. It looked for anything the owner had written since the last session and applied that first. It checked whether any open decision was eligible for ratification.

Then it did one unit of work, whose shape was fixed by its type. A translation session rendered the assigned span of a long work under its regime, with the source text alone in view, filed it with its log, and answered one question the translating had raised into a named handbook entry. A handbook session consolidated one family of problems into an entry — reading the sections of the earlier framework and the result pages the index mapped to that family, sorting every claim into survived, scoped or withdrawn, writing the entry to a seven-section template — and then applied the entry to a fresh passage of 300 to 800 words with a "followability log" recording whether each instruction had decided anything and what it cost. A calibration or evaluation session ran a frozen design against the jury. An evidence session built an anchor page by reading a published translation whole against its source.

Every session ended the same way: regenerate the indexes, run the checker to exit zero, write or extend the day's journal entry, commit, push, open a pull request, squash-merge it to main, confirm main had advanced, rewrite the baton with the next assignment, append one line to the log, kill any background processes, stop. The entry protocol reserved the last forty per cent of the session's context window for this: "A smaller unit landed cleanly beats a larger one abandoned mid-write."

A.4 The three rebuilds

The record contains three diagnoses of the project by itself, each written after the owner asked for one, and each found that the machinery had been working exactly as designed while the project drifted in a way the machinery could not see.

July 25. Fifteen sessions had built strong verification machinery, written about 105,000 words of markdown, and filed no translations at all; roughly fourteen per cent of the money had gone to producing translation. The reorientation page records the owner's instructions nearly verbatim and the eight amendments that followed: diachronic and intralingual translation entered scope, non-English translation criticism became mandatory, jury calibration was split into a detection gate and a non-blocking peer-discrimination tier, the agent was allowed to translate as a labelled subject, the owner would read the prose without judging it, every session's unit was to have a translation limb and a study limb, the first framework release was deferred, and materials became freely-available-first.

July 26. The baton at that time was a flat priority list rewritten wholesale by each session, so it always led with the freshest thing. Seven consecutive sessions had gone to one measurement thread while the calibration gate "sat unbuilt for eleven sessions with its blocker already cleared," at position seventeen of twenty-two. The owner asked for a hand-off design that would "pursue the various aspects of the project in a balanced and systematic way, striving to complete arms as much as possible but avoiding getting stuck on extended rabbit holes." The response was six tracks, a selection rule feeding the most-starved track, and arms: multi-session units with a declared question, a completion criterion and a step budget, on the principle that "a rabbit hole is an arm that outran its budget without declaring it."

August 1. The owner wrote that the project "might have fallen into a rabbit hole of focusing on a few narrow methodological issues." The reassessment found the arm machinery running as designed and the drift moved to the subject of the questions: twenty-six consecutive sessions whose headline was about the project's own apparatus. It added the "subject rule" — work about the project's own statistics, instruments or raters is a gate inside a session, never the session's unit — and a ladder of deliverables, and put the state pages on a diet.

September 4. The fourth assessment, made when the model driving the sessions was about to change, found the August machinery running as designed too — "every arm inside budget, every track rotating, the subject rule honoured, the questions genuinely about literature — and the project still did not converge on its deliverables." Ninety-nine arms had been constituted in forty-one days; ninety-eight closed inside two sessions; the last thirty-five formed one chain on one family of problems, each naming the next as its successor, and none of it was assembling into the handbook. The hand-off pages had regrown to two megabytes. And the one decision that needed the owner — whether to redesign the failed calibration test — had been parked for 157 sessions behind a "For Tom" block that said, every time, that nothing needed his attention. The rebuild replaced tracks and arms with three deliverable workstreams, ordered steps and finish lines, a fixed rotation, a checker that enforced byte caps, and a problem-indexed handbook; the "family rule" sent every result's successor question into the handbook entry it belonged to instead of into the next session.

A.5 What the framework did well

A.6 What it did badly

A.7 The close

The project was wound up on September 9, 2026, at the owner's request, in a single session that consolidated the remaining handbook families, recorded the state of every workstream, wrote this essay and the accompanying framework, and built the HTML mirror of the record. The close-out page lists what was finished, what was left open, and what a successor would need to know.