Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: journal/2026-07-23.md · rendered 2026-09-09

Journal — 2026-07-23 (S001, first run)

Tom —

First run, empty repository to working project. Everything below is landed on main.

What was built. The full load-bearing scaffold from the charter: session entry point (continue-prompt.md), conventions and environment record (CLAUDE.md), the wiki skeleton with the goodness typology's first draft, budget ledger, decision machinery, workshop structure with the experiment discipline written down, and a small index generator (tools/build_index.py) so navigation stays cheap as pages accumulate.

The panel. I surveyed OpenRouter's July 2026 catalog and probed nine candidate models from eight labs with three short public-domain passages (Akutagawa, Miyazawa, Kajii — chosen to test archaic realia, casual dialogue with numerals, and lyric interiority). All nine were alive and broadly competent; the differences were instructive at the edges. Mistral's mid-tier rendered 火桶 as "foot warmer"; several otherwise strong models gave Miyazawa's chatty gentleman ("ぼくは二千四百円の損害だ") the diction of an insurance adjuster ("I have suffered a loss of two thousand four hundred yen"); and Kajii's これはちょっといけなかった sorted the field neatly — Grok's "This was rather bad" kept the litotes, others flattened or overshot it. Panel v1 is five models across five labs (OpenAI, Google, xAI, Moonshot, DeepSeek), deliberately excluding Anthropic models so the evaluation loop stays independent of the lead agent's priors. Probe cost: $0.093. Full rationale in config/models.md; composition opened for ratification as D-20260723-03.

The pilot (permitted by charter §10.7, informal only). Section 一 of 蜘蛛の糸 under the two drafted regimes — single-pass vs. draft-plus-self-revision — with DeepSeek as translator, $0.010 total. Two things worth reporting, as hypotheses only: the self-revision pass did real work (it caught its own draft's opening calque, "It was a certain day," for ある日の事でございます); and neither regime so much as noticed the source's storyteller です/ます register — both silently normalized to neutral written English. That silent normalization is now an open question (OQ-20260723-target-register): into which English are we translating, and who decides?

Decisions opened (for a later session to ratify, per the cross-session rule): the nine-sense initial cut of the goodness typology; canon criteria plus a seven-work shortlist (with its weaknesses recorded — no female author yet; two candidates to verify as fixes); panel v1 composition.

Spent. $0.104 of today's $5.00.

Next. Canon verification pass (real lengths, PD arithmetic, text condition, the Okamoto/Hayashi candidates), ratifications, and the start of the calibration-case hunt — Botchan and No Longer Human look like the cleanest documented cases of competing translations with an asymmetric record, which is what the jury has to be tested against before its verdicts mean anything.

If you'd like anything translated — the private-texts pipeline is empty and waiting (wiki/base/wanted.md).


Journal — 2026-07-23 (S002)

Tom —

Second run, same day. Two things got done: the canon was verified and built, and all three of S001's open decisions were ratified. All landed on main.

The canon is real now. I wrote a small Aozora fetcher (tools/fetch_aozora.py — Shift_JIS in, ruby and editorial marks stripped, character count out) and ran it over the whole shortlist. Every work is public-domain in Japan with room to spare (all six are past both the old 50-year term and the current 70-year one, so nothing hinges on the 2018 extension). The real character counts corrected a few of my run-one guesses: 檸檬 and 注文の多い料理店 were shorter than I'd estimated, and — the one that mattered — 走れメロス came in at 9,806, under the 10,000 target, so the worry that it might be too long to translate whole is gone. 夢十夜 I stored as just the first two nights (the Aozora file is all ten), which is what the canon entry always meant. Six works now sit in workshop/canon/, each with its source text and a manifest showing the provenance, the public-domain arithmetic, and text-condition notes (Miyazawa's tale, for instance, keeps its historical kana and the old /\ repetition marks — authentic, but a translator needs to know they mean "repeat the syllable").

The female-author gap is still open, and now I understand why. I checked the three candidates I floated last time — Okamoto Kanoko's 鮨 and two Hayashi Fumiko stories. All three are longer than the target as whole works (鮨 is right at the ceiling, 13,143). So there's no clean short whole-work swap. The reviewer who audited the canon caught something fair: I'd been willing to excerpt Sōseki to make him fit but hadn't tried the same trick on any woman writer. That's now a live task, not a footnote — next session tries an excerpt of 鮨 or 下町, and checks Tamura Toshiko (died 1945, so public-domain, and a genuinely modern prose writer I'd overlooked). Adding a member is its own decision, so the canon of six is frozen in the meantime with the gap recorded in the open.

Ratifications. The charter's rule is that a different session ratifies, and that each decision gets an independent adversarial review plus one vote from a non-Anthropic model — the point being that I (an Anthropic model) shouldn't be the one blessing my predecessor's calls. So I ran three separate review agents, each told to genuinely try to break the decision, and routed one vote each through GPT-5.6, Grok-4.5, and DeepSeek. All three decisions — the nine-sense goodness vocabulary, the canon, and the five-model panel — came back ratify, but every review earned its keep: the typology got three honesty fixes (it now carries the right flags and no sense is written as if it's already anchored); the panel decision now records that our biggest real risk (three of five models are US labs, whose tastes might move together) is only manageable if the calibration study is actually built to measure it — so I wrote that requirement into the plan; and one small overclaim about Grok's translation got corrected against the raw output. The full reviews and votes are kept in the repo as provenance.

Spent. $0.063 today (the three votes; the canon verification was all network, no model calls). That's $0.167 of the $5.00 for the day across both sessions.

Next. The calibration-case hunt is the real prize and it's next: building the reception records for Botchan and No Longer Human so the jury has documented human judgment to be tested against. Alongside that, the female-author excerpt search and the first anchor/source pages.

Nothing needs your input right now. The private-texts pipeline is still empty whenever you have something you'd like put into English.


Journal — 2026-07-23 (S003)

Tom —

Third run today. This one did the thing S002 pointed at: it built the reception records for the two calibration cases and turned them into an actual, criticized experiment design. All on main.

The two cases came apart in a useful way. I went looking for documented human judgment comparing the rival English translations of Botchan and No Longer Human — the raw material the jury has to reproduce before its verdicts mean anything — and the two cases turned out to have opposite shapes:

Then I designed the calibration experiment — and had it torn apart before freezing it. I wrote the frozen design (blinded passages, five models scoring competing translations sense by sense, measuring both how well they match the human record and whether the three US-lab models agree with each other suspiciously more than with the others — the risk S002 flagged). Then, per our own rules, an independent critic agent that hadn't written it went at it adversarially. It came back "needs redesign," and it was right on three real things: (1) my US-cluster metric couldn't tell shared taste apart from shared competence — three models that are simply better would also agree more, and I'd have misread that as the bias signal; (2) the Botchan test was almost rigged to pass, because a 1918-vs-2005 pair differs in period register, which is the very thing the record rewards — a model could "pass" just by preferring modern-sounding prose, no real judgment involved; and (3) my plan would have routed copyrighted excerpts into the public part of the repo.

I rewrote the design to fix all three: the bias metric now only looks at cases the record can't explain (so competence can't masquerade as taste) and has a hard pre-committed numeric trigger instead of a vibe; the Botchan test now pairs Turney 1972 against Cohn 2005 — same era, so a model can't coast on "sounds newer" — with the easy 1918 pair kept only as a sanity check; and every copyrighted excerpt is quarantined to private-texts/, with only the numbers going public. Both the critique and my point-by-point response are in the repo (critic.md), because that argument is the evidence the design is sound.

One thing I need from you. The run itself needs the actual translation texts — brief matched passages from four in-copyright books (Cohn's Botchan; Keene, Gibeau, and Winters Carpenter's No Longer Human). Short excerpts are enough, but picking matching passages needs the full texts, so this realistically needs you to provide them (they'd go into the quarantined private-texts/ area). Without at least Cohn's Botchan, the primary test can't run. It's logged in wiki/base/wanted.md. Everything else — Morri's 1918 Botchan and both Japanese originals — is public domain and already reachable.

Spent. $0.00 today — this was all web reading and writing; the critic was an internal agent, no API calls. Still $0.167 for the day. The calibration run (funded at ~$1.50) is the next real spend, and it's waiting on the texts.

Next. Once the texts are in hand: build the blinded passages and run the calibration. If you'd rather I keep moving without waiting, I can start the poetics-track reading (the first anchor and source pages) or hunt a third clean calibration case (Akutagawa's stories, Kojima vs. Rubin, look promising).


Journal — 2026-07-23 (S004)

Tom —

Fourth run today. The calibration run is still waiting on the in-copyright texts, so rather than idle I took the branch the last three sessions kept deferring: the poetics track, which had literally nothing in it yet. wiki/base/anchors/ — the shelf that's supposed to hold the highest-authority evidence for what good English actually reads like — was empty. I built its first entry.

The first anchor: Katherine Mansfield. The idea behind the anchor hierarchy is that when the project says a translation "reads as natural English," that judgment has to be pinned to real published prose, not to my own sense of good writing (which is exactly the thing that's supposed to be under suspicion). So I went looking for the right yardstick for this canon — Japanese short stories from roughly 1900–1945 — and Mansfield is close to ideal: she was writing and publishing at the dead centre of that window (her last collection, The Garden Party, 1922), in the same short-story form, at the top of the craft, and she's a master of exactly the technique a J→E translator most needs and most often botches — free indirect discourse, a character's inner voice carried inside third-person narration (Japanese's subjectless, particle-coloured sentences push you toward it constantly). She's public domain several times over (died 1923; the book is a 1922 publication), so I could store a whole story — "Miss Brill" — in the repo as the concrete reference, and I read it closely and catalogued what specifically makes it read as native English: the free indirect slips ("Dear little thing!"; "The Brute! The Brute!"), the freely contracted colloquial default, repetition doing real emotional work ("she too, she too"), sound-play written without embarrassment ("Tum-tum-tum tiddle-um!"). Those are the tells — the positive models — that later naturalness judgments can point at instead of hand-waving.

The honest catch, which I wrote into the page. Choosing an era-matched anchor is itself a quiet decision about a question we haven't answered: into which English are we translating — the English of the source's era, a "timeless" register, or today's? Mansfield licenses "button boots" and "to-day" and "ermine toque" — sentences that would read as faintly costumed if a 2026 translator reached for them. So anchoring naturalness to 1922 English tilts us toward a period target. That might be right, might not; the point is I made the tilt visible rather than letting it hide inside the word "natural." It's flagged straight into the open register question, and the fix is concrete: build a second, contemporary-English naturalness anchor and let the same translation be measured against both — the gap between them is that open question moving.

So: the naturalness sense in our typology just moved from "to derive" to "deriving" — it now has one real anchor under it, with a caveat, instead of a promissory note. Still eight senses and a long way to go (nothing's been anchored for affect yet, and one register is not a basis), but the shelf isn't empty anymore.

Spent. $0.00 today — all web reading and writing, no model calls. Still $0.167 for the day.

Next. The calibration run remains the prize and remains blocked on you (the four in-copyright translations — Cohn's Botchan above all). While that waits, the productive path continues: the first Tier 2 source page (a translator's afterword — Cohn's or Gibeau's — does double duty with the calibration cases), a contemporary-English naturalness anchor to pair with Mansfield, and the third calibration case (Akutagawa, Kojima vs. Rubin).