Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260801d-name-ground-truth/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260801d-name-ground-truth
statusfrozen
created2026-08-01
updated2026-08-01
senses—
internal-judgment-onlytrue
linkswiki/arms/ARM-figure-audit.md, tools/ngram_overlap.py, wiki/findings/results/RS-20260726c-name-tokens-repair.md, wiki/decisions/resolved/D-20260726-08-name-tokens-exemption-set.md, workshop/translations/njala/R04-v1/translation.md, workshop/canon/njala/manifest.md, config/models.md

E-20260801d — what the name-excluded column actually excludes

Frozen before any rater call. Session S080, track T3, ARM-figure-audit step 3.

1. The question, and why it is this session's

Step 3 of ARM-figure-audit sweeps every published figure computed by tools/ngram_overlap.py or tools/dependence_check.py before the S031 name_tokens repair and recomputes it under the repaired rule, or lists it as un-recomputed with the reason. That sweep is retrospective and mechanical, and it can return exactly one kind of verdict: this figure moves or this figure stands.

A figure that stands under the repaired rule has been shown to be stable, not shown to be right. The arm says so itself, in its constraints declared at birth: "A figure that survives the check is not thereby validated. Reachability and body-selection are necessary conditions, not sufficient ones." The repaired rule has never been checked against a ground truth. RS-20260726c reports, as a finding it could not pursue, that after all three repairs eleven function words are still admitted as proper names across the seven S026 cells and twenty in beowulf-ingeld.

So: of the tokens name_tokens calls proper names, how many are proper names; of the proper names in the text, how many does it call? Nobody has asked. Everything the sweep leaves standing rests on the answer.

2. Why this material

T-njala-R04-v1 — Brennu-Njáls saga chapters 1–2, translated in session under R04 v1.0 — against Dasent 1861. Chapter 1 is very largely genealogy and chapter 2 is a betrothal negotiation at the Althing: twenty-three people and fifteen places in 1,151 source words. The passage was chosen for that density before it was translated (workshop/canon/njala/manifest.md §Why this passage).

The name rule has power to be wrong here in a way it does not on ordinary narrative, and the consequence is measurable on this session's own numbers: the gate below computes 34 shared 7-grams, of which the rule removes 17, leaving 17.

[CORRECTED IN PLACE, post-run] This sentence read "removes 20, leaving 14" when the design was frozen. Those were the unrepaired materials' figures, quoted after the repair had been applied — see the correction in §3.

3. The contamination gate, run and recorded BEFORE this design was written

CLAUDE.md's standing rule: measure before designing, as a selection gate. gate/build_gate.py extracts both English bodies to plain text and runs tools/dependence_check.py unmodified.

pair shared 7-grams 12-grams 15-grams longest run verdict
dasent1861~lead (T-njala-R04-v1) 34 0 0 11 clean
dasent1861~draft (T-njala-R06-v1) 37 0 0 11 clean

The 11-token run is he is here at the thing and his daughter too and — ten function words and the assembly noun, and not one proper name in it, which is not what the translator's log (L01–L03) predicted convergence would look like. The translation's declaration therefore stays at suspected, unchanged from what was written before the measurement, on the precedent of T-son-makara-R04-v1 at 8 tokens. It is not none: nothing here establishes independence, and the project has three measurements against the premise that a canonical comparator gives the longer run.

Two materials defects were found by reading the resolved name list and repaired before this design was frozen; the unrepaired run is preserved at gate/gate-run1-unrepaired.json. Dasent's edition carries a literal ENDNOTES: label and the reference marker (1), which tokenise() folds into the preceding word: the token thorgerda1 was a name occurring once, and the token he was admitted as a proper name because Thorgerda.(1) did not look like a sentence ending — defect A's exact class, arising from the material rather than from the rule, three sessions after the rule was repaired. Stripping the apparatus removed endnotes, thorgerda1 and he from the resolved set (115 → 112).

[CORRECTED IN PLACE, post-run] This paragraph asserted that the repair "moved no overlap figure at all (34/0/0/11 both times)", and that is false. It moved none of the four columns dependence_check prints — shared 7-, 12- and 15-grams and the longest run were 34/0/0/11 before and after — and it moved the name-excluded count from 14 to 17, which that summary table does not print. Removing three spurious names freed three shared 7-grams. The claim was stronger than the evidence it was read off, which is note (bgx)'s exact shape, committed inside the design of the experiment about it, and it is corrected here rather than argued with.

One materials decision is declared rather than hidden: Dasent's endnote TEXT is kept. He renders the whole chapter-1 genealogy as a footnote where the Icelandic has it in the running text. Dropping it would remove the densest proper-name material in the passage, in the direction that flatters the lead.

4. The candidate set, and why it makes recall computable

[A4] A proper name in English is normally capitalised at least once, so the set of token types that are ever capitalised anywhere in the cell is a superset both of what name_tokens can resolve and of the capitalised true name set. It is not a superset of the true name set simpliciter: a name the texts never capitalise — a typo, an OCR slip, a stylistic lowercasing — falls outside it. That gap is the instrument's and not the design's: R1 and R2 both require capitalisation, so such a name is invisible to name_tokens too. Every recall figure below is therefore labelled recall over capitalised proper names, and the one detectable part of the gap — a GT_NAME type that also occurs lowercase somewhere — is counted and reported.

Computed on the two frozen texts: 593 distinct token types, 138 ever capitalised, 112 resolved by name_tokens as names — and resolved ⊆ ever-capitalised is asserted, not assumed (build_items.py).

Classifying all 138 therefore yields precision and recall-over-capitalised-names exactly, with no sampling and no generation task for the raters to be noisy at.

5. Procedure

Raters. Three seats, config/models.md P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5. Declared reserve for every seat: P5 deepseek/deepseek-v4-pro. Non-Anthropic by charter §5. This is a factual-adjudication task, not a quality judgement — the class on which config/models.md's S015 instrument note records the panel performing strongly (20/20 planted false claims rejected, discrimination 0.886–0.917) — so Tier D's failure does not bear on it, and no quality figure about any translation is computed anywhere in this design.

Payload. One call per seat. Both English texts in full, then the 138 candidates as N001–N138, each in a different pseudo-random order per seat (seeded by seat index, seed recorded), each accompanied by up to three quoted occurrences in context taken verbatim from the frozen texts. max_tokens 6,000, temperature 0, attempts capped at 2 per slug (S079's overrun cause).

Response format. One line per item, Nnnn | VERDICT | brief reason, then a literal END terminator. Verdicts:

EITHER exists because the alternative is to force a binary and then read the forced choices as data. It is scored separately and never silently folded into either side.

6. Controls

the repeat's cell-level agreement must be ≥ the mean pairwise cell-level agreement of the three seats. Below that, every three-seat agreement figure on this page is descriptive only — the instrument disagrees with itself more than the seats disagree with each other. Independently, a repeat below an absolute 0.80 means the instrument is not usable for this task and the precision and recall figures are withdrawn, not downgraded.

Both legs are reported with their numbers whether they fire or not. The absolute 0.90 this section carried before the critic pass was looser than its own stated proposition — note (bgx), third instance in three sessions, and the critic found it from the design text alone.

7. Failure criteria, registered

8. Analysis, specified before the run

8.1 Ground truth. GT_NAME = the types with a majority NAME among the surviving seats; GT_NOT = majority NOT; GT_EITHER = majority EITHER or a three-way split. Reported with the counts of each.

8.2 Precision and recall of name_tokens, over the 138 candidates, computed twice — once counting GT_EITHER as NAME (the lenient reading) and once as NOT (the strict one), because the choice is a judgement and hiding it inside one number is the defect this project keeps finding:

precision = |resolved ∩ GT_NAME| / |resolved| · recall = |resolved ∩ GT_NAME| / |GT_NAME|

[A3] and computed twice again, type-level and token-weighted, each type weighted by its occurrence count across the two texts. §1 asks a question about tokens and a type-level figure does not answer it: a name occurring fifty times and a false positive occurring once are one item each. Where the two disagree, the disagreement is the finding.

[A5] The spread between the lenient and the strict reading is itself a registered output. It is the measure of how much the answer depends on an orthographic convention rather than on a fact about the text, and it — not either endpoint — is what a later reader should carry away.

8.3 The consequential arm. Recompute the shared-7-gram counts of §3 with the resolved set replaced by GT_NAME, and report the change in the name-excluded count (17 as computed by the tool on the repaired materials; 14 was the unrepaired figure).

8.4 Attribution. For every false positive, report which rule admitted it — R1 (capitalised at a non-sentence-opening) or R2-only (never occurs lowercase in the cell) — by re-running name_tokens with each rule alone. This is what makes the finding actionable rather than a complaint.

9. Predictions, registered before dispatch

10. Budget

Declared worst case $0.76, built from max_tokens per note (abc) and from the caller's real retry structure per S079 — the estimate is (seats × attempts × slugs), not seats:

line worst case
pre-run critic, qwen/qwen3.7-max, cap 12,000 $0.16
three seats × 2 attempts, payload 11,288 input tokens measured, cap 6,000, P3 priced at twice list for routing (S022 caution) $0.48
byte-identical repeat, one seat × 2 attempts $0.12

Day headroom at session open $2.472159157 (three prior sessions, $2.527840843 of $5.00). The translation limb, the gate, the item build and every analysis are $0.00.

11. What this design cannot show