Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260729-drift-window-verify/design/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260729-drift-window-verify
statusfrozen
created2026-07-29
updated2026-07-29
sensesstyle-correspondence, accuracy
provisionaltrue
internal-judgment-onlytrue
linkswiki/base/anchors/A-beowulf-ingeld/A-beowulf-ingeld.md, wiki/findings/claims/CL-20260726-drift-window.md, wiki/arms/ARM-evidence-audit.md, workshop/experiments/E-20260729-drift-window-verify/design/sites-frozen.md, workshop/translations/beowulf-parting/R04-v1/translation.md, config/models.md, config/budget.md

E-20260729-drift-window-verify — second-read the drift window, and test it forward on a held-out passage

Frozen before dispatch. Amendments after the independent pre-run critic pass are marked [A-n] and dated; nothing else is edited after the first API call.

1. Why this exists

CL-20260726-drift-window is the project's most quantitative close-reading result and its evidence class is X1b — single-reader counting. Both anchors that carry it (A-beowulf-ingeld, A-yosano-yomogiu) print a standing warning at the top of the page saying they have never been second-read. The one verification pass this project has ever run (RS-20260725-anchor-verification) retracted one claim and corrected three, and A-beowulf-ingeld §3 records a scoring defect found by re-auditing its own table one session after it was written — the quotations were all correct and the arithmetic over them was not.

The claim is unusually exposed for three separate reasons, and this design attacks all three:

  1. The scoring is one reader's. 60 cells of TAKE / REFUSE, by the lead, never checked.
  2. The independent variable is one reader's. Whether a lemma's drift is total or partial is a judgment, and the entire threshold claim (0/12 against 19/40) is a comparison between those two classes. The anchor's own §6 says "a second reader who classes drift differently would get different numbers" and stops there.
  3. The site list is one reader's, and it was drawn after reading all four renderings. So is the denominator. The anchor never states an inclusion rule at all.

And it is retrospective: the window was formulated on the same 56 lines it was measured on.

2. Question

Four questions, in decreasing order of how much they cost the claim if they come out badly.

3. Materials

Old passage (the anchor's own): Bēowulf ll. 2015–2070a, H&S, with Morris & Wyatt 1895, Gummere 1909, Kirtlan 1913 and T-beowulf-ingeld-R04-v1, all four stored whole at wiki/base/anchors/A-beowulf-ingeld/. 15 lemmas × 4 renderings = 60 cells.

New passage (held out): Bēowulf ll. 1800–1866, H&S — dawn at Heorot, Hrunting returned, Beowulf's speech of thanks, Hrothgar's answer. 67 lines. Disjoint from the old passage; not among the commonly anthologised episodes; chosen from the Old English alone for realia and reflex density before any modern rendering of it was opened. Renderings by the same three published translators, extracted from the same Gutenberg editions, plus T-beowulf-parting-R04-v1. 23 CLEAR lemmas × 4 renderings = 92 cells, plus 14 ARGUABLE lemmas carried but not sent.

The freeze chain, and it is checkable in git rather than asserted:

commit what was frozen what it protects
825039e design/sites-frozen.md — the reflex census the site list is blind to all four renderings, including the lead's own, which is strictly better than the anchor, whose sites were drawn after reading everything
46e0c74 T-beowulf-parting-R04-v1 and its log the lead's rendering is source-only
(this commit) this design, design/items.json, the blinding map predictions and failure criteria precede any API call

Six of the 23 new lemmas are lemmas the anchor also scored (wine, dugan, mōd, þegn, folc, hyrde). That is a property of the poem's vocabulary, not a design choice, and it makes part of Q3 a within-lemma replication in different sentences.

4. Procedure

Roles, and no model both produces data and criticises it.

role model why
raters R1, R2 P1 openai/gpt-5.6-terra, P3 x-ai/grok-4.5 the two panel members with no failure history on note (b)
independent census (Q4) reserve qwen/qwen3.7-max must not be a rater: a model that enumerated the sites cannot be a blind scorer of them
pre-run critic P2 google/gemini-3.6-flash the only panel member that is a subject in nothing here

Slugs resolved from config/models.md and logged as provenance. temperature: 0, reasoning: {"effort":"low"}, worst case priced from max_tokens (note (abc)).

Authorship-stripping, which is not blinding [A-5]. Renderings are presented as A / B / C / D under a map frozen in items.json, different for the two passages so a rater cannot carry a guess across calls, and no rater is told that one rendering is the project's own. This is not blinding and the design no longer calls it that: Morris 1895 and Gummere 1909 are public-domain texts that sit in these models' training data, and the setup cannot enforce non-recognition. So recognition is measured instead of assumed — every rater is asked, after scoring, to name any rendering it recognises and to say on what basis, and the recognition rate is reported. Note (bcq): an instrument whose confound is never elicited has not been shown to lack it.

Anchoring on the source, not on the lead's target text (note (bdi)). Raters are not given per-site excerpts chosen by the lead. They receive each rendering whole (511–641 words) together with the Old English and the line numbers, and locate the site themselves. This is the direct application of the S052 finding that an anchor set written in the lead's own words misses precisely at the sites where a comparator diverged.

Every judgment must carry a verbatim quotation from the rendering it is about. verify.py checks every quotation by exact string match against the stored file and any judgment whose quotation does not attest is discarded before the totals are computed — the discipline RS-20260725-anchor-verification imposed and A-beowulf-ingeld §3 followed at 180/180.

Call plan — 9 calls.

# call prompt contains returns
1 critic this design verdict + findings
2, 3 drift classification, R1 and R2 38 lemmas: OE form, H&S sense, the modern reflex, the Old English line. No renderings. total / partial / marked / none per lemma + one-line reason
4, 5 reflex scoring, OLD passage, R1 and R2 OE 2015–2070a + four renderings whole + the 15 sites + the rule verbatim TAKE / REFUSE / OMITTED + quotation per cell
6, 7 reflex scoring, NEW passage, R1 and R2 OE 1800–1866 + four renderings whole + the 23 sites + the rule verbatim same
8 [A-1] presentation-order repeat of call 6 — identical items, identical blinding map, identical temperature, site order permuted by frozen seed 20260729 same as 6, reordered order-invariance of the cell judgments
9 independent census, qwen OE 1800–1866 + the inclusion rule verbatim. No renderings, no site list. its own list of sites + drift class

Call 8 is the partial discharge of the S049 backlog row "a rating design that repeats every condition, not only the treatment". [A-1], and the amendment withdraws a claim rather than adding one. The design first specified a byte-identical repeat at temperature: 0 and called it "within-rater reproducibility"; the independent critic pointed out before dispatch that greedy decoding makes such a repeat a test of backend determinism and nothing else. It is therefore a presentation-order repeat: the same 23 sites in a permuted order. It measures whether the cell judgments are invariant to presentation. It does not measure sampling variance, and this project still has no estimate of that.

The lead scores the new passage too, by hand, and those scores are committed before any rater output is read. They are a comparison, not the datum.

5. Predictions, registered before dispatch

# prediction what it is testing
P1 Inter-rater raw agreement on drift class across 38 lemmas ≥ 0.70 Q2 — is the independent variable reproducible at all
P2 Inter-rater cell agreement on TAKE / REFUSE, all 152 cells, ≥ 0.85 Q1 — the scoring is close to mechanical and should be near-ceiling
P1b [A-4] Inter-rater agreement on the binary total vs not total split ≥ 0.80 Q2 — this, not the 4-class figure, is the boundary the threshold claim rests on
P3 Rater majority agrees with the anchor's stored 60 cells at ≥ 0.85 Q1 — the second read proper
P4 [A-2] On the OLD passage, recomputed under each rater's own drift classification: a one-sided Fisher exact test of (total vs partial) × (TAKE vs REFUSE) is significant at α = 0.05 in the direction fewer TAKEs at total drift Q2 — does the threshold survive re-classification
P5 [A-2] On the NEW passage, Morris + Gummere + Kirtlan only (the authorship-stripped published rows): the same Fisher exact test is significant at α = 0.05 in the same direction Q3 — does the window predict forward
P6 Morris's partial-drift acceptance on the NEW passage exceeds both Gummere's and Kirtlan's §4.4, archaism as an inheritance technology, tested prospectively
P7 Jaccard overlap between the independent census and the lead's CLEAR ∪ ARGUABLE set is below 0.60 Q4 — registered as an expected failure of the lead's census. Note (bcl): a criterion is not shown to define a set until a non-author applies it
P9 [A-5] (measured, not assumed) Raters recognise at most one of the four renderings per passage the confound finding 5 named; reported whatever it shows
P8 (local, no API call) Where the H&S glossary supplies its own English rendering of a site, it refuses the reflex at ≥ 0.80 of them the priming declared at T-beowulf-parting-R04-v1 log §1 — is the apparatus itself pushing both lead rows down

P7 is predicted to fail in the direction that costs this design something, and it is registered that way on purpose. A design whose only registered predictions are ones it expects to pass is not carrying risk.

6. Failure criteria — what stops a number being reported

5b. Power, tabulated before dispatch [A-2]

Note (bci): before freezing a gate, tabulate its power against the effect it exists to catch. One-sided Fisher exact at α = 0.05, exact enumeration over the binomial outcome space.

arm cells (total / partial) true partial acceptance power
NEW, 3 published rows 12 / 48 0.60 (the rate the anchor reported for these three) 1.000
NEW, 3 published rows 12 / 48 0.45 0.999
NEW, 3 published rows 12 / 48 0.30 0.819
NEW, 3 published rows 12 / 48 0.20 0.241
NEW, with one stray TAKE at total 12 / 48 0.45 (total 0.08) 0.781
OLD, four rows 12 / 40 0.475 (the anchor's own 19/40) 0.997
OLD, four rows 12 / 40 0.25 0.416
NEW, if raters keep only 2 total-drift lemmas 6 / 48 0.60 0.992

So the design can detect the effect the anchor reported and cannot detect a weak version of it. A null at P5 will therefore mean "not the effect the anchor reported" and will not mean "no effect", and the result page must say so.

7. What this design cannot do

8. Budget

Pre-flight below; actuals recorded in config/budget.md after the run. Worst case is built from max_tokens at each model's list out-price plus the prompt at its list in-price (note (abc)), and list prices are not what gets billed (S022 caution) — the routed price can be several times list.

calls model max_tokens worst case
1 critic P2 gemini-3.6-flash 8,000 $0.075
2 drift P1, P3 6,000 $0.132
4 scoring (old + new × 2 raters) P1, P3 12,000 $0.540
1 repeat P1 12,000 $0.210
1 census qwen3.7-max 8,000 $0.100
total ≈ $1.06

Against a $5.00 UTC-day cap with $5.00 unspent at session start (2026‑07‑29 has no rows). Every run in this ledger since S043 has landed between 15% and 49% of a max_tokens worst case. Fall-through: qwen/qwen3.7-max for a rater that returns finish_reason: length with no content (note (b)) — except on call 9, where qwen is the subject; its reserve is P2, which by then has already made its critic call and is a subject in nothing else.