Repository path: workshop/experiments/E-20260729-drift-window-verify/design/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260729-drift-window-verify |
| status | frozen |
| created | 2026-07-29 |
| updated | 2026-07-29 |
| senses | style-correspondence, accuracy |
| provisional | true |
| internal-judgment-only | true |
| links | wiki/base/anchors/A-beowulf-ingeld/A-beowulf-ingeld.md, wiki/findings/claims/CL-20260726-drift-window.md, wiki/arms/ARM-evidence-audit.md, workshop/experiments/E-20260729-drift-window-verify/design/sites-frozen.md, workshop/translations/beowulf-parting/R04-v1/translation.md, config/models.md, config/budget.md |
E-20260729-drift-window-verify — second-read the drift window, and test it forward on a held-out passage
Frozen before dispatch. Amendments after the independent pre-run critic pass are marked [A-n] and
dated; nothing else is edited after the first API call.
1. Why this exists
CL-20260726-drift-window is the project's most quantitative close-reading result and its evidence
class is X1b — single-reader counting. Both anchors that carry it (A-beowulf-ingeld,
A-yosano-yomogiu) print a standing warning at the top of the page saying they have never been
second-read. The one verification pass this project has ever run
(RS-20260725-anchor-verification) retracted one claim and corrected three, and
A-beowulf-ingeld §3 records a scoring defect found by re-auditing its own table one session after
it was written — the quotations were all correct and the arithmetic over them was not.
The claim is unusually exposed for three separate reasons, and this design attacks all three:
- The scoring is one reader's. 60 cells of TAKE / REFUSE, by the lead, never checked.
- The independent variable is one reader's. Whether a lemma's drift is total or partial is a judgment, and the entire threshold claim (0/12 against 19/40) is a comparison between those two classes. The anchor's own §6 says "a second reader who classes drift differently would get different numbers" and stops there.
- The site list is one reader's, and it was drawn after reading all four renderings. So is the denominator. The anchor never states an inclusion rule at all.
And it is retrospective: the window was formulated on the same 56 lines it was measured on.
2. Question
Four questions, in decreasing order of how much they cost the claim if they come out badly.
- Q1. Do two independent readers, blind to the anchor, reproduce its 60 TAKE / REFUSE cells?
- Q2. Do they reproduce its drift classification, and does the 0/12-vs-19/40 threshold survive being recomputed under each of their classifications rather than the lead's?
- Q3. Does the window predict forward? On 67 lines of the same poem, in the same edition, by the same three published translators, with the site list frozen from the Old English alone before anything was translated or opened — is acceptance still ~0 at total drift and materially above 0 at partial drift?
- Q4. Is the site census a census? An independent reader given the same 67 lines and the same inclusion rule, blind to the lead's list, enumerates sites; how far apart are the two sets?
3. Materials
Old passage (the anchor's own): Bēowulf ll. 2015–2070a, H&S, with Morris & Wyatt 1895,
Gummere 1909, Kirtlan 1913 and T-beowulf-ingeld-R04-v1, all four stored whole at
wiki/base/anchors/A-beowulf-ingeld/. 15 lemmas × 4 renderings = 60 cells.
New passage (held out): Bēowulf ll. 1800–1866, H&S — dawn at Heorot, Hrunting returned,
Beowulf's speech of thanks, Hrothgar's answer. 67 lines. Disjoint from the old passage; not among the
commonly anthologised episodes; chosen from the Old English alone for realia and reflex density
before any modern rendering of it was opened. Renderings by the same three published translators,
extracted from the same Gutenberg editions, plus T-beowulf-parting-R04-v1.
23 CLEAR lemmas × 4 renderings = 92 cells, plus 14 ARGUABLE lemmas carried but not sent.
The freeze chain, and it is checkable in git rather than asserted:
| commit | what was frozen | what it protects |
|---|---|---|
825039e |
design/sites-frozen.md — the reflex census |
the site list is blind to all four renderings, including the lead's own, which is strictly better than the anchor, whose sites were drawn after reading everything |
46e0c74 |
T-beowulf-parting-R04-v1 and its log |
the lead's rendering is source-only |
| (this commit) | this design, design/items.json, the blinding map |
predictions and failure criteria precede any API call |
Six of the 23 new lemmas are lemmas the anchor also scored (wine, dugan, mōd, þegn,
folc, hyrde). That is a property of the poem's vocabulary, not a design choice, and it makes part
of Q3 a within-lemma replication in different sentences.
4. Procedure
Roles, and no model both produces data and criticises it.
| role | model | why |
|---|---|---|
| raters R1, R2 | P1 openai/gpt-5.6-terra, P3 x-ai/grok-4.5 |
the two panel members with no failure history on note (b) |
| independent census (Q4) | reserve qwen/qwen3.7-max |
must not be a rater: a model that enumerated the sites cannot be a blind scorer of them |
| pre-run critic | P2 google/gemini-3.6-flash |
the only panel member that is a subject in nothing here |
Slugs resolved from config/models.md and logged as provenance. temperature: 0,
reasoning: {"effort":"low"}, worst case priced from max_tokens (note (abc)).
Authorship-stripping, which is not blinding [A-5]. Renderings are presented as A / B / C / D
under a map frozen in items.json, different for the two passages so a rater cannot carry a guess
across calls, and no rater is told that one rendering is the project's own. This is not blinding and
the design no longer calls it that: Morris 1895 and Gummere 1909 are public-domain texts that sit in
these models' training data, and the setup cannot enforce non-recognition. So recognition is
measured instead of assumed — every rater is asked, after scoring, to name any rendering it
recognises and to say on what basis, and the recognition rate is reported. Note (bcq): an instrument
whose confound is never elicited has not been shown to lack it.
Anchoring on the source, not on the lead's target text (note (bdi)). Raters are not given per-site excerpts chosen by the lead. They receive each rendering whole (511–641 words) together with the Old English and the line numbers, and locate the site themselves. This is the direct application of the S052 finding that an anchor set written in the lead's own words misses precisely at the sites where a comparator diverged.
Every judgment must carry a verbatim quotation from the rendering it is about. verify.py checks
every quotation by exact string match against the stored file and any judgment whose quotation does
not attest is discarded before the totals are computed — the discipline
RS-20260725-anchor-verification imposed and A-beowulf-ingeld §3 followed at 180/180.
Call plan — 9 calls.
| # | call | prompt contains | returns |
|---|---|---|---|
| 1 | critic | this design | verdict + findings |
| 2, 3 | drift classification, R1 and R2 | 38 lemmas: OE form, H&S sense, the modern reflex, the Old English line. No renderings. | total / partial / marked / none per lemma + one-line reason |
| 4, 5 | reflex scoring, OLD passage, R1 and R2 | OE 2015–2070a + four renderings whole + the 15 sites + the rule verbatim | TAKE / REFUSE / OMITTED + quotation per cell |
| 6, 7 | reflex scoring, NEW passage, R1 and R2 | OE 1800–1866 + four renderings whole + the 23 sites + the rule verbatim | same |
| 8 | [A-1] presentation-order repeat of call 6 — identical items, identical blinding map, identical temperature, site order permuted by frozen seed 20260729 |
same as 6, reordered | order-invariance of the cell judgments |
| 9 | independent census, qwen |
OE 1800–1866 + the inclusion rule verbatim. No renderings, no site list. | its own list of sites + drift class |
Call 8 is the partial discharge of the S049 backlog row "a rating design that repeats every
condition, not only the treatment". [A-1], and the amendment withdraws a claim rather than adding
one. The design first specified a byte-identical repeat at temperature: 0 and called it
"within-rater reproducibility"; the independent critic pointed out before dispatch that greedy
decoding makes such a repeat a test of backend determinism and nothing else. It is therefore a
presentation-order repeat: the same 23 sites in a permuted order. It measures whether the cell
judgments are invariant to presentation. It does not measure sampling variance, and this project still
has no estimate of that.
The lead scores the new passage too, by hand, and those scores are committed before any rater output is read. They are a comparison, not the datum.
5. Predictions, registered before dispatch
| # | prediction | what it is testing |
|---|---|---|
| P1 | Inter-rater raw agreement on drift class across 38 lemmas ≥ 0.70 | Q2 — is the independent variable reproducible at all |
| P2 | Inter-rater cell agreement on TAKE / REFUSE, all 152 cells, ≥ 0.85 | Q1 — the scoring is close to mechanical and should be near-ceiling |
P1b [A-4] |
Inter-rater agreement on the binary total vs not total split ≥ 0.80 |
Q2 — this, not the 4-class figure, is the boundary the threshold claim rests on |
| P3 | Rater majority agrees with the anchor's stored 60 cells at ≥ 0.85 | Q1 — the second read proper |
P4 [A-2] |
On the OLD passage, recomputed under each rater's own drift classification: a one-sided Fisher exact test of (total vs partial) × (TAKE vs REFUSE) is significant at α = 0.05 in the direction fewer TAKEs at total drift | Q2 — does the threshold survive re-classification |
P5 [A-2] |
On the NEW passage, Morris + Gummere + Kirtlan only (the authorship-stripped published rows): the same Fisher exact test is significant at α = 0.05 in the same direction | Q3 — does the window predict forward |
| P6 | Morris's partial-drift acceptance on the NEW passage exceeds both Gummere's and Kirtlan's | §4.4, archaism as an inheritance technology, tested prospectively |
| P7 | Jaccard overlap between the independent census and the lead's CLEAR ∪ ARGUABLE set is below 0.60 | Q4 — registered as an expected failure of the lead's census. Note (bcl): a criterion is not shown to define a set until a non-author applies it |
P9 [A-5] |
(measured, not assumed) Raters recognise at most one of the four renderings per passage | the confound finding 5 named; reported whatever it shows |
| P8 | (local, no API call) Where the H&S glossary supplies its own English rendering of a site, it refuses the reflex at ≥ 0.80 of them | the priming declared at T-beowulf-parting-R04-v1 log §1 — is the apparatus itself pushing both lead rows down |
P7 is predicted to fail in the direction that costs this design something, and it is registered that way on purpose. A design whose only registered predictions are ones it expects to pass is not carrying risk.
6. Failure criteria — what stops a number being reported
- F1.
[A-4]If inter-rater agreement on the binarytotalvsnot totalsplit is below 0.60, the threshold statistic is not reported under any classification, and the finding is thatCL-20260726-drift-window's independent variable is not reproducible between readers. This is a live and serious outcome and it is named first for that reason. - F2. If inter-rater agreement on TAKE / REFUSE is below 0.75, the re-score is reported as a failure of the scoring rule to define a set, and no corrected table is published.
- F3. Any cell whose supporting quotation does not attest verbatim in the stored file is discarded, not repaired. If more than 10% of cells are discarded for a given rater, that rater's whole return is set aside and the reason recorded.
- F4. The repeat (call 8 against call 6) is reported whatever it shows. If byte-identical inputs move cell agreement by more than the inter-rater gap, the inter-rater gap is not interpretable and is reported as such.
- F5.
[A-3]If the independent census names CLEAR-eligible sites the lead's list missed, every ratio in this result is reported over the frozen site set and explicitly not over "the passage", and the missed sites are listed by name. They are not retro-fitted into the totals. And if P7 fails, P5 and P6 are demoted from primary to statements about the frozen site set — a forward-prediction rate over a site list shown to be one reader's does not support a claim about the passage. A lead-only sensitivity recomputation over the union of both censuses is reported beside them and labelled lead-only. - F6. The lead's own row on the new passage is excluded from P5 and P6 — it is primed twice (site list known before translating; glossary priming, log §1). It is reported beside them.
5b. Power, tabulated before dispatch [A-2]
Note (bci): before freezing a gate, tabulate its power against the effect it exists to catch. One-sided Fisher exact at α = 0.05, exact enumeration over the binomial outcome space.
| arm | cells (total / partial) | true partial acceptance | power |
|---|---|---|---|
| NEW, 3 published rows | 12 / 48 | 0.60 (the rate the anchor reported for these three) | 1.000 |
| NEW, 3 published rows | 12 / 48 | 0.45 | 0.999 |
| NEW, 3 published rows | 12 / 48 | 0.30 | 0.819 |
| NEW, 3 published rows | 12 / 48 | 0.20 | 0.241 |
| NEW, with one stray TAKE at total | 12 / 48 | 0.45 (total 0.08) | 0.781 |
| OLD, four rows | 12 / 40 | 0.475 (the anchor's own 19/40) | 0.997 |
| OLD, four rows | 12 / 40 | 0.25 | 0.416 |
| NEW, if raters keep only 2 total-drift lemmas | 6 / 48 | 0.60 | 0.992 |
So the design can detect the effect the anchor reported and cannot detect a weak version of it. A null at P5 will therefore mean "not the effect the anchor reported" and will not mean "no effect", and the result page must say so.
7. What this design cannot do
- It cannot make
CL-20260726-drift-windowa general claim about translation. It is one poem, one pair, four translators, three of whom published within nineteen years of each other. The anchor's §6 period confound is untouched and this design does not address it. - It cannot establish the site census is complete. Q4 measures the distance between two enumerations; two readers agreeing would not make a third redundant.
- No model here is calibrated. Tier D has not passed, so nothing on this page carries evidential weight as a quality judgment. What it relies on is the S015 instrument note: on factual-adjudication tasks — is this string present in this text — the panel scored 0.886–0.917 discrimination against planted false claims. Does the modern reflex appear in this rendering is a task of that kind and not of the other kind. This distinction is the whole licence for the run and if it is wrong the run is worthless.
- It says nothing about whether any rendering is good. No sense is being scored for quality.
8. Budget
Pre-flight below; actuals recorded in config/budget.md after the run. Worst case is built from
max_tokens at each model's list out-price plus the prompt at its list in-price (note (abc)), and
list prices are not what gets billed (S022 caution) — the routed price can be several times list.
| calls | model | max_tokens |
worst case |
|---|---|---|---|
| 1 critic | P2 gemini-3.6-flash |
8,000 | $0.075 |
| 2 drift | P1, P3 | 6,000 | $0.132 |
| 4 scoring (old + new × 2 raters) | P1, P3 | 12,000 | $0.540 |
| 1 repeat | P1 | 12,000 | $0.210 |
| 1 census | qwen3.7-max |
8,000 | $0.100 |
| total | ≈ $1.06 |
Against a $5.00 UTC-day cap with $5.00 unspent at session start (2026‑07‑29 has no rows). Every
run in this ledger since S043 has landed between 15% and 49% of a max_tokens worst case.
Fall-through: qwen/qwen3.7-max for a rater that returns finish_reason: length with no content
(note (b)) — except on call 9, where qwen is the subject; its reserve is P2, which by then has already
made its critic call and is a subject in nothing else.