Repository path: workshop/experiments/E-20260729b-graded-drift/design/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260729b-graded-drift |
| status | frozen |
| created | 2026-07-29 |
| updated | 2026-07-29 |
| senses | accuracy, style-correspondence |
| internal-judgment-only | true |
| links | wiki/arms/ARM-graded-typology.md, wiki/findings/claims/CL-20260726-drift-window.md, wiki/findings/results/RS-20260729-drift-window-verify.md, workshop/experiments/E-20260729-drift-window-verify/design/design.md, workshop/translations/alfred-preface/R04-v1/translation.md, config/models.md, wiki/method-notes.md |
E-20260729b-graded-drift — is the drift window two classes, or a boundary drawn on a continuum?
ARM-graded-typology step 1. Frozen before any lemma was scored, before the Alfred census was
written, and before a word of the translation limb was drafted. Commit hash of this freeze is recorded
in the arm log.
1. The question
CL-20260726-drift-window is a contrast between two classes of semantic drift: total (the modern
reflex has no overlap with the Old English sense) and partial (it overlaps wrongly — a false friend).
RS-20260729-drift-window-verify §4 measured the classes and they do not reproduce: two capable
readers, given the same written decision tree and no translations, agree on 0.654 of 52 lemmas
four-way, and the published contrast is absent under one of the two classifications (p = 0.373
against 0.0093). The disagreements are not carelessness — both readers give nearly the same sentence
for opposite classes (hyrde → herd: "a herd is animals, not a keeper", filed once as total and
once as partial).
Two explanations are live and this project has not separated them.
- (A) The boundary is real and the readers are noisy. There are two kinds of drift; the decision tree is under-specified; a better rule would recover the classes.
- (B) The boundary is imposed. Semantic distance is continuous; "no overlap at all" is not a line two readers can be relied on to draw in the same place, because it depends on how coarse a sense is. The window is a two-class model laid over a slope.
A graded measure needs no boundary at all, and nothing in this project has ever tried one.
And a second hypothesis this design commits to before seeing data: the four-class scheme
(total/partial/marked/none) conflates two axes. marked is not a degree of semantic drift;
it is a fact about register. duguð → doughty drew total from one S053 reader
("doughty does not mean retainers") and marked from the other ("doughty still means valiant but
is archaic"), and both are right, because they answered different questions. If that is what is
happening, a two-number instrument — semantic retention and currency of the modern form — should be
more reproducible than a one-dimensional four-class label, and should explain which S053
disagreements happened.
2. What is measured
For every lemma, three independent raters supply two integers and a reason:
- S — semantic retention, 0–100. How much of the Old English word's meaning is still carried by the modern English form? 0 = none of it; 100 = the modern word means what the Old English word meant.
- C — currency, 0–100. How ordinary is the modern form in present-day written English? 0 = obsolete or found only in historical/poetic usage; 100 = everyday modern English.
No translations are shown for the S/C task. Items are presented exactly as E-20260729's
classification task presented them (Old English lemma, the 1893 dictionary sense, the modern reflex,
the line reference), with Beowulf and Alfred lemmas shuffled together and no source text named, so
no text-level cue can operate.
3. Materials
3.1 Set B — the 52 frozen Beowulf lemmas (existing)
../../E-20260729-drift-window-verify/design/items.json, unchanged. Carries, already on disk:
- the lead's four-class label for all 52;
- R1 (
openai/gpt-5.6-terra) and R2 (x-ai/grok-4.5) four-class labels for all 52; - 152 TAKE/REFUSE cells — 38 lemmas × 4 renderings (Morris 1895, Gummere 1909, Kirtlan 1913, lead 2026), every cell carrying an attesting quotation verified verbatim (396/396 at S053).
3.2 Set A — Alfred's Preface to the Pastoral Care (new)
King Alfred's prose preface to the Old English Cura Pastoralis, c. 890. 874 words of Old English
prose, ninth century, non-fiction, no verse constraint — every one of which the Beowulf material is
not. RS-20260729 §12's first stated limit is "it does not test the window outside Beowulf"; this is
the cheapest available attack on it, and the text is additionally the founding statement of translation
method in English (hwīlum word be worde, hwīlum andgit of andgiete), which is why the arm sits on T2.
- Source text:
materials/alfred-preface-oe.txt— Bright's Anglo-Saxon Reader (public domain) via Wikisource, collated against Sweet 1871 (EETS 45), Hatton 20 text, with three emendations recorded inmaterials/collation.md. Sweet's Old English pages only were read; see §8. - Reflex census:
design/sites-alfred.md, written from the Old English alone, with the inclusion rule stated verbatim, frozen and committed before the lead's S/C scores and before translating. - Lead's own S/C scores:
design/lead-scores-alfred.json, assigned from the Old English and the dictionary alone, frozen and committed before a word was translated. Declared non-independent. - Renderings: the lead's
T-alfred-preface-R04-v1(this session) and Sweet 1871's facing English translation, extracted after the lead's translation and log were frozen and committed.
4. Procedure, in order, each step committed before the next begins
- This design frozen and committed.
- Independent pre-run critic (P2
google/gemini-3.6-flash), dispositions written, design amended only by addition in a dated## Amendmentssection. - Alfred census written from the Old English alone; committed.
- Lead S/C scores for the Alfred sites; committed.
- R06 draft of the whole preface; committed. Contamination gate run on ¶1 against Sweet 1871
before ¶2–¶4 are revised (
CLAUDE.md's standing selection-gate rule). - R04 self-revision and the translator's log; committed and frozen.
- Rater calls dispatched (§5).
- Analysis, then an independent verifier that imports nothing from the analysis.
5. Calls
| # | role | model | task |
|---|---|---|---|
| 1 | pre-run critic | P2 google/gemini-3.6-flash |
this design |
| 2 | G-R1 | P1 openai/gpt-5.6-terra |
S/C on all lemmas, shuffled |
| 3 | G-R2 | P3 x-ai/grok-4.5 |
S/C on all lemmas, shuffled |
| 4 | G-R3 | P5 deepseek/deepseek-v4-pro |
S/C on all lemmas, shuffled |
| 5 | K-R3 | P5 deepseek/deepseek-v4-pro |
the S053 four-class task, decision tree verbatim, Set B only — a third categorical reader |
| 6 | T-R1 | P1 openai/gpt-5.6-terra |
TAKE/REFUSE on the Alfred sites × 2 renderings, attesting quotations |
| 7 | T-R2 | P3 x-ai/grok-4.5 |
same |
Call 5 exists because S053's 0.654 rests on two readers. A third reader answering the identical question is the cheapest test of whether that figure is a property of the scheme or of one pairing.
temperature: 0. A byte-identical repeat measures backend determinism, not judgment
(RS-20260729 §2, note (bdi)); no such repeat is run and no sampling-variance claim is made.
6. Registered predictions
Set B = the 52 Beowulf lemmas. S̄, C̄ = mean across the three graded raters.
| # | prediction | statistic | threshold |
|---|---|---|---|
| G1 | the graded scale is more reproducible between readers than the four-class label | mean pairwise Kendall τ-b on S across the 3 raters, against mean pairwise Cohen κ on the four-class labels (R1, R2, and the new R3) | τ̄(S) > κ̄(4-class) |
| G1b | and the same holds for C | τ̄(C) > κ̄(4-class) | — |
| G2 | the S053 four-class disagreements sit in the middle of the graded scale | Mann–Whitney U on |S̄ − 50|, the 18 R1/R2-disagreed lemmas against the 34 agreed, one-sided | p < 0.05, disagreed lower |
| G2b | disagreements that involve marked are separated by C, not by S |
descriptive: mean S̄ and C̄ for marked-involving vs semantic-only disagreements | exploratory, no threshold |
| G3 | the crux. Once the graded score is known, the class label adds nothing | logistic TAKE ~ translator + S̄, against TAKE ~ translator + S̄ + 1[lead class = total]; likelihood-ratio test on the class term, 152 cells | LR p > 0.05 and ΔAIC < 2 ⇒ the boundary is not carrying information the slope does not |
| G3b | and the graded model is not worse than the class model | AIC(translator + S̄) vs AIC(translator + 1[total]) | AIC(S̄) ≤ AIC(class) |
| G4 | take-rate rises monotonically with retention | Spearman ρ between S̄ and lemma-level take fraction, 38 lemmas | ρ > 0, p < 0.05 |
| G5 | the graded scale transfers to prose 900 years older than the comparison set | τ̄(S) on Set A ≥ 0.50 | — |
| G6 | the lead's frozen pre-translation S predicts the lead's own choices | Spearman between the lead's frozen S and its own TAKE on Set A | ρ > 0 — and this row is non-blind by construction and is excluded from every pooled statistic (F4) |
| G7 | the lead's frozen S agrees with the independent raters | Spearman between lead frozen S̄ and rater S̄ on Set A | ρ ≥ 0.60 |
What each outcome of G3 means, written before the data exists so it cannot be rewritten after:
- G3 passes (class term adds nothing): explanation (B). The window is a boundary drawn on a
continuum.
CL-20260726-drift-windowshould be restated on a graded predictor, and the project's habit of two-class typologies takes a hit that generalises past Beowulf. - G3 fails (class term significant beyond S̄): explanation (A). There is structure at the boundary that the slope misses; the two-class model earns its keep and the S053 reproducibility failure is a specification problem, not a modelling one.
- G1 fails as well as G3 — both instruments unreliable: the honest result is a double null, and the claim's evidence class falls further rather than being repaired.
7. Failure criteria (fired as registered, not renegotiated)
- F1. If mean pairwise τ(S) < 0.40 on Set B, the graded measure is not shareable either. G3 is still computed but no model comparison is reportable as evidence about the world, only about the instruments.
- F2. Any rater returning fewer than 90% of the requested ids, or values outside 0–100, is re-run once and then dropped, with the drop reported.
- F3. Every TAKE/REFUSE cell must carry an attesting quotation that matches the stored rendering by literal string search after whitespace normalisation. Cells failing attestation are discarded, not repaired, and the discard rate is reported (S053's procedure; 396/396 there).
- F4. The lead's Alfred row is not blind — the lead wrote the census and its own S/C scores before translating. It is excluded from every statistic pooling translators, and appears only in G6.
- F5. The Alfred census is the lead's, and
RS-20260729§7 measured the lead's Beowulf census against an independent one at Jaccard 0.4375. No independent Alfred census is run here; every Set A figure is therefore a statement about the frozen site set, not about the preface, and must be worded that way. - F6. If the Sweet extraction cannot be aligned cell-for-cell to the census sites, the Set A take arm is dropped whole and the drop is reported. G5 and G7 do not depend on it.
8. Declared primings and contaminations
- Sweet 1871 is a facing-page edition, so its English sits beside its Old English. The Old English
pages were read for collation (§3.2). One clause of Sweet's English was seen incidentally while
locating the passage by
grep: "which is called in Latin Pastoralis, and in English Shepherd's Book,". It is declared here, before translating, and again on the artifact. BeowulfSet B take cells are S053's, produced by two of the three raters used here. The graded task is a different question on the same items by the same models; this is not independent replication of the take cells and nothing here re-tests them.- Contamination on the Alfred limb is measured, not asserted — one longest-common-run call against
Sweet 1871 after ¶1 is drafted and before the rest is revised, per
CLAUDE.md's standing rule. - The critic model (P2) is a subject in nothing here, unlike at S053 where the census fell through
to it. If any rater call fails twice the reserve is P4
moonshotai/kimi-k3, not P2.
9. Pre-flight cost
Worst case is built from max_tokens, not from expected output (note (abc)).
| call | max_tokens | worst case at list out-price + prompt in-price |
|---|---|---|
| critic | 8,000 | $0.075 |
| G-R1 | 8,000 | $0.135 |
| G-R2 | 8,000 | $0.060 |
| G-R3 | 8,000 | $0.015 |
| K-R3 | 4,000 | $0.010 |
| T-R1 | 12,000 | $0.195 |
| T-R2 | 12,000 | $0.085 |
| total | $0.575 |
Today's headroom before this run: $4.680262. The run fits at 12% of the cap even at worst case.
Amendments (2026-07-29, after the pre-run critic, before any lemma was scored)
critic/dispositions.md carries the reasoning. Only additions; nothing above is deleted or reworded.
A1 — G1/G1b restated on one metric. Krippendorff's α, ordinal for S and C, nominal for the
four-class labels, over the same 52 lemmas; plus a like-for-like nominal comparison, α on S binarised at
each rater's own median against α on total vs not-total. Threshold: α_ord(S) > α_nom(4-class) and
α(S binarised) > α(total vs not). The result states that no ordinal-vs-nominal comparison is exactly
like-for-like.
A2 — the crux moves. G3 is replaced by a direct test of latent structure in S. The critic
established that the old G3 could not fail informatively: p > 0.05 on a class term arises both when
the boundary is imposed and when the boundary is real but the label is noisy. New G3:
- G3-dip. Hartigan's dip statistic on the 52 pooled S̄ values, p by uniform-null bootstrap, 2,000 draws. Registered: dip p > 0.10 ⇒ no departure from unimodality ⇒ no gap ⇒ explanation (B).
- G3-mix. One- against two-component Gaussian mixture by EM. Registered: ΔBIC in favour of one component, and a parametric-bootstrap LR p > 0.05 (500 draws from the fitted one-component null; the asymptotic χ² is invalid here and is not used).
- If the two disagree, both are reported and neither is called the answer.
- The old prediction survives as G3c, secondary, carrying the critic's objection in the text.
A3 — clustering. All inference on the 152 behaviour cells is by lemma-level bootstrap: resample the 38 lemmas with replacement, 2,000 replicates. Applies to G3b, G3c and G4.
A4 — the predictor is a holdout model. Primary S for every behaviour analysis is R3
deepseek/deepseek-v4-pro alone, which supplied nothing at S053. The three-rater mean is secondary
and is labelled confounded wherever it appears.
A5 — S and C in separate calls. Six graded calls, one axis each, no shared context.
A6 — Set A. G6 and G7 become exploratory and are not confirmatory evidence. G5 stays
confirmatory — it is a reliability figure among three independent raters and does not pass through
the lead. The two rater calls scoring TAKE/REFUSE on the Alfred renderings are cut; Set A take data
is produced mechanically by literal string match under a rule frozen in design/sites-alfred.md,
for the lead's rendering and for Sweet 1871.
A7 — revised call list. critic · G-R1-S · G-R1-C · G-R2-S · G-R2-C · G-R3-S · G-R3-C · K-R3. Eight calls. Worst case $0.665; today's headroom before the run $4.680262.