Repository path: workshop/experiments/E-20260801d-name-ground-truth/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260801d-name-ground-truth |
| status | frozen |
| created | 2026-08-01 |
| updated | 2026-08-01 |
| senses | — |
| internal-judgment-only | true |
| links | wiki/arms/ARM-figure-audit.md, tools/ngram_overlap.py, wiki/findings/results/RS-20260726c-name-tokens-repair.md, wiki/decisions/resolved/D-20260726-08-name-tokens-exemption-set.md, workshop/translations/njala/R04-v1/translation.md, workshop/canon/njala/manifest.md, config/models.md |
E-20260801d — what the name-excluded column actually excludes
Frozen before any rater call. Session S080, track T3, ARM-figure-audit step 3.
1. The question, and why it is this session's
Step 3 of ARM-figure-audit sweeps every published figure computed by
tools/ngram_overlap.py or tools/dependence_check.py before the S031 name_tokens repair and
recomputes it under the repaired rule, or lists it as un-recomputed with the reason. That sweep is
retrospective and mechanical, and it can return exactly one kind of verdict: this figure moves or
this figure stands.
A figure that stands under the repaired rule has been shown to be stable, not shown to be right.
The arm says so itself, in its constraints declared at birth: "A figure that survives the check is
not thereby validated. Reachability and body-selection are necessary conditions, not sufficient
ones." The repaired rule has never been checked against a ground truth. RS-20260726c reports, as
a finding it could not pursue, that after all three repairs eleven function words are still
admitted as proper names across the seven S026 cells and twenty in beowulf-ingeld.
So: of the tokens name_tokens calls proper names, how many are proper names; of the proper names
in the text, how many does it call? Nobody has asked. Everything the sweep leaves standing rests
on the answer.
2. Why this material
T-njala-R04-v1 — Brennu-Njáls saga chapters 1–2, translated in session under R04 v1.0 — against
Dasent 1861. Chapter 1 is very largely genealogy and chapter 2 is a betrothal negotiation at the
Althing: twenty-three people and fifteen places in 1,151 source words. The passage was chosen
for that density before it was translated (workshop/canon/njala/manifest.md §Why this passage).
The name rule has power to be wrong here in a way it does not on ordinary narrative, and the consequence is measurable on this session's own numbers: the gate below computes 34 shared 7-grams, of which the rule removes 17, leaving 17.
[CORRECTED IN PLACE, post-run] This sentence read "removes 20, leaving 14" when the design was frozen. Those were the unrepaired materials' figures, quoted after the repair had been applied — see the correction in §3.
3. The contamination gate, run and recorded BEFORE this design was written
CLAUDE.md's standing rule: measure before designing, as a selection gate. gate/build_gate.py
extracts both English bodies to plain text and runs tools/dependence_check.py unmodified.
| pair | shared 7-grams | 12-grams | 15-grams | longest run | verdict |
|---|---|---|---|---|---|
dasent1861~lead (T-njala-R04-v1) |
34 | 0 | 0 | 11 | clean |
dasent1861~draft (T-njala-R06-v1) |
37 | 0 | 0 | 11 | clean |
The 11-token run is he is here at the thing and his daughter too and — ten function words and
the assembly noun, and not one proper name in it, which is not what the translator's log
(L01–L03) predicted convergence would look like. The translation's declaration therefore stays
at suspected, unchanged from what was written before the measurement, on the precedent of
T-son-makara-R04-v1 at 8 tokens. It is not none: nothing here establishes independence, and the
project has three measurements against the premise that a canonical comparator gives the longer run.
Two materials defects were found by reading the resolved name list and repaired before this design
was frozen; the unrepaired run is preserved at gate/gate-run1-unrepaired.json. Dasent's edition
carries a literal ENDNOTES: label and the reference marker (1), which tokenise() folds into the
preceding word: the token thorgerda1 was a name occurring once, and the token he was admitted as
a proper name because Thorgerda.(1) did not look like a sentence ending — defect A's exact
class, arising from the material rather than from the rule, three sessions after the rule was
repaired. Stripping the apparatus removed endnotes, thorgerda1 and he from the resolved set
(115 → 112).
[CORRECTED IN PLACE, post-run] This paragraph asserted that the repair "moved no overlap figure at
all (34/0/0/11 both times)", and that is false. It moved none of the four columns
dependence_check prints — shared 7-, 12- and 15-grams and the longest run were 34/0/0/11 before
and after — and it moved the name-excluded count from 14 to 17, which that summary table does
not print. Removing three spurious names freed three shared 7-grams. The claim was stronger than
the evidence it was read off, which is note (bgx)'s exact shape, committed inside the design
of the experiment about it, and it is corrected here rather than argued with.
One materials decision is declared rather than hidden: Dasent's endnote TEXT is kept. He renders the whole chapter-1 genealogy as a footnote where the Icelandic has it in the running text. Dropping it would remove the densest proper-name material in the passage, in the direction that flatters the lead.
4. The candidate set, and why it makes recall computable
[A4] A proper name in English is normally capitalised at least once, so the set of token types
that are ever capitalised anywhere in the cell is a superset both of what name_tokens can
resolve and of the capitalised true name set. It is not a superset of the true name set
simpliciter: a name the texts never capitalise — a typo, an OCR slip, a stylistic lowercasing — falls
outside it. That gap is the instrument's and not the design's: R1 and R2 both require
capitalisation, so such a name is invisible to name_tokens too. Every recall figure below is
therefore labelled recall over capitalised proper names, and the one detectable part of the gap
— a GT_NAME type that also occurs lowercase somewhere — is counted and reported.
Computed on the two frozen texts: 593 distinct token types, 138 ever capitalised, 112 resolved by
name_tokens as names — and resolved ⊆ ever-capitalised is asserted, not assumed
(build_items.py).
Classifying all 138 therefore yields precision and recall-over-capitalised-names exactly, with no sampling and no generation task for the raters to be noisy at.
5. Procedure
Raters. Three seats, config/models.md P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash,
P3 x-ai/grok-4.5. Declared reserve for every seat: P5 deepseek/deepseek-v4-pro. Non-Anthropic by
charter §5. This is a factual-adjudication task, not a quality judgement — the class on which
config/models.md's S015 instrument note records the panel performing strongly (20/20 planted false
claims rejected, discrimination 0.886–0.917) — so Tier D's failure does not bear on it, and no
quality figure about any translation is computed anywhere in this design.
Payload. One call per seat. Both English texts in full, then the 138 candidates as
N001–N138, each in a different pseudo-random order per seat (seeded by seat index, seed
recorded), each accompanied by up to three quoted occurrences in context taken verbatim from the
frozen texts. max_tokens 6,000, temperature 0, attempts capped at 2 per slug (S079's overrun
cause).
Response format. One line per item, Nnnn | VERDICT | brief reason, then a literal END
terminator. Verdicts:
NAME— a proper noun in this text: a personal name, a place name, or a fixed byname or epithet that forms part of a name (the Red, Flatnose, Fiddle).NOT— not a proper noun in this text.EITHER— genuinely both: the type occurs as a name in one place and as a common word in another, or its proper-noun status is an editorial convention rather than a fact about the text (the Thing, the Law Council).
EITHER exists because the alternative is to force a binary and then read the forced choices as
data. It is scored separately and never silently folded into either side.
6. Controls
CTRL-POS, twelve items. Types no competent reader can call anything butNAME: nine personal names —hrut,hoskuld,hallgerd,mord,unn,thorgerd,kjartan,bolli,eyvind— and three place names,hordaland,norway,bergen. A seat scoring below 12 of 12 is not reading the task.CTRL-NEG, twelve items. Types no competent reader can call anything butNOT, all of them present in the candidate set only because they open a sentence:the,and,he,she,so,but,then,there,they,when,where,well. A seat scoring below 12 of 12 on these is not reading either. Both control blocks are scored before any figure is computed and are stated at the head of the prompt, per note (bgo).- Order control. The three seats see three different permutations. If a seat's verdicts correlate with position rather than with item, that is visible and reported.
- Byte-identical repeat, note (bfz). One seat's payload is re-issued unchanged after the primary run and the cell-level agreement between the two bodies is reported. [A1, BLOCKING] The criterion is relative, because the proposition it implements is:
the repeat's cell-level agreement must be ≥ the mean pairwise cell-level agreement of the three seats. Below that, every three-seat agreement figure on this page is descriptive only — the instrument disagrees with itself more than the seats disagree with each other. Independently, a repeat below an absolute 0.80 means the instrument is not usable for this task and the precision and recall figures are withdrawn, not downgraded.
Both legs are reported with their numbers whether they fire or not. The absolute 0.90 this section carried before the critic pass was looser than its own stated proposition — note (bgx), third instance in three sessions, and the critic found it from the design text alone.
7. Failure criteria, registered
- F1 — any seat scoring below 12 of 12 on
CTRL-POSor below 12 of 12 onCTRL-NEG: that seat is dropped from the majority and the remaining figures are descriptive only. - F2 — three-seat Fleiss' κ over the three-way verdict below 0.60: the majority is not a ground truth, and every precision and recall figure on this page is descriptive only.
- F3 — the runner's item guard. A body is a seat failure unless it carries exactly one line per
item id in that seat's own payload, matched against that payload's own ids, plus the
ENDterminator. Note (bgw): the previous session's guard rejected four good dispatches because the item-id pattern was carried forward hard-coded.call.pyhere therefore takesline_reas a required argument with no default. - F4 — fewer than 20 shared 7-grams in the gate, which would leave the consequential arm (§8.3) with nothing to move. Not fired: 34.
8. Analysis, specified before the run
8.1 Ground truth. GT_NAME = the types with a majority NAME among the surviving seats;
GT_NOT = majority NOT; GT_EITHER = majority EITHER or a three-way split. Reported with the
counts of each.
8.2 Precision and recall of name_tokens, over the 138 candidates, computed twice — once
counting GT_EITHER as NAME (the lenient reading) and once as NOT (the strict one), because the
choice is a judgement and hiding it inside one number is the defect this project keeps finding:
precision = |resolved ∩ GT_NAME| / |resolved| · recall = |resolved ∩ GT_NAME| / |GT_NAME|
[A3] and computed twice again, type-level and token-weighted, each type weighted by its occurrence count across the two texts. §1 asks a question about tokens and a type-level figure does not answer it: a name occurring fifty times and a false positive occurring once are one item each. Where the two disagree, the disagreement is the finding.
[A5] The spread between the lenient and the strict reading is itself a registered output. It is the measure of how much the answer depends on an orthographic convention rather than on a fact about the text, and it — not either endpoint — is what a later reader should carry away.
8.3 The consequential arm. Recompute the shared-7-gram counts of §3 with the resolved set
replaced by GT_NAME, and report the change in the name-excluded count (17 as computed by the
tool on the repaired materials; 14 was the unrepaired figure).
8.4 Attribution. For every false positive, report which rule admitted it — R1 (capitalised at a
non-sentence-opening) or R2-only (never occurs lowercase in the cell) — by re-running name_tokens
with each rule alone. This is what makes the finding actionable rather than a complaint.
9. Predictions, registered before dispatch
- [A2]
P1a— type-level precision under the strict reading (EITHER=NOT) falls in [0.75, 0.90].P1b— the number of false positives under the lenient reading (EITHER=NAME) is between 10 and 25 of the 112 resolved types. (The pre-criticP1, "precision < 1.00", was withdrawn as unfalsifiable: the resolved list visibly containsthat,what,now,my,one,here, so the existing data already proved it.) - P2 — recall ≥ 0.95. Basis: in name-dense narrative every name occurs mid-sentence at least once, so R1 fires on it.
- P3 — the false positives are predominantly R2-only, not R1. Falsifiable by §8.4: if most are R1, the sentence-boundary rule is still the weak part after its repair.
- P4 — the name-excluded shared-7-gram count rises above the tool's own figure when the resolved set is replaced
by
GT_NAME, because false positives are common words that sit inside many 7-grams. A fall, or no change, falsifies it. - P5 —
GT_EITHERis non-empty, i.e. at least one type is genuinely both. If it is empty the three-way scheme was unnecessary and the label-reachability lesson of this arm's own step 1 (note (bdq), note (bfq)) applies to this design too, and will be recorded as applying.
10. Budget
Declared worst case $0.76, built from max_tokens per note (abc) and from the caller's real
retry structure per S079 — the estimate is (seats × attempts × slugs), not seats:
| line | worst case |
|---|---|
pre-run critic, qwen/qwen3.7-max, cap 12,000 |
$0.16 |
| three seats × 2 attempts, payload 11,288 input tokens measured, cap 6,000, P3 priced at twice list for routing (S022 caution) | $0.48 |
| byte-identical repeat, one seat × 2 attempts | $0.12 |
Day headroom at session open $2.472159157 (three prior sessions, $2.527840843 of $5.00). The translation limb, the gate, the item build and every analysis are $0.00.
11. What this design cannot show
- Nothing about translation quality. No quality judgement of
T-njala-R04-v1, Dasent, or anything else is made or elicited. Tier D's failure is therefore not a limit on any figure here, and this is stated so that no later reader borrows these numbers for a purpose they do not carry. - Nothing about other cells. Precision and recall are measured on one name-dense cell in one language pair. A rule that is 0.8 precise here may be better or worse on ordinary narrative; §8.4's rule attribution is the part that generalises, because it names a mechanism.
- Nothing that clears a shared run. The gate figure is a measurement, not a clearance, and the
six FORCED verdicts of
RS-20260728bare untouched. - [A4] Nothing about names the texts never capitalise. Recall here is recall over capitalised proper names. The restriction is shared by the instrument under test, which is why it is a stated scope and not a repaired defect.
- [A5]
GT_EITHER's membership is convention-dependent. Whether the Thing or the Law Council is a proper noun is an editorial convention, and three models will apply one without being told which. The design does not stipulate a convention — doing so would decide by instruction the question the experiment exists to measure — so the lenient/strict spread is the honest output and a single precision figure quoted from this page without its spread is a misquotation. - The seats are not a gold standard. They are three models adjudicating a linguistic fact, with two control blocks and a self-consistency repeat to say how much that is worth. Charter §4 forbids reading their agreement as validation; F1 and F2 are what make the disagreement readable.