Repository path: workshop/experiments/E-20260731b-figure-audit/design/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260731b-figure-audit |
| status | frozen |
| created | 2026-07-31 |
| updated | 2026-07-31 |
| senses | accuracy |
| links | wiki/arms/ARM-figure-audit.md, wiki/method-notes.md, wiki/findings/results/RS-20260727e-sense-boundary.md, wiki/findings/results/RS-20260728f-nonlead-items.md, wiki/findings/results/RS-20260729-drift-window-verify.md, wiki/findings/results/RS-20260729b-graded-drift.md, wiki/findings/results/RS-20260729c-neutral-summary.md, wiki/findings/results/RS-20260729g-graded-senses.md |
E-20260731b-figure-audit — the label-reachability sweep, and a positive control for the one label it cannot settle
ARM-figure-audit step 1. Frozen before any figure was recomputed and before the
translation limb's source was opened for translation.
1. The question, and the two halves it divides into
Note (bdq) says: any future use of a label set must check that every label is reachable before quoting agreement — which is what a positive control is for. Note (beb) says: before quoting an agreement statistic over a fixed option set, print the realised distribution and say how many options were actually used, and adds a consequence — a chance-corrected agreement figure computed over a five-category vocabulary is corrected against categories that do not exist in the data.
Those are two different claims and this design separates them, because they have different truth conditions and only one of them is about arithmetic.
Half A — the arithmetic claim, and it is decidable by computation alone. Does an unused category in the nominal vocabulary change the chance-corrected statistic? This is a property of Fleiss' κ, Cohen's κ and Krippendorff's α, not of any particular run. It is settled by recomputing every published figure twice — once over the nominal vocabulary, once over the realised categories only — and printing the difference.
Half B — the validity claim, and it is not decidable by computation. If a label never fires, the figure is agreement about the other labels, whatever its arithmetic. Whether that voids what the figure was read as evidence for depends on what the figure was quoted for. This half is answered cell by cell, in prose, against what each result page actually claims.
Half C — what neither half can settle, and which is the reason for the translation limb. A label that never fires in the data may be unreachable in the instrument (raters will not use it whatever they are shown) or merely absent from the material (no site called for it). Stored outputs cannot distinguish these. A positive control can, and it needs material built to elicit the label. That material does not exist and is made here.
2. Materials — the six runs and the seven published figures
Every figure below is recomputed from the stored raw bodies. Nothing is re-run against the API for Halves A and B; the whole sweep is local computation.
| # | published figure | page | run | nominal vocabulary |
|---|---|---|---|---|
| 1 | Fleiss κ, 3 strata × 2 conditions (0.550/0.413, 0.556/0.286, 0.036/0.647) | RS-20260727e §3 |
E-20260727c |
4 offered + off-list carried |
| 2 | Fleiss κ, 4 conditions (0.648/0.354/0.377/0.599) | RS-20260728f §2 |
E-20260728f |
5 |
| 3 | four-class agreement 0.654 | RS-20260729 |
E-20260729 |
4 |
| 4 | four-class α 0.515, Cohen κ 0.518, raw 0.654 | RS-20260729b §3 |
E-20260729 (same bodies) |
4 |
| 5 | binarised α 0.5139 / 0.5160 | RS-20260729b §3 |
E-20260729b + E-20260729 |
2 |
| 6 | the five-option ratification vote | RS-20260729c |
E-20260729c |
5 (A–E) |
| 7 | CAT α 0.553 / 0.478 (and every figure derived against it) | RS-20260729g §1 |
E-20260729g |
5 |
Figure 7 is not on the arm's list of five and is swept anyway, because note (beb) names
E-20260729g explicitly and its CAT label is the only published α in the project computed over
a five-way vocabulary.
3. Procedure — frozen
- Recover the offered vocabulary from the runner source, not from the result page. For each run, quote the literal option block sent in the prompt, by file and line. A vocabulary read off a summary is the thing this sweep exists to distrust.
- Recover the realised distribution from the stored bodies, by a parse written for this sweep
and importing nothing from the original
analyse.py. - Per cell, where a cell is the exact (items × raters × condition) set the published figure was computed on: list offered labels, list realised labels, name the unused ones.
- Recompute each published statistic twice — over the nominal vocabulary and over the realised categories only — and print both, to 6 dp, alongside the published value.
- Reproduce the published value first. A recomputation that does not land on the published figure is a discrepancy to be reported before anything else is said about it.
- Correct in place anything that moves, and correct the notes if the notes are what moved.
4. Predictions, pre-registered
- P1 (the arithmetic). For every figure in the table, nominal and realised-only computations are identical to 1e-9, because every one of these statistics derives its expected agreement from observed marginals and a category with zero observations contributes zero. If P1 holds, note (beb)'s consequence clause is false and is corrected, and no published number moves on those grounds.
- P2 (reproduction). Every published figure reproduces from the stored bodies to 3 dp. Any that does not is reported as a discrepancy and investigated before P1 is read.
- P3 (reachability, the one the sweep is really for). At least one cell among figures 1–7 has an offered label with zero realised uses. Named, with what the figure was quoted for.
- P4 (the note audit). Note (beb)'s factual claims about which options went unused in
which runs are checked one by one against the stored bodies. This is the prediction the lead
expects to fail, because the first pass over
E-20260727c's runner already shows a four-option prompt where the note asserts a five-option one.
Failure criterion for the sweep as a whole. If fewer than three of the seven figures can be reproduced from stored bodies, the sweep is not a check on those figures and says so instead of reporting differences.
5. The positive control — Half C
What it tests. composite and both are the two labels the project has watched go
near-unused. The question the stored outputs cannot answer: would a rater use composite on a
site where the marked element genuinely separates into a style part and a culture part?
Materials. Built by the translation limb: 蒲松齡〈促織〉, classical Chinese, translated in
session, with a site census frozen before a word of the translation is written. The census
declares, per site, which label the lead's own reading says the site calls for — that declaration
is internal-judgment-only and is not the measure. The measure is whether the label fires at
all.
Why this material. Late-imperial Chinese officialdom fuses the two senses at a single
lexical site by construction: 里正 is at once an office in a social world (cultural-mediation)
and a piece of bureaucratic-register diction (style-correspondence), and rendering it requires
two separable decisions — what to call the office, and at what register. If composite is
reachable anywhere, it is reachable here.
Design. Three raters, one condition, the same five-option block used in E-20260728f,
verbatim, on a mixed item set: the composite-candidate sites plus filler sites the census
declares single-sense. Filler is not decoration — a run where every item is a composite candidate
measures compliance, not reachability.
Pre-registered reading, and it is deliberately weak because reachability is a weak question.
- C1 — reachable.
compositeis used on ≥ 3 distinct items by ≥ 2 of 3 raters. Then the zero rates inE-20260728fB/B2 andE-20260729gare about the material, and the correct repair to those runs is item selection, not the option list. - C2 — unreachable.
compositeis used on ≤ 1 item across all three raters. Then the label is an option in name only and every α computed over a vocabulary containing it is, in (beb)'s own sense, reporting a vocabulary the raters do not have. - C3 — anything between is reported as between, with the count, and settles nothing. This is named in advance so that a middling result is not written up as either.
What this control cannot do. It cannot show the raters are right about which sites are composite; the lead's census is one reader's judgment and the run is not scored against it for correctness. It cannot generalise past this material. And three raters agreeing to use a label is not evidence the label carves anything.
Cost gate. The control runs only if the sweep leaves the reachability question live, and only
inside the day's remaining headroom, estimated from max_tokens per note (abc). If it does not
run, the sweep reports Half C as unrun and the arm's step 2 inherits it.
6. Verification
verify.py, importing nothing from sweep.py, re-parsing every stored body by a different path,
re-implementing Fleiss' κ, Cohen's κ and Krippendorff's α from first principles, and mutation-tested
— at least three mutations that must each be caught.