Repository path: workshop/experiments/E-20260731f-catalogue-reach/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260731f-catalogue-reach |
| status | frozen |
| created | 2026-07-31 |
| updated | 2026-07-31 |
| senses | naturalness, voice, style-correspondence |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-anchor-second-read.md, wiki/base/anchors/A-mansfield-garden-party/A-mansfield-garden-party.md, wiki/base/anchors/A-doctorow-little-brother/A-doctorow-little-brother.md, wiki/base/anchors/A-mchugh-presence/A-mchugh-presence.md, wiki/findings/results/RS-20260730d-register-centre.md, wiki/findings/results/RS-20260731-anchor-second-read.md, workshop/regimes/R10-anchored-register.md, workshop/translations/karen/R10c-v1/translation.md, workshop/translations/karen/R10p-v1/translation.md, config/models.md |
E-20260731f — can a feature catalogue be second-read?
ARM-anchor-second-read step 2, the arm's last. Frozen before any check is written and before any call is dispatched. Claims frozen at ca602a7; the matched translation pair and its logs frozen at 836f58c, both before this file existed.
1. The question
The arm's condition 2: the naturalness anchors "are either second-read under a brief that can actually check a feature catalogue, or the arm records why a feature catalogue is not second-readable and what that means for the three naturalness anchors that rest on one."
Step 1 discharged condition 1 with a deterministic instrument: 193 of 194 quoted correspondences attest, 33 of 33 counts recompute, over three anchors that make quotable and countable claims. The arm then wrote that "the three naturalness anchors carry no quantitative claims at all, so the deterministic instrument used here reaches none of them." That sentence is the first thing this design tests, because step 1 had to begin by correcting an equally confident sentence in the same arm page.
Two questions, in order:
- How much of a naturalness anchor's feature catalogue does a deterministic instrument reach, and what is the yield?
- For the residue — the propositions no script can settle — can an independent reader attest them, and does the attestation carry evidence?
And the control the project has repeatedly found itself without: a text into which the catalogued features were deliberately installed. If a blind reader cannot find, in prose written to a catalogue, the features that catalogue names, then ABSENT verdicts on the published anchors mean nothing either.
2. Materials
Claim set — claims.json, 41 atomic propositions, decomposed by the lead from the three anchors' ## What it grounds sections. Propositions are lead work and are declared as such. Quotations are not retyped: analysis/build_claims.py extracts every quoted span from the anchor page mechanically and resolves each claim's quotations against that pool by unique substring, so a mistyped quotation cannot enter the file. Counts: A-mansfield-garden-party M01–M11, A-doctorow-little-brother D01–D11, A-mchugh-presence C01–C19.
Passages.
| tag | text | words | provenance |
|---|---|---|---|
| A | Mansfield, "Miss Brill" (1922) | 1,891 | A-mansfield-garden-party/miss-brill.txt |
| B | Doctorow, Little Brother (2008), opening | 1,760 | A-doctorow-little-brother/little-brother-excerpt.txt |
| C | McHugh, "Presence" (2005), opening | 1,511 | A-mchugh-presence/presence-excerpt.txt |
| Tc | T-karen-R10c-v1 — Kielland «Karen» ¶1–15, target A-mchugh-presence |
862 | frozen 836f58c |
| Tp | T-karen-R10p-v1 — same source unit, target A-mansfield-garden-party |
950 | frozen 836f58c |
Tc and Tp are a matched pair: one source unit, one translator, one session, single pass each, differing in the target catalogue. Both contamination-gated against the Collier 1907 English before either was committed (5 and 11 tokens, 0 shared 12-grams both).
3. Seats
Per config/models.md. All non-Anthropic. The panel is used only on the factual side of the S015 line — is this stated property of English present in this passage — never is this prose good. Tier D has not passed; every verdict below is internal-judgment-only and provisional.
- Rater P1 —
openai/gpt-5.6-terra. Rater P3 —x-ai/grok-4.5. The same pair asRS-20260730d, for comparability. - Pre-run critic —
moonshotai/kimi-k3(P4). Not a rater, per the S053 role-collision fix. - Reserve for every stage —
deepseek/deepseek-v4-pro(P5), declared before dispatch, note (bfc).
4. Procedure
Stage 2 — deterministic reach (no API, no cost)
For each of the 41 claims, the lead writes a check in analysis/checks.py after this design is frozen. A check is admissible only if it (i) runs against the stored passage file with no run-time judgment, (ii) returns PASS / FAIL / UNREACHED, and (iii) is written before its result is looked at.
Reachability is demonstrated, not classified. A claim is reached iff a check settles it. There is no lead-assigned type, and therefore nothing for the lead to fudge: the script either settles the claim or reports UNREACHED. Every FAIL is a yield item, applied to the anchor page.
Stage 3 — the contrast second-read of the residue (6 calls)
Every claim UNREACHED at stage 2 goes to two blind raters, one call per (passage, rater).
The brief is CONTRAST, not SALIENCE — the arm's most important inherited constraint, which cost E-20260730d its primary statistic. The rater is given the target passage X, a named comparison passage Y of a different register, and the claim list, and is asked whether each statement is true of X. Comparison assignment, fixed here: A↔C, B↔A, C↔B.
Verdicts are evidence-bearing. PRESENT must be accompanied by the single strongest instance, quoted verbatim from X. A PRESENT whose quotation is not verbatim in X is recorded as unattested and is the statistic of prediction P4. This is the upgrade on RS-20260730d, where PRESENT was unfalsifiable.
Options are PRESENT / ABSENT only. No third option: RS-20260730d offered UNCLEAR on 266 cells and it was used zero times, and RS-20260729g offered two third options across 240 cells and they were used none — note (ber). Offering one again would buy nothing.
Stage 4 — the crossover, with planted negatives (8 calls)
Four cells: passage ∈ {Tc, Tp} × claim set ∈ {C-set, M-set}. In every cell the comparison passage Y is the other translation, unlabelled. Two raters, one call per cell per rater.
- C-set (17 items) — C01–C06, C08, C10–C13, C15–C19 (the 16 on the frozen opportunity list) plus C07, planted.
- M-set (11 items) — M01–M06, M08–M10 (the 9 on the frozen opportunity list) plus M07 and M11, planted.
- Excluded from stage 4 entirely: C09 and C14, which are counts over the anchor's own 1,511 words and cannot hold of any other passage.
The planted negatives are the internal control. C07 (one consciousness without exception) is refused by the source, which leaves Karen for the wind, the postboy and the fox; M07 (written-out sound play) has no material, the source naming its noises and spelling none; M11 (an irony sprung by an overheard cruelty at the close) cannot hold of a passage that is a story's opening. All three were placed on the frozen opportunity list before either translation was committed and nothing was done to simulate any of them. A rater who attests them is attesting what is not there.
At the end of each stage-4 call, and not announced in advance, the rater is asked whether it recognises passage X — the probe RS-20260730d §5 made informative.
Stage 5 — the byte-identical repeat floor (2 calls)
One stage-4 cell — (X = Tc, C-set) — is re-sent byte-identically, same session, same parameters, to both raters. This is the drift gate's option 1. It measures how much the primary statistic moves when nothing changes, and it is the reason S073's headline survived: the floor it measured was larger than the effect the design was looking for.
5. The registered statistics
reach= reached / 41, and per anchor.yield= claims a checkFAILs.att(P)= attestation rate:PRESENTby both raters / items, on passage P.ev= evidence rate:PRESENTverdicts whose quotation is verbatim in X / allPRESENTverdicts. Verbatim means after case-folding, whitespace collapse and removal of markdown emphasis; nothing else.- The primary is the crossover interaction
I = [att_C(Tc) − att_C(Tp)] + [att_M(Tp) − att_M(Tc)], computed over opportunity items only, planted negatives excluded, consensus (both ratersPRESENT). floor= the sameI-component recomputed with the stage-5 repeat substituted for its original cell; reported as|Δ|onatt_C(Tc).plant= attestation rate on the three planted negatives.
6. Predictions, registered
- P1 —
reach≥ 12 of 41. The arm's premise that a deterministic instrument reaches none of the three naturalness anchors is false. Declared non-blind: the lead has read all three anchor pages and knowsA-mchugh-presenceprints two counts. What has not been computed is the size of the reachable set, which is what the number tests. - P2 —
yield≥ 1. Three verification passes on this shelf, three non-zero yields (S015, S053, S058) and a fourth at S069. A fifth zero would be the first. - P3 —
atton own-passage stage-3 items ≥ 0.80 by consensus. The catalogues were written by close reading of these exact files; an independent reader given a contrast should confirm most of them. - P4 —
ev≥ 0.80. - P5 —
I> 0 andI>floor. The catalogue transfers: features installed for one target are found in that arm and not the other. - P6 —
plant<atton the opportunity items, on both raters.
7. Failure criteria, registered
- F1 — raw inter-rater agreement < 0.70 over all stage-3 and stage-4 cells → the instrument is not reportable and condition 2 closes on reliability, negative.
- F2 —
ev< 0.80 →PRESENTverdicts are not attested and no attestation rate in this design is reportable as evidence. - F3 —
floor≥ |I| → the crossover is reported as a null inside its own floor, exactly as S073's framework effect was. - F4 —
reach= 0 → the arm's premise stands and condition 2 closes on the written reason it names. - F5 —
plant≥ 0.50 → the raters attest absent features and everyPRESENTin this design is suspect.
A failure on any of F1, F2 or F5 makes the crossover unreportable regardless of I. Stated now so that a favourable I cannot be rescued after the fact.
8. What this design cannot do
- It is not calibration. Tier D has not passed and nothing here is a quality judgment.
- The two arms differ in orthographic standard as well as in register — Tc is American, Tp is British, declared in
T-karen-R10p-v1log 30 — because the two anchors are a 2005 American text and a 1922 British one. A rater could separate the pair on grey/gray alone. The crossover therefore cannot attribute a separation to register as against orthography without inspecting which propositions carried it, and the per-claim table is reported for that reason. - The opportunity list was written after the drafts were composed, which is a deviation from
R10§6 and is recorded onworkshop/translations/karen/opportunity.mdwith its direction of bias. n= 1 source unit, 1 translator, 2 raters. Nothing here generalises to catalogues in general; it is one catalogue pair, executed once.- Recognition is not controlled for A and B, which
RS-20260730dmeasured as recognised by both raters. It is controlled within the crossover, because both arms are the same story: recognition is constant across the two arms and cannot produce the predicted interaction.
9. Cost
Declared worst case $1.30, built from max_tokens and not from an assumed output length (note (abc)).
| stage | calls | slug | max_tokens |
worst case |
|---|---|---|---|---|
| pre-run critic | 1 | moonshotai/kimi-k3 |
12,000 | $0.20 |
| stage 3 | 6 | P1 ×3, P3 ×3 | 6,000 | $0.32 |
| stage 4 | 8 | P1 ×4, P3 ×4 | 6,000 | $0.42 |
| stage 5 | 2 | P1 ×1, P3 ×1 | 6,000 | $0.11 |
| routing margin (note (x), S022 caution) | — | — | — | $0.25 |
The critic's cap is 12,000 because S073's qwen3.7-max critic call burned 4,019 reasoning tokens and returned nothing at 8,000 — note (b), nineteenth firing. UTC day 2026-07-31 stands at $1.012808407 of $5.00 across five sessions; headroom $3.987191593.
10. Order of operations, and what is already frozen
ca602a7—R10v1.0 andclaims.json, before the source was read for translation.836f58c— both translations, both logs, the opportunity list, the contamination gate.- This file — frozen before
analysis/checks.pyexists and before any call. - Pre-run critic; findings accepted or refused in writing, amendments recorded in this file with an
A#id. - Stage 2, then stages 3–5.
analysis/verify.py— an independent recomputation of every reported number, importing nothing from the analysis path, with mutation tests.
Amendments, made in session and recorded before the dispatch each affects
A1 — the evidence rule is polarity-aware, and ev is redefined over quotation-required verdicts. Made by the lead from a dry render of a stage-4 prompt, before any rater call was dispatched.
The design as frozen said PRESENT must carry a verbatim instance and defined ev over PRESENT verdicts. Nine of the 41 claims are negative or universal — C01–C05 (no archaism, no period lexis, no era-marking slang, no dialect spelling, no learned display), C07 (without exception), C10 (every figure … none extended), C12 (not glossed), C19 (contains no sound play …). For these a PRESENT verdict asserts that nothing is there, so no instance can be quoted, and ev as frozen would have counted every one of them unattested by construction — firing failure criterion F2 for a reason with nothing to do with the raters.
The amended rule, in the frozen prompt: for a statement that something occurs, PRESENT carries the instance; for a statement that something does not occur or holds without exception, ABSENT carries the counter-instance and PRESENT carries NONE.
evis now the fraction of quotation-required verdicts whose quotation is verbatim in X, where quotation-required is determined by the frozen polarity list below crossed with the verdict returned. F2's threshold of 0.80 is unchanged.- The change buys a second yield channel. An
ABSENTon a negative claim now arrives with the counter-instance the rater found — a rater-supplied candidate correction to an anchor's own negative assertion, which the frozen design had no way to collect. - Frozen polarity list (negative or universal):
C01 C02 C03 C04 C05 C07 C10 C12 C19. Every other claim is treated as asserting occurrence. Polarity is a property of the claim's wording, fixed atca602a7; it is recorded here rather than inclaims.jsonso that the frozen file is not touched.
A2 — the header-stripping rule was wrong and was corrected before stage 2 ran. The three stored passage files delimit their provenance headers three different ways (--- EXCERPT BEGINS ---; a bare rule of hyphens; a rule of hyphens behind a #), and the rule first written matched none of them. Corrected in analysis/checks.py and independently re-implemented in analysis/verify.py, before any check was run and before any result was looked at. Had it not been caught, every check and every rater prompt would have carried the provenance headers.
A3 — the word-count rule is declared as whitespace tokens. Claims C09 and C14 assert 1,511 words. wc -w over the stripped body of presence-excerpt.txt returns 1,511; a letter-token count returns 1,506, the difference being numerals and P&G. The whitespace rule is the one the anchor's figure was produced by and is the rule the checks use; the letter-token count is printed alongside as sensitivity, so a reader can see that the claim's truth depends on a tokenisation choice.
A4 — the pre-run critic seat moved to the declared reserve. moonshotai/kimi-k3 was dispatched at max_tokens 12,000 — raised from 8,000 because of note (b) — and returned finish_reason: length, 11,997 reasoning tokens, zero characters of content, billed $0.195747. Note (b), twentieth firing, and the first time the cap raised in response to it was still not enough. The pass was re-dispatched to deepseek/deepseek-v4-pro, the reserve declared for every stage in §3 before any dispatch. The reserve is not a rater in this design, so the S053 role-collision fix is not breached.
Pre-run critic: NEEDS-REDESIGN, six findings, all six accepted
deepseek/deepseek-v4-pro (the declared reserve), Novita, stop, 5,223 in / 10,390 out of which 9,346 reasoning, 135 s, $0.030371504. Raw at runs/critic-reserve.raw, verdict at runs/critic-reserve.txt. No finding was refused.
A5 — Finding 1 (BLOCKING) is accepted and answered with an independent typing stage. The critic: "the lead writes all 41 atomic propositions, then later writes the checks.py script that decides which of those propositions a deterministic script can settle … reach is not an independent measurement." This is correct, and publishing the checks' source does not answer it. New stage 2b, 2 calls: two seats that are neither raters nor the critic — google/gemini-3.6-flash (P2) and qwen/qwen3.7-max (probed-but-not-selected, the seat S073 used) — are given only the 41 propositions, no passages, no checks, no anchor names, and asked of each: could a computer program decide this about a passage of English prose, using only string search, counting and arithmetic, with no reading judgment? reach_ind is their consensus YES rate, and reach is reported against it. If the lead's demonstrated reach materially exceeds both independent seats', the critic's objection is sustained by measurement and the reach figure is reported as the lead's own and not as a property of the catalogue.
A6 — Finding 2 (BLOCKING) is accepted and the primary statistic changes. The critic: the opportunity list was written after the drafts, so the items entering I are post-hoc and "any observed I is unreliable." Correct. The registered primary becomes I_full, computed over the whole C-set and M-set as fixed in claims.json at ca602a7 — before the source was read for translation and before either draft existed — with only C09 and C14 excluded, and those excluded on a ground fixed at the same commit (they are counts over the anchor's own 1,511 words). That item set cannot have been selected on the drafts, because it predates them. I_opp, over the opportunity subset, becomes secondary and is reported beside it. The full set carries the three planted negatives, which enter both arms symmetrically and can therefore only shrink |I_full|: the change is conservative. P5 is now a prediction about I_full.
A7 — Finding 4 (SERIOUS) is accepted and answered with a registered decomposition. The critic: I > 0 could be produced by the American/British orthographic split alone. Registered now, before any rater call: I_full is decomposed by whether the translator's log codes the claim A (aimed) for the arm in question. The A codes were frozen at 836f58c, before this design existed, and are extracted mechanically from the two logs. The orthography account predicts a uniform shift across all items of a set; the catalogue-transfer account predicts the effect concentrates on the A-coded items. P5 is supported only if I_full on the A-coded subset exceeds I_full on the rest. If the two are equal, the design reports that it cannot separate register transfer from orthography, which is the honest outcome the critic asked for.
A8 — Finding 3 (SERIOUS) was found independently and is already fixed. The critic's finding is amendment A1 above, which was written from a dry render before the critic returned. Two independent readings of the same design reached the same defect; the fix is unchanged.
A9 — Findings 5 and 6 (MINOR) accepted. (5) If an anchor's stage-3 residue is empty, that cell is skipped and P3/P4 are evaluated over the cells that ran; the count of skipped cells is reported. (6) The three planted negatives are pre-checked deterministically against both translations before the raters run — C07's third-person component, M07's syllable-play detector, and M11, for which no deterministic check exists and which is declared as such. A rater PRESENT on a planted item whose pre-check shows the feature genuinely crept in is a translation fact, not a rater error, and the two are separated in the report.
Cost, raised in session with the reason written before the next dispatch. Declared worst case rises $1.30 → $1.60, on three grounds: the wasted kimi-k3 dispatch billed $0.195747 against nothing; the reserve critic added $0.030372; and A5 adds two calls. Note (abc). Spent so far $0.226118504. UTC day would stand at $2.61 of $5.00 in the worst case.