Repository path: workshop/experiments/E-20260801c-anchor-instance/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260801c-anchor-instance |
| status | frozen |
| created | 2026-08-01 |
| updated | 2026-08-01 |
| senses | naturalness, voice, style-correspondence |
| provisional | true |
| internal-judgment-only | true |
| links | wiki/arms/ARM-anchor-instance.md, wiki/base/anchors/A-mchugh-presence/A-mchugh-presence.md, workshop/experiments/E-20260731f-catalogue-reach/design.md, workshop/experiments/E-20260731f-catalogue-reach/claims.json, workshop/translations/pelsen/R10c-v1/translation.md, workshop/translations/pelsen/opportunity.md, config/models.md, config/budget.md |
E-20260801c — does the anchor's second instance attest the same catalogue?
Frozen before any stage-A or stage-F call. The gate (gate/) is a separate registration and ran
first; it is not part of this design.
1. The question
A-mchugh-presence is the project's only Tier 1 anchor built on two texts, and it states its own
reason: "one text cannot distinguish a register with no markers from an author with no mannerisms."
The catalogue was then written as if from one text — E-20260731f/claims.json labels every one
of the nineteen propositions C01–C19 with passage: presence-excerpt.txt — and
snake-girl-excerpt.txt has never been read by any instrument, deterministic or panel.
So the anchor's stated design has never been executed. Two questions:
- Which of
C01–C19do both stored instances attest, and which rest on one? - Whether the shared runs between
T-pelsen-R10c-v1and its published comparator are forced by the Swedish — the mandatory contamination gate on this session's translation limb returned the largest overlap this project has recorded for a lead translation, and a run length alone does not say what produced it.
The wire between the limbs, in one sentence. The translation was written to this anchor's catalogue as a target register, so the propositions that actually decide its wordings are on record; the study limb then measures which of those propositions the anchor's two instances both attest, and therefore whether the translation was aimed at a register or at one author's mannerisms.
2. Materials
Claim set — C01–C19, reused verbatim from E-20260731f/claims.json
(sha256 1e9430e4fbfab8d1129e286226a6321341d6c621288c2ecdcf7384a6af692534), read programmatically,
never retyped. Not re-decomposed: re-writing the propositions would make the two runs
incommensurable, and the point is to run the same instrument on the text it was never run on.
Passages.
| tag | text | words | role |
|---|---|---|---|
| P | McHugh, "Presence" (2005), opening — presence-excerpt.txt |
1,511 | the anchor's primary instance |
| S | Kessel, "The Snake Girl" (2008), opening — snake-girl-excerpt.txt |
1,505 | the anchor's second instance, never read by any instrument |
| Y | Doctorow, Little Brother (2008), opening | 1,760 | the fixed contrast, identical in both arms |
Y is held fixed across the two arms deliberately. E-20260731f assigned contrasts
A↔C, B↔A, C↔B; its arm for P used Y. Using Y for S as well keeps the only
difference between the arms the target passage, and keeps arm P comparable to the published run.
Translation limb. T-pelsen-R10c-v1 — Söderberg «Pelsen», 1,273 Swedish words → 1,298 English,
R10 against this anchor, frozen at 217ce02 with its opportunity list frozen before it at
9fb7ba2. Contamination high, measured after the freeze: longest run 24 tokens, 46 shared
12-grams, 404 of 1,299 tokens (31.1%) inside a shared run ≥ 7 with Stork 1923.
3. Seats
Per config/models.md. All non-Anthropic. Used only on the factual side of the S015 line —
is this stated property of English present in this passage — never is this prose good. Tier D has
not passed; every verdict here is internal-judgment-only and provisional.
- Raters, both arms and stage F — P1
openai/gpt-5.6-terra, P2google/gemini-3.6-flash, P3x-ai/grok-4.5. P1 and P3 areE-20260731f's pair, for comparability. - Pre-run critic —
qwen/qwen3.7-max, probed-but-not-selected, not a rater and not the raters' reserve (S053 role-collision fix). - Reserve for every stage — P5
deepseek/deepseek-v4-pro, declared before dispatch, note (bfc).
4. Procedure
Stage 0 — deterministic, free, and RUN BEFORE THIS DESIGN WAS FROZEN
Two things, both by importing E-20260731f/analysis/checks.py unmodified and running it against
snake-girl-excerpt.txt, which that module already knows how to load and was never asked about:
- 0a — the ten machine-reachable claims, on both texts. Results are printed in the result page and are the source of the control block below.
- 0b — the hand read. Every quoted span on the anchor page is attributed to P, to S, to
both, or to neither, by unique-substring match against the two stored bodies. This is the backlog
row
wiki/backlog.mdhas carried since S074.
Stage 0's results were obtained before this design was written and are declared here rather than predicted, because a design cannot register a prediction about a computation it has already run. They are used in two ways only: to build the control block, and as the thing the panel arms are compared against.
Stage A — the per-text attestation (6 calls)
One call per (passage ∈ {P, S}, rater ∈ {P1, P2, P3}). Each call carries the target passage X,
the fixed contrast Y, and 22 items: C01–C19 in their frozen wording plus three planted
claims, interleaved and unlabelled.
- Verdicts are
PRESENT/ABSENTonly. No third option —RS-20260730dofferedUNCLEARon 266 cells and it was used zero times; note (ber). PRESENTmust carry the single strongest instance, quoted verbatim from X. APRESENTwhose quotation is not verbatim in X is recordedunattestedand is prediction P4's statistic.- No rater is told the two passages belong to one anchor, that a second arm exists, or that any claim was written about either text.
The control block — six cells per arm, every ground truth from stage 0's DATA, per note (bfy):
| id | claim | expected on P | expected on S | mechanical basis |
|---|---|---|---|---|
| K1 | C14 — forty-eight paragraphs across 1,511 words |
PRESENT | ABSENT | 48 / 1,511 against 36 / 1,505 |
| K2 | C13 — very short paragraphs, one-sentence paragraphs structural |
PRESENT | PRESENT | mean 31.4 / 41.6 words (≤ 45); one-sentence 8 / 5 (≥ 5) |
| K3 | C04 — no dialect spelling |
PRESENT | PRESENT | declared dialect patterns: none in either |
| K4 | planted: the narration is in the first person | ABSENT | ABSENT | 0 first-person tokens in narration in either |
| K5 | planted: at least one word is set in italic or bold type | ABSENT | ABSENT | 0 emphasis spans in either stored file |
| K6 | planted: at least one paragraph runs to more than three hundred words | ABSENT | ABSENT | longest paragraph 122 / 169 words |
K1 is the only cell whose correct answer DIFFERS between the arms, and it is therefore the only one that can catch a rater answering from a general impression of unmarked contemporary prose rather than from the passage in front of it. K1 is declared to be doing double duty — it is both a control and one of the nineteen findings — and that is stated here rather than discovered later.
Hairline cells excluded from the control block, and why. Stage 0 puts C08 on S at
say-rate 0.812 against a PASS threshold of 0.80, and C17 on S at lag-1 autocorrelation
0.303 against a threshold of 0.30. Both are within 0.015 of their thresholds. They are the two
most interesting cells in the run and they are not used as ground truth for anything.
Stage F — is the overlap forced? (3 calls)
The instrument is RS-20260728b-forced-run-ru's, unchanged. Three Swedish paragraphs — the ones
carrying the four longest shared runs (source ¶4, ¶16, ¶34) — are given alone, with no English,
no comparator, and no mention that any published translation exists, and each seat is asked for a
contemporary English rendering.
Frozen verdict rule, inherited: a shared run is FORCED iff ≥ 2 of 3 independent seats reproduce ≥ 7 contiguous tokens of it; ELECTIVE otherwise.
Reported alongside, and it is the more informative number: each seat's own longest common run against Stork 1923 and against the lead, over the same three paragraphs.
5. Registered predictions
| # | prediction | what falsifies it |
|---|---|---|
| P1 (primary) | ≥ 4 of 19 propositions get different majority verdicts on P and S | ≤ 3 differ — the pair attests one catalogue and the anchor's two-instance design is doing its job |
| P2 | C08 (say-attribution, no ornamental verbs) is PRESENT on P and ABSENT on S |
any other pattern. Registered because stage 0 finds gasped and whispered six times in S and zero ornamental verbs in P, and because the hand read finds one of that item's four quotations is S's |
| P3 | C14 differs (K1) |
it does not — in which case the raters are not reading the passages and F1 has probably already fired |
| P4 | the unattested rate — PRESENT whose quotation is not verbatim in X — is ≤ 0.15 on both arms |
above 0.15 on either |
| P5 (closure-defeating, registered as such) | — | if the arms differ at ≥ 8 of 19, the anchor's sentence "one text cannot distinguish a register with no markers from an author with no mannerisms" must be restated, not annotated, and this arm may not close without doing so |
| P6 | the 24-token run is FORCED | fewer than 2 of 3 seats reproduce ≥ 7 tokens of it |
| P7 | the 15-token final-line run ("it has given me the last seconds of happiness i have known in my life") is NOT FORCED | 2 or 3 seats reproduce ≥ 7 tokens of it |
P1 is the primary. P3 is a control restated as a prediction and buys nothing on its own.
6. Failure criteria — pre-committed
- F1. On a given arm, fewer than 2 of 3 raters pass all six controls → that arm is descriptive only and no cross-arm difference count is reported as an estimate. (If it fires on one arm only, the comparison is descriptive; there is no one-arm result to keep.)
- F2. Fewer than 15 of 19 items receive a majority verdict on either arm → that arm is descriptive only.
- F3. A seat whose
PRESENTquotations are non-verbatim in X at a rate > 0.5 is void for that arm and the arm falls to two raters, with the two-rater reading declared post hoc. - F4. A stage-F seat returning fewer than 3 paragraphs of English is a seat failure, in
the runner, not in the analysis. (Note (bgt), bought last session:
finish_reasonis not a truncation check.) - F5. Any deviation from §7's ordering → the affected stage is void, not adjusted.
7. Order of operations
- Target propositions frozen — S074,
sha256 1e9430e4…. - Source ingested; opportunity list frozen —
9fb7ba2. - Translation and translator's log frozen —
217ce02. - Comparator fetched and contamination measured —
7a3b39a. (3 before 4 isR10§1.) - Stage 0 run.
- This design frozen.
- Pre-run critic; amendments written and committed before any stage-A or stage-F call.
- Stage A, then stage F. Raw bodies preserved.
analysis/verify.py, which imports nothing fromanalyse.pyand re-parses every answer from the stored.rawbytes, with mutation tests.
8. What this design cannot do
- It cannot say the anchor is wrong about the register. It can say whether its two instances attest the same propositions. A register can be real and two texts can still differ on how they instantiate it.
- It judges nothing. No prose is assessed as good or bad anywhere, and Tier D has not passed.
- The propositions are the lead's decomposition (S074) and remain so. A difference between arms could be a difference in the texts or a proposition that is badly worded for one of them, and this design cannot separate those. Where a difference turns on wording, the result page says so.
- Stage F cannot clear the lead.
RS-20260730festablished that this project's memory controls do not work on published material; a FORCED verdict withdraws one explanation and supplies none. - One anchor, two texts, three raters, one arm of one language pair.
9. Cost — pre-flight, note (abc), and built from the RETRY STRUCTURE as well as from max_tokens
The gate this session ran overran its own declared worst case by 3%, and the reason was not the token cap: it was the runner's fall-through, which can dispatch a seat up to 3 attempts × 2 slugs. An estimate built from one dispatch per seat is not a worst case. Attempts are therefore capped at 2 per slug for this design and the estimate is built accordingly.
| stage | shape | max_tokens |
worst case |
|---|---|---|---|
| pre-run critic | 1 call, qwen/qwen3.7-max |
12,000 | $0.16 |
| stage A | 2 arms × 3 seats, in ≈ 4,600 | 8,000 | $0.38, ×2 attempts = $0.76 |
| stage F | 3 seats, in ≈ 900 | 4,000 | $0.088, ×2 attempts = $0.18 |
| one reserve dispatch | — | 8,000 | $0.06 |
Declared worst case $1.16. Spent so far today across all sessions $2.008982423; headroom $2.991017577. The gate's $0.105042406 is already inside that figure.