Repository path: workshop/experiments/E-20260725-anchor-verification/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260725-anchor-verification |
| status | frozen |
| created | 2026-07-25 |
| updated | 2026-07-25 |
| links | wiki/base/anchors/A-shaw-spider-thread/A-shaw-spider-thread.md, wiki/base/anchors/A-garnett-vanka/A-garnett-vanka.md, wiki/base/anchors/A-baudelaire-chat-noir/A-baudelaire-chat-noir.md, wiki/findings/theory/TH-20260724-translation-distance-axes.md, wiki/decisions/resolved/D-20260724-04-pair-relative-sense-weights.md, workshop/experiments/E-20260725-anchor-verification/critic.md, config/models.md, config/budget.md, workshop/experiments/README.md |
Frozen design (v2) — second-reader verification of the three precedent close readings
v1 was rejected by the independent critic pass (critic.md, verdict NEEDS-REDESIGN, twelve blockers). This is the rewrite. All twelve blockers and nineteen of twenty-two lesser points are implemented; the dispositions are tabulated in critic.md. A run may proceed only against this version.
No senses: field: this page evaluates no translation and invokes no goodness sense (critic G5).
1. Question
TH-20260724-translation-distance-axes states its own revision trigger #5:
Any of the three close readings is found to misread its source text. These are the lead agent's readings of texts in three languages; the Russian and French readings in particular have had no second pair of eyes.
This experiment supplies the second pair of eyes, for all three readings — the critic's blocker G1 was that v1 covered two of the three while claiming to dispose of a trigger scoped to all three. It asks:
Do the source↔translation correspondences asserted in
A-shaw-spider-thread,A-garnett-vankaandA-baudelaire-chat-noirhold against the stored texts, when checked by parties that did not write them?
It is also revisit trigger (a) of the ratified split in D-20260724-04: promoting the two-axis pair profile from optional descriptive metadata to a weighting requirement needs a second-reader check of the three close readings.
2. What is and is not under test
Under test: the textual substrate of the three anchors — that the quoted strings exist as quoted, that the linguistic descriptions attached to them are correct, and that the source↔target pairings are the pairings the pages claim.
Not under test: the generalisations. Whether C1–C4 hold of translation at large is not settleable by model calls over three short stories. A clean pass leaves C1–C4 exactly as under-powered as they were (n=1 per cell); it removes one specific way they could be baselessly under-powered.
Not evidence of quality. The panel is NOT CALIBRATED (config/models.md, RS-20260725-calibration-caseA). No result here is a quality verdict.
And not a licence to upgrade anything (critic D1). Charter §4 holds that AI-only convergence is weak evidence — QA, not validation — and that rule is not scoped to matters of taste. So agreement among verifiers, on descriptive items as much as evaluative ones, cannot upgrade a claim's status. Only two things can: the deterministic audit (stage 1) and an external non-AI reference. What the panel arms can do is fail — surface a contradiction that sends the lead back to the text. The instrument is built to be asymmetric on purpose, and §6's P-C makes the asymmetry explicit and bounded.
Shared-corpus caveat (critic B1). Lab diversity buys independence from the lead, not from the corpus. All four verifiers and the lead trained on these six texts and on the critical commonplaces about them ("the T/V slip in Vanka", "Poe's PERVERSENESS"). Against a shared commonplace, four labs are one witness. This is why the design plants held-out items (true claims about passages the anchor pages never discuss) and witness-only decoys (refutable only by reading this file): they measure whether a verifier is reading or recalling.
Item tagging (critic D2, D3, D4). Every adjudication item carries:
- settled_by: audit · external-reference · native-competence · taste — what could actually settle it, which decides how much weight the verdict may carry;
- scope: this-text · this-pair · language-general — so a claim is put at the granularity the anchor argues it, not at the granularity of its headline. (The critic showed the cost of getting this wrong: "French does not mark animacy in pronouns" is false as stated — French marks it in qui/que, in à lui/y, in relatives — while the anchor's actual claim, that third-person clitic reference to the two cats cannot carry Poe's him/it contrast, is defensible. The compressed form would have earned a correct CONTRADICTED and forced a spurious retraction.)
- evidence_mode: quote · enumerate · none — absence claims cannot be quote-evidenced, so they require an exhaustive enumeration instead, which is machine-checkable and a stronger test (critic A5).
Conjunctive assertions are split: one verdict cannot address two propositions (critic D4).
3. Materials
Six public-domain texts already stored in the repository, unchanged:
| anchor | pair | source | translation |
|---|---|---|---|
A-shaw-spider-thread |
JA→EN | workshop/canon/kumonoito/source.txt |
wiki/base/anchors/A-shaw-spider-thread/spider-thread-shaw-1930.txt |
A-garnett-vanka |
RU→EN | .../A-garnett-vanka/vanka-chekhov-1886-ru.txt |
.../A-garnett-vanka/vanka-garnett-1922.txt |
A-baudelaire-chat-noir |
EN→FR | .../A-baudelaire-chat-noir/black-cat-poe-1843.txt |
.../A-baudelaire-chat-noir/chat-noir-baudelaire-1857.txt |
All PD on both author and translator bases (verified in the anchor pages). No quarantine applies; every artifact of this run is public-tree material.
A single loader strips the provenance header (=== TEXT BEGINS === / === BEGIN TEXT ===) and is used by both stages, so a count can never be scoped one way in the audit and another way in what the verifiers see (critic F7).
The Poe witness's known corrupt line ("made to resemble the red of the cellar", where other editions read "the rest of the cellar") is pre-listed; a verifier contradiction resting on it is dispositioned WITNESS-ARTEFACT, not counted against a claim (critic F9).
4. Procedure
Stage 1 — deterministic string audit (no API, no cost)
verify_anchor_claims.py check tests every quotation the anchor pages assert against the stored text.
Normalisation folds, by named codepoint: apostrophes and quotes (U+2018/2019/201C/201D/02BC/00AB/00BB/2039/203A), dashes (U+2010/2011/2013/2014/2212), spaces (U+00A0/2009/202F/2007), …→..., markdown emphasis (**, backticks), and outer quotation furniture the page added around a truncated quotation. Italic markers _x_ are stripped except where the claim is about typography, in which case they are matched literally.
Matching rules, each fixed before the run:
- Multi-word quotations match as substrings.
- Single-token quotations in space-delimited scripts match on Unicode word boundaries — али must not match inside отбивали, gin must not match inside imagination (critic F6). CJK needles fall back to substring matching, since Japanese is written without spaces.
- Elision is forbidden inside a claim string (critic F3): an elided quotation is entered as separate fragments, each matched independently.
- Lemma citations (lemma: true) may match an inflected surface form via a prefix that is both ≥4 characters and ≥60% of the lemma, reported as ATTESTED-INFLECTED with the surface form, and every such verdict is hand-reviewed in verification.md (critic F5).
Verdicts: ATTESTED · ATTESTED-CASEFOLD · ATTESTED-INFLECTED · NOT-FOUND · COUNT-MISMATCH.
Every count assertion carries an explicit executable rule in claims.json — scope, regex, case-sensitivity — committed before the run, so no count can be reinterpreted after the number is seen (critic F7).
Stage 2a — competence screen (critic C1)
config/models.md documents Japanese competence only. This run needs Russian sub-standard morphology and 19th-century French register too. Before any payload is sent, each of the four verifiers answers a six-item screen in each of the three languages, with no text attached:
- RU — give the standard literary equivalent of вчерась, ейной, отседа, кажное, стоющие, plus one already-standard control (однако →
SAME). - FR — for hyperdiabolique, antihumain, nonobstant, maison, indicible, décharger: established standard French or not, and a register label from a fixed list.
- JA — categorise でございます, するすると, 極楽, 翡翠, 御覧になる, plus one non-Japanese control that should return
UNKNOWN.
Scored deterministically against a key fixed in the tool. Cut: ≥4 of 6. A verifier below the cut on a language is excluded from that language's anchor — the screen gates inclusion exactly as the decoy control does. The screen is published in the run record whatever it says.
Stage 2b — blind elicitation arm (critic B2)
The critic's central objection to v1 was that a SUPPORTED/CONTRADICTED format on a stated proposition is a leading question, and can only fail to contradict — it cannot discover a misreading the lead did not think to assert. The fix is an open-elicitation arm.
It is a separate call, not an earlier section of the same call: a single forward pass sees the whole prompt, so putting open questions above a claim list blinds nothing. This arm ships the two texts with no claims at all and asks five open questions per anchor — "list every second-person pronoun form the boy uses to address his grandfather"; "list every word TEXT S prints in full capitals, and quote TEXT T's rendering of each"; "list every place-name of the Buddhist afterworld and give TEXT T's rendering of each, and say which strategy the translator used". The questions are written so that the anchors' headline findings are what a correct answer would state, without stating them.
Three verifiers (P1, P3, P5) × three anchors. Scored first and independently, before the adjudication results are opened, so it cannot be contaminated by them.
Stage 2c — claim adjudication, two mutually exclusive forms
Every assertion stage 1 cannot settle goes to the four verifiers as a claim-adjudication task.
Two forms (critic A1). v1 put a decoy and the true claim it inverts in the same shuffled list, so a verifier could reject one on pure consistency grounds with zero source competence. Items are now split into form A and form B, mutually exclusive: a claim and its inversion never appear in the same call. Some true claims live in A with their inversions planted in B, and others live in B with their inversions planted in A, so neither form is "the true one". Form assignment rotates by anchor so no verifier is locked to one form.
Decoys, tiered by what refutes them (critic A2). world-knowledge decoys are refutable by anyone who has read the story; witness-only decoys are refutable only against this file (a name spelled as it is not spelled here; a closing sentence that says something else). Each form of each anchor carries 6–8 decoys, of which 2 are witness-only. P-B is gated on the witness-only tier.
True controls (critic A3, B1). Each form also carries a held-out item (a true claim about a passage the anchor pages never discuss) and a surprising-true item (a true claim that reads as implausible). Together with the decoys these give a two-armed instrument: a verifier that simply doubts everything scores high on decoys and wrecks the true controls, and the discrimination metric catches it.
Task. Per item: SUPPORTED / CONTRADICTED / NOT-DETERMINABLE, plus verbatim evidence (or an enumeration, for absence claims), which text it came from, a one-line rationale, and confidence. The system prompt states that the list contains both deliberate falsehoods and true claims that sound implausible, and that the two error directions are equally bad — v1's instruction pushed only against acquiescence and would itself have induced a doubt-everything bias.
Mechanics. Temperature 0. max_tokens 14000 (adjudicate) / 9000 (elicit) / 2500 (screen). Raw JSON preserved per call. Idempotent, and a cached response is refused if its recorded payload hash differs from the file on disk (critic E7). Every run file records the SHA-256 of claims.json; the analyzer refuses to score a run whose recorded hash differs from the file on disk (critic E6). claims.json was committed before any API call (commit 81a2fe1).
5. Metrics (pre-committed)
- Stage-1 attestation rate = ATTESTED(any) / quotation assertions, per anchor, with
NOT-FOUND-rawdistinguished fromNOT-FOUND-after-review. - Discrimination, per verifier = decoy-rejection rate − real-claim-contradiction rate. Reported alongside its two arms, the witness-only and world-knowledge rejection counts separately, and the held-out / surprising-true control outcomes.
- Abstention rate, per verifier — a first-class metric, the best single diagnostic of whether a verifier engaged with the payload (critic E2).
NOT-DETERMINABLEis excluded from metric 4's denominator and reported here instead. - Support rate per real claim, over eligible verifiers (non-flagged and screen-passed for that anchor's language), reported with and without P5 as a pre-committed sensitivity (critic C2).
- Evidence validity: each verifier's own quotation is re-checked by stage 1's matcher. A
SUPPORTEDbacked by a quotation that is not in the text is dropped from metric 4. The audit applied to the auditors. - Confidence, reported, with a low-confidence-excluded sensitivity re-analysis of metric 4 (critic E5).
6. Predictions and failure criteria (frozen before the run)
| # | prediction | what failure means |
|---|---|---|
| P-A | Stage-1 attestation ≥ 0.90 per anchor. | Below 0.90 → the anchor misquotes. Item-scoped stop rule (critic E1): a claim whose quotation is NOT-FOUND is pulled from stage 2 and corrected; the run continues on the rest. A global halt only if an anchor falls below 0.70, or a load-bearing quotation fails. |
| P-B | Per verifier: discrimination ≥ 0.50, and at least 1 of 2 witness-only decoys rejected on every anchor it is scored on. | A verifier failing either is flagged and excluded from metric 4. Metric 4 requires ≥2 eligible verifiers (critic A6); with ≤1 the arm is reported as instrument failure and corroborates nothing. Chance baseline: indiscriminate guessing over three labels gives ≈0.33 rejection and ≈0.33 false alarm, i.e. discrimination ≈0. |
| P-C | Among eligible verifiers, mean support on real descriptive claims ≥ 0.75, and no real claim carries ≥2 CONTRADICTED. | A claim with ≥2 CONTRADICTED is flagged → mandatory lead re-read against the stored text; the amendment (or its refusal, with reasons) follows the re-read, not the vote (critic D5). Pre-committed reading of the other branch (critic E3): support < 0.75 with low contradiction counts = high abstention = instrument under-powered; nothing corroborated, not "the claims are shaky". |
| P-D | On evaluative items, support is predicted to be indistinguishable from the false-alarm-adjusted baseline of P-B — i.e. verifier agreement on taste stops tracking the text. | If evaluative support is clearly higher than that baseline, that is a real instrument finding (the panel tracks something on taste items after all) and is reported as such. Either way this is a stated, falsifiable expectation, not a decoration (critic E4). |
| P-E | On the blind elicitation arm, verifiers' unprompted answers name the anchors' headline observations (the single V-form slip; the two full-capital words and their lower-case renderings; the several-way handling of the Buddhist place-names). | Failure here is the most informative outcome available: if a blind reader with the texts in front of it does not find what the anchor says is there, the anchor's salience claim is wrong even if its string claims are right. |
What would make this experiment worthless, stated in advance: if discrimination is near zero across verifiers, the instrument has no power and the run tells us only that models agree with what they are shown. That is recorded as such.
What a pass does not license: an upgrade of TH-20260724 from draft, or of the anchors from internal-judgment-only. A pass disposes of revision trigger #5 and satisfies one of the three D-04 revisit conditions. The other two — a non-J→E workshop result, and separating the near/near confound — are untouched.
7. Budget
Pre-flight estimate $1.00–1.60; hard cap $2.00, enforced before dispatch: the worst-case cost of each call (max_tokens × that model's output price + the input estimate) is added to the running total and the call is refused if it would pass the cap (critic G4). Worst case per adjudication call: P1 $0.24, P2 $0.13, P3 $0.11, P5 $0.02.
Basis: 25 calls — 12 screen (tiny), 9 elicit, 8 adjudicate — over payloads of ~13k (Vanka), ~16k (Poe/Baudelaire) and ~8k (Akutagawa/Shaw) input tokens. Today's ledger stands at $2.248649 of $5.00; this fits the $2.75 headroom. Actual per-response usage.cost is summed and ledgered.
8. Steps
- ~~Freeze v1~~ → critic pass (
critic.md, NEEDS-REDESIGN) → this v2. Re-frozen 2026-07-25. claims.jsoncommitted before any API call (commit81a2fe1); extended to three anchors and two forms and re-committed before stage 2. SHA recorded in every run file.- Stage 1 →
runs/stage1.json. Item-scoped stop rule per P-A. - Stage 2a screen → gate. 2b elicitation → scored first. 2c adjudication.
- Analysis recomputing every §5 metric from raw files →
runs/analysis.json+analysis.md. - Post-run verification (
verification.md): independent recomputation of every reported number, hand-review of everyATTESTED-INFLECTED, and a lead re-read of every flagged claim against the stored text. - Public finding
wiki/findings/results/RS-20260725-anchor-verification.md. Pre-committed flags (critic G6):analysis.mdand the result page carryinternal-judgment-only: trueandprovisional: true, since the panel is not calibrated and the synthesis is the lead's. Amend the anchors and the theory page as the flags require; record the disposition of trigger #5 either way.
9. Change log
- 2026-07-25 (S015) — v1 created,
draft; two anchors, single-form adjudication, one-sided decoy control. - 2026-07-25 (S015) — v2,
frozen, after the independent critic pass returned NEEDS-REDESIGN. Adds the third anchor (A-shaw-spider-thread, so the run matches the scope of the trigger it disposes of); splits adjudication into two mutually exclusive forms; re-tiers decoys by what refutes them and adds witness-only decoys; adds held-out and surprising-true controls and a two-armed discrimination metric; adds a gating competence screen in three languages; adds a genuinely blind open-elicitation arm as a separate call; states the charter §4 limit on descriptive convergence; addssettled_by/scope/evidence_modetags and splits conjunctive claims; makes the stop rule item-scoped; meters abstention; requires a lead re-read before any retraction; hashes and freezes the claims file; and moves the budget check before dispatch.