Repository path: workshop/experiments/E-20260725-anchor-verification/critic.md · rendered 2026-09-09
Page metadata (front matter)
| type | note |
|---|---|
| id | E-20260725-anchor-verification-critic |
| status | active |
| created | 2026-07-25 |
| updated | 2026-07-25 |
| links | workshop/experiments/E-20260725-anchor-verification/design.md |
Independent pre-run critic pass — anchor verification
Experiment discipline step 2 (experiments README). An independent adversarial critic (a separate agent instantiation, not the orchestrator that drafted the design) pressure-tested design.md v1 before freeze. Verdict: NEEDS-REDESIGN, on twelve blockers spread across the instrument (the decoy control), independence (leading-question format), language competence, the descriptive/evaluative split, the failure criteria, the string matcher, and the scope of the disposition claimed.
The critique is recorded verbatim below, followed by the disposition of every numbered point. design.md was then rewritten as v2; a run may proceed only against v2.
Two notes on reading it. (1) The critic reviewed design.md; several of its §F predictions about the matcher were made against the design's prose spec, and the implementation (tools/verify_anchor_claims.py, written in parallel) already handled some of them — where so, the disposition says which, and says how it was confirmed. (2) Its §F contains measurements of the stored files that are correct and useful independently of the design, and two of them (F5, F6) name real defects that the orchestrator's own smoke test had independently hit.
Critique (verbatim)
Independent pre-run critic pass — E-20260725-anchor-verification
I read the design, the charter, both anchor pages, the theory page,
config/models.md,config/budget.md,D-20260724-04, the experiments README, and the calibration-v1 critic pass. I also ran the Stage-1 matcher's job by hand against all four stored texts. Several points below are not hypotheses — they are measured facts about the files this experiment will run on.
A. Does the instrument actually verify anything?
A1. The decoys are defeatable by list-internal inconsistency, without reading a word of the payload — blocker. The design's own hard-decoy examples are "the reversed direction of the T/V slip" and "the reversed contraction relation of a patronymic." Those are inversions of real claims that the design also puts in the same shuffled list. A verifier that notices item 7 says ты→вы and item 23 says вы→ты rejects one of them on pure consistency grounds with zero source competence. That is not a control, it is a puzzle. Fix: a decoy and its true counterpart must never appear in the same call. Put the true claim in the Vanka call and its inversion in a separate pass, or — cheaper — construct hard decoys that are internally self-consistent and refutable only against the witness (e.g. a claim about a form that does not occur; a claim about a count).
A2. Decoy rejection measures memorisation, not reading — blocker. "The letter reaches the grandfather" is refutable by anyone who has ever encountered "Vanka," and all four texts are canonical and certainly in every verifier's pretraining. A model that never opens the payload passes the easy tier and plausibly the hard tier too (the T/V slip is a commonplace of Chekhov criticism). Fix: tier the decoys by what they require, not by difficulty:
world-knowledge-refutablevswitness-only-refutable(a false count of italic spans; a claim that a string occurs which does not; a claim about this witness's corrupt "red of the cellar" line). Report rejection separately per tier and gate P-B on the witness-only tier alone.A3. No false-alarm arm; decoy-rejection rate alone cannot separate discrimination from a blanket-CONTRADICT bias — blocker. Metric 2 is a one-sided hit rate. A verifier with a 40% indiscriminate CONTRADICTED rate clears 0.60 on four hard decoys roughly one time in five, and simultaneously wrecks metric 3 on real claims — and the design would read that as "one flagged claim needs retraction." The instrument has no way to tell a good detector from a grumpy one. Fix: add planted-true-but-surprising items (true assertions that read as implausible — e.g. that Garnett writes "grouse and woodcocks and fish and hares" where Chekhov has three birds and no fish; that Poe's narrator calls the second cat "it"), and score discrimination as a 2×2 (decoy-hit rate vs real-claim false-alarm rate), reporting the difference. Without both arms P-B is decorative.
A4. The 0.60 threshold is arbitrary and its denominator does not exist yet — blocker. §4 says "six planted false claims… in two tiers" but never fixes how many are hard. P-B is "pooled hard-decoy rejection ≥ 0.60 per verifier." With 4 hard items that means ≥3/4 (0.75); with 6 it means ≥4/6 (0.667). The stated number is unreachable exactly and the tier split is unspecified, so the gate is set after the denominator is known. There is also no stated null: over three labels, indiscriminate guessing yields 0.33, so 0.60 on 4–6 items is roughly two lucky calls from chance. Fix: fix tier counts before freeze; restate P-B as an integer cut on a named denominator (e.g. "≥5 of 6 decoys, including ≥2 of 3 witness-only"); state the chance baseline and require the binomial lower bound to clear it.
A5. Metric 5 (evidence-quotation re-check) is the best idea in the design and is currently unimplementable for a large fraction of items — should-fix. The task demands "a verbatim quotation from the named text" for every item, but the highest-stakes claims are absence claims: English has one second person; French has no animacy contrast in third-person clitics; кобелек is dropped. There is no quotation that evidences an absence. Verifiers will supply an unrelated span or fabricate, and metric 5 will downgrade sound verdicts to UNEVIDENCED. Fix: pre-tag each item as quote-evidenceable or not; for absence claims require instead an exhaustive enumeration ("list every second-person pronoun form in the letter") which is machine-checkable and is a stronger check than a quote.
A6. "Instrument failure" has no reachable middle branch — should-fix. P-B defines the all-four-flagged case. If exactly one verifier survives, metric 3 becomes one model's opinion and nothing in the design forbids reporting it as "convergent corroboration." Fix: require ≥2 non-flagged verifiers for metric 3 to be computed at all; n≤1 → instrument failure.
B. Circularity and independence
B1. Non-Anthropic ≠ independent for this failure mode — should-fix (charter-grounded). The exclusion in
config/models.mdtargets shared priors between the lead and the panel. The threat here is different: the lead and all four verifiers were trained on the same corpus, which contains these four texts and the secondary commentary about them. "The T/V slip in Vanka," "Poe's PERVERSENESS," "Baudelaire's Poe as appropriation" are critical commonplaces. Lab diversity buys nothing against a shared commonplace. Fix: say so in §2 "Not under test," and add a held-out control: two or three true claims about passages the anchor pages never discuss. If verdicts on held-out items are as confident as on famous ones, the verifiers are reading; if not, they are recalling.B2. Showing the claim removes the design's only chance of discovering an error — blocker. SUPPORTED/CONTRADICTED on a stated proposition is a leading-question format, and the sole mitigation is a prompt instruction ("your job is to find them"), which is the weakest control available and precisely what acquiescence-bias work says fails. Structurally, this instrument can only fail to contradict; it cannot surface a misreading the lead did not think to assert. Fix, and it is nearly free because the payload dominates token cost: add an open-elicitation arm for the 8–12 highest-stakes items in the same call, asked before the claim list: "List every second-person pronoun form in Vanka's letter and flag any anomaly." · "Which of the two cats does Poe's narrator refer to as it, and where?" · "How many words in this text are set in full capitals? List them." · "Which words does Poe mark with italics as foreign?" Score the free answers against the anchor. That arm is genuinely blind, and it is the only part of this experiment that could produce a finding rather than a non-refutation.
B3. A cheaper, genuinely external check exists for the subset that matters most and is not considered — should-fix. Charter §4 pre-labels AI-only convergence as "QA, not validation." Several "descriptive" items are settleable at zero cost by non-AI external reference: whether Макарыч is a contracted patronymic, whether French third-person clitics encode animacy, whether décharger carries a register label — all are reference-grammar / dictionary facts reachable by WebFetch. Spending $0.60–1.20 to obtain weak-by-charter AI convergence on items an external source settles outright is bad allocation. Fix: route reference-settleable items to an external lookup (logged in
wiki/base/consulted.md), and spend the API budget only on items where no external source can adjudicate.
C. Language competence
C1. No competence screen for either language under test — blocker.
config/models.mddocuments Japanese only: three passages, one call each. The competences this run actually requires are 19th-century Russian sub-standard peasant morphology (вчерась, ейной, отседа, кажное, стоющие) and Second-Empire literary French register. The design offers P-B as the screen, but per A2 decoy rejection is satisfiable from memory. A clean result would read "four verifiers corroborated the Russian reading" when the supportable statement is "four models of undocumented Russian competence agreed with a claim shown to them." Fix — cheap and adequate: one short call per model per language, no payload, ~$0.02–0.05 total. Russian: give five sub-standard forms lifted from the stored text and require the standard equivalent plus a gloss. French: give five items from the Baudelaire and require a register label (courant / soutenu / littéraire) with a known-elevated control word to catch flattery. Score it, publish it in the run record, and gate verifier inclusion on it exactly as P-B gates on decoys. This is a rounding error against a $1.50 hard stop and it is the difference between a competence claim and an assumption.C2. P5 is included at full weight against the panel's own most recent instrument finding — should-fix.
config/models.mdrecords, fromRS-20260725-calibration-caseA(yesterday's run), thatdeepseek-v4-proshowed a 0.44 order-flip rate and should be "down-weight[ed] until reps increase." The design cites the S010/S014 cost lesson to exclude kimi and is silent on the same page's instrument lesson about P5. With kimi gone the panel is P1/P2/P3 (all US) + P5 (order-unstable) — the lab-diversity argument thins exactly where independence is load-bearing. Fix: either exclude P5, down-weight it explicitly in metric 3, or substitute the documented first reserveqwen/qwen3.7-max; and state the resulting 3-US/1-other composition as a limitation.
D. The descriptive/evaluative split
D1. The split as written is a side door around charter §4 — blocker (charter-forced fix). §2 says convergent agreement on descriptive items is "corroboration of a factual claim, which is not a calibration-gated judgment." Charter §4's rule is not scoped to taste: "AI-only convergence is weak evidence. … agreement among them is QA, not validation." Calibration is the gate on jury verdicts; §4 is a separate and broader constraint that the descriptive route sidesteps. Fix: keep the split for scoring, but state that descriptive convergence is likewise QA-grade AI convergence, and that the only things that can upgrade a descriptive item are (i) the deterministic audit or (ii) an external non-AI reference. One paragraph, and it removes the loophole.
D2. The two-way tag is drawn by convenience; a four-way tag is the principled cut — should-fix. "Is décharger register-flatter than unburthen" is filed as evaluative, but it is largely lexicographic (Robert/Littré vs OED register labels) and externally settleable. "Does French mark animacy in pronouns" is filed as descriptive but is a typological generalisation checkable against no stored text. Fix: tag by what could settle it —
audit-settleable/external-reference-settleable/native-competence/taste— and route each tag to its proper adjudicator.D3. At least one named "descriptive" item is false as compressed and will generate a spurious retraction — blocker. "Does French mark animacy in pronouns" is, as stated, wrong: French marks animacy in the interrogative (qui vs que/quoi), in oblique pronominalisation (je pense à lui vs *j'y pense), and in relatives. The anchor's actual claim is narrower and defensible — that le chat is il* regardless, so Poe's him/it contrast has no third-person-clitic counterpart. A competent French-reading verifier should CONTRADICT the compressed form, and P-C then fires ("≥2 CONTRADICTED → the anchor and the theory claim resting on it are amended or retracted in the same session"), forcing an amendment for a wholly artefactual reason on the sole evidence for C1's most important row. Fix: claims must be stated at the granularity the anchor argues** them, not the granularity of its headline, and every claim must carry an explicit scope field (
this-text/this-pair/language-general).D4. Bundled assertions yield uninterpretable single verdicts — should-fix. The anchor's patronymic entry asserts two things at once: Макарыч is the colloquially contracted form of Макарович (descriptive, true), and "the informality it signals is invisible" in Garnett (evaluative). One SUPPORTED/CONTRADICTED cannot address both. Fix: mechanically split every conjunctive assertion in
claims.json; a claim with an "and"/"but" clause is two claims.D5. P-C's retraction rule gives uncalibrated disagreement power that uncalibrated agreement is denied — should-fix. Agreement cannot upgrade a claim; disagreement can retract one. That asymmetry is defensible (falsification is cheap) only if the contradiction is itself verified. Step 6 mentions a re-read of flagged claims, but P-C does not. Fix: restate P-C's trigger as "≥2 CONTRADICTED → mandatory lead re-read against the stored text; the amendment follows the re-read, not the vote."
E. Predictions and failure criteria
E1. The step-3 stop rule is a trip-wire on a normalisation bug, and it will fire — blocker. Per §F below, the current normalisation spec will produce spurious NOT-FOUNDs on the Baudelaire anchor from the apostrophe character alone. P-A then fails, and step 3 halts the entire run before stage 2. Fix: (a) run the normaliser over both anchor pages as a dry run before freeze, fix normalisation, then freeze P-A; (b) make the stop rule item-scoped: quotation-dependent claims whose quotations are NOT-FOUND are pulled from the stage-2 list and the run proceeds on the rest; a global halt only if attestation is catastrophic (say <0.70) or a load-bearing quotation fails. (c) P-A must distinguish
NOT-FOUND-rawfromNOT-FOUND-after-manual-review— only the latter is evidence that the lead misquoted.E2. NOT-DETERMINABLE is unmetered and can move P-C either side of 0.75 by definition alone — blocker. Metric 3's denominator is "verifiers returning a verdict." If abstention counts, support rate is depressed by caution; if it doesn't, a verifier that abstains on everything hard scores perfectly on the residue. Fix: define it explicitly, and report abstention rate as a first-class metric — it is the single best diagnostic of whether the verifiers are engaging with the payload.
E3. P-C's first conjunct feeds no decision — should-fix. "Mean support ≥ 0.75 on descriptive claims" has a consequence stated only for the second conjunct (≥2 CONTRADICTED). If mean support lands at 0.55 with no claim reaching two contradictions, the design says nothing happens. Fix: pre-commit the reading — most likely "support < 0.75 with low contradiction counts = high abstention = instrument under-powered; nothing is corroborated" — so it cannot be interpreted after the fact.
E4. P-D is not a prediction, and the evaluative items are measured while deciding nothing — should-fix (repeat defect). The calibration-v1 critic flagged exactly this shape twice (B3 "cannot fail in any defined way"; F3 "collected but unused") and both fixes were accepted. Reintroducing it is a process regression. Fix: either drop evaluative items from the payload (saving tokens) or give them a falsifiable job — e.g. predict evaluative support is indistinguishable from the false-alarm-adjusted baseline from A3, which would be a real instrument finding about where model agreement stops tracking text.
E5. The
confidencefield feeds no metric — minor (repeat of calibration-v1 F3). Fix: as dispositioned there — report it as a diagnostic and run a low-confidence-excluded sensitivity re-analysis of metric 3.E6. The claims file — the actual instrument — is authored after the critic pass and is never frozen — blocker. Step 1 freezes
design.md; step 2 authorsclaims.json. Which assertions are tested, which are tagged descriptive, and which are decoys all live in step 2, unreviewed. §2's promise that the split is "fixed before the run and not revised after seeing results" is enforced by nothing. Fix:claims.jsonis committed in its own commit before any API call, its SHA recorded in every run file, and the analyzer refuses to score a run whose recorded claims-hash differs from the file on disk. Ideally the critic (or a second pass) sees the claims file, since that is where the experiment actually lives.E7. Idempotence keyed on (verifier × anchor) will silently reuse stale responses — should-fix. "An existing output file is not re-requested" is correct until
claims.jsonchanges; then a re-run scores old answers against a new claim list. Fix: include the payload hash in the filename or refuse reuse on hash mismatch.
F. Stage 1 — the deterministic string audit
I ran the audit's job by hand. These are measurements, not conjectures.
F1. The apostrophe alone will fail every French quotation containing one — blocker.
chat-noir-baudelaire-1857.txtcontains 277 U+2019 (’) and zero ASCII'.A-baudelaire-chat-noir.mdcontains 62 ASCII'and zero U+2019. So"l'homme naturel","l'épouse de mon cœur","l'Archidémon","je n'essaierai pas de les élucider"— every one — returns NOT-FOUND. §4's "straight/curly quote variants folded" is at best ambiguous about the apostrophe use of U+2019. This single omission is enough to sink P-A on the Baudelaire anchor and halt the run at step 3. Fix: fold U+2019/U+02BC/U+2018 →'on both sides; enumerate the exact codepoint classes in the spec, not a prose gloss.F2. Markdown emphasis leaks into quoted strings — blocker.
A-garnett-vanka.mdcarries 268**markers, many inside quotations:«…И пишу **тебе** письмо. Поздравляю **вас** с Рождеством…»,"I had a **wigging**","a **forty-pound** sheat-fish","the **phantasm** of the cat". Stage-1 normalisation lists NFC, quotes, hyphens, whitespace, and_x_italics — not**. Every bolded quotation → false NOT-FOUND, including the headline T/V evidence. Fix: strip markdown emphasis (**,*, backticks) from claim strings while preserving_x_where the claim is about typography; the two conventions collide and the spec must say which wins.F3. Ellipsis-elided quotations cannot be matched and cannot be distinguished from authorial ellipsis — blocker. The Russian file contains 16 authorial
…. The anchor also uses…for editorial elision: «Теперь, наверно, дед стоит у ворот… Бабы нюхают и чихают… А погода великолепная.» I verified all three fragments occur — the concatenation does not. The matcher has no way to tell Chekhov's…from the lead's. Fix: forbid elision inside claim strings; inclaims.json, an elided quotation becomes an ordered conjunction of sub-strings, each matched independently, with an optional order/adjacency constraint. Separately: Garnett uses spaced. . .(12 occurrences) and zero…, so ellipsis normalisation must be per-file, not global.F4. Truncated quotations carry an editorial closing quote mark the text does not have — blocker. The Garnett file reads
…all blessings from\nGod Almighty. I have neither father nor mother…; the anchor quotes it as ending…God Almighty.". That terminal"is the lead's, not Garnett's — exact match fails on the anchor's single most load-bearing English quotation. The Russian«…»framing around the letter opening is likewise editorial. Fix: strip outer quotation furniture (" ",« »,“ ”) from claim strings before matching, and record the stripping in the audit output so it is auditable.F5. Lemma-vs-inflected: a confirmed false NOT-FOUND on a theory-bearing item — blocker.
Ванюшкаoccurs zero times invanka-chekhov-1886-ru.txt; the text has the instrumentalВанюшкой(«посмеивается над озябшим Ванюшкой»). The anchor's diminutive table lists the lemma, and the theory page cites Ванюшка twice — in C1's table and in "the Иван/Ванька/Ванюшка gradient levelled to two strings." Exact substring matching marks a true and load-bearing observation as NOT-FOUND. (Garnett does render it "laugh at frozen Vanka," so the reading is sound.) Fix: claims about inflected languages must quote the surface form as it occurs, with the lemma recorded in a separate field; the audit matches the surface form, and a lemma-only claim is aSPEC-ERROR, not aNOT-FOUND.F6. Naive substring matching produces confirmed false ATTESTEDs in both inflected languages — blocker. Measured, naive-substring vs Unicode word-boundary: Russian
али7 vs 1 (matches inside отбивали, вешали, заливается);образ2 vs 1 (matches inside вообразил);ты4 vs 2 (matches inside тыкать). Frenchgin8 vs 2 — six spurious hits inside imaginai, originelle, imagination, imaginaires, imaginables, imaginer, so the anchor's "Baudelaire keeps gin" is attested six times for the wrong reason and any count claim on it is wrong. Fix: Unicode-aware word-boundary matching ((?<!\w)…(?!\w)withre.UNICODE) for single-token claims; whole-string matching only for multi-word quotations.F7. Every count claim in §4 is determined by an unstated scope/tokenisation decision, and the provenance header is stripped only for Stage 2 — blocker. Measured on the stored files: - "Poe's only two full-capital words": the file yields 13 (header adds PROVENANCE, HEADER, PUBLIC, DOMAIN, TEXT×2, WITNESS, CAVEAT, BEGINS); the body after
=== TEXT BEGINS ===yields 5 (the title lineTHE BLACK CAT.contributes THE, BLACK, CAT); only "prose body, excluding the title" yields the claimed 2. §3 says headers are stripped "before they are sent" — that is Stage 2. Stage 1 is specified as matching "against the stored text." - "nineteen italic spans": the Poe body has exactly 19_…_spans — but the file has 20, because the header's own explainer line contains_word_. Two of the 19 span a newline and one is_a brute beast _with a trailing space inside the markers, so a_[^_\n]+_regex and a newline-permitting one return different numbers. - "three occurrences of Gentlemen": the French hasGentlemen×1 andgentlemen×2. A case-sensitive count returns COUNT-MISMATCH (1≠3) on a claim the anchor states correctly.Fix: (a) strip the provenance header in Stage 1 too, from a single shared
load_text()used by both stages; (b) every count claim inclaims.jsoncarries an explicit, executable counting rule — scope (body/excluding-title), tokenisation regex, case-sensitivity — committed before the run so no count can be reinterpreted after seeing the number; (c) report both raw and rule-applied counts.F8. NBSP in the French is unhandled by name — should-fix. The French body contains 64 U+00A0, systematically before
!,?,;and after«(gibet\xa0!,cœur\xa0!,gentlemen\xa0?). The anchor writes these with ASCII spaces. Python's\sfolds\xa0in Unicode mode, so are.sub(r'\s+',' ')implementation survives — but the spec should not depend on an implementation accident. Fix: name U+00A0/U+202F/U+2009 explicitly in the normalisation list.F9. The corrupt witness line has no verdict category — should-fix. Stage 2 ships the whole Poe text including "made to resemble the red of the cellar," and the whole French including "le reste de la cave" (present, ×1). A verifier may reasonably flag a contradiction that is a witness artefact the anchor already documents. Fix: add a
WITNESS-ARTEFACTdisposition to the analyzer and pre-list the known corrupt loci.
G. Scope and value for money
G1. The design covers two of the three readings but claims to dispose of a trigger scoped to all three — blocker. TH trigger #5 reads "any of the three close readings is found to misread its source text"; D-20260724-04's revisit condition (a) reads "a second-reader/panel check of the three close readings."
A-shaw-spider-threadis absent from this design entirely — and it is the anchor carrying the most downstream load (origin of C1, C4, the handling typology, the scaffold gradient). Yet §6 states a pass "removes revision trigger #5 as an open risk" and "satisfies one of the three D-04 revisit conditions." As written the run would overclaim its own disposition. Fix: add the Shaw anchor — at ~1,000 words it is the cheapest of the three payloads and would raise the run by roughly a third — or restate the disposition throughout as partial, with trigger #5 and D-04(a) explicitly left open.G2. Stage 1 is excellent value and should not be held hostage to Stage 2 — should-fix. The audit is free, reproducible, and — on the evidence in §F — will surface real defects in the anchor pages (quotation truncation, lemma-vs-inflected, count-scope). Fix: run Stage 1 (after the F-fixes) plus the ~$0.05 competence screen from C1 as a standalone unit, land the anchor corrections, and only then decide whether Stage 2 is worth $1.20.
G3. As currently scoped, Stage 2 buys the weakest evidence class the charter recognises, at the highest price in the design — should-fix. Four models with no documented competence in either language, on texts they have memorised, adjudicating claims shown to them, against a decoy control defeatable by list-internal consistency, producing AI-only convergence that charter §4 calls "QA, not validation." Fix: re-scope Stage 2 to the two things it can uniquely do — the open-elicitation arm (B2) and the two-armed discrimination control (A3) — on 8–12 high-stakes items. Same dollars, a check that can actually fail.
G4. The $1.50 abort is checked after the overspend — should-fix. Both prior lessons (S010: est. $1.9–3.0 → actual $4.49; S014: $0.90 → $1.83; $0.25 → $0.40) are reasoning-token overruns. With
max_tokens16000 on four reasoning models over 13–18k-token payloads, a single grok/gemini call can be large, and "prints running cost and aborts above it" trips after the call that breaks the ceiling. Fix: checkrunning_total + worst_case_next_callbefore dispatching each call; state the worst-case per-call figure in §7.G5. Front matter lists five goodness senses on a page that evaluates none — minor (repeat of calibration-v1 G1). §2 states nothing here is a quality verdict and evaluative items change no status. Fix: trim
senses:to what a prediction actually invokes, or drop the field with a note.G6. The design does not pre-commit the flags its own outputs must carry — minor.
config/models.mdrecordsNOT CALIBRATED; CLAUDE.md requiresprovisional: trueon assessments made before calibration passes, andinternal-judgment-onlyon unanchored judgment.analysis.mdandRS-20260725-anchor-verificationwill need both. Fix: state it in §8 now, so it is not decided after the numbers are in.
Verdict: NEEDS-REDESIGN — the Stage-1 normalisation will spuriously fail P-A and halt the run at step 3 (F1–F7), the decoy control cannot discriminate reading from memory or detection from bias (A1–A4), no language competence is established for either language under test (C1), one named "descriptive" item will force a spurious retraction (D3), and the scope does not match the trigger it claims to dispose of (G1); these require new items, a new task arm, and a rewritten matcher spec, not tweaks.
Disposition
All twelve blockers accepted. Nineteen of the twenty-two should-fix/minor points accepted; three modified; one rejected with reason. design.md v2 implements them.
| point | disposition |
|---|---|
| A1 decoy/true pair in one call | ACCEPTED. Claims are split into two mutually exclusive forms. Every decoy that inverts a real claim is placed in the form that does not carry its twin; each verifier sees one form per anchor, two verifiers per form. No verifier can solve a decoy by list-internal contradiction. |
| A2 decoys refutable from memory | ACCEPTED. Decoys are re-tiered by what refutes them: world-knowledge vs witness-only (a string asserted to occur that does not; a false count; a claim about this witness's documented corrupt line). P-B is gated on the witness-only tier. |
| A3 no false-alarm arm | ACCEPTED. Added planted-true-but-surprising items, and metric 2 becomes a two-armed discrimination score: decoy-rejection rate − real-claim-contradiction rate. A blanket-CONTRADICT verifier scores ≈0 and is flagged. |
| A4 arbitrary threshold, floating denominator | ACCEPTED. Tier counts fixed in claims.json before the run; P-B restated as an integer cut on a named denominator with the chance baseline stated. |
| A5 absence claims cannot be quote-evidenced | ACCEPTED. Each item carries evidence_mode: quote / enumerate / none. Absence claims require an enumeration, which the analyzer checks. |
| A6 no middle branch for instrument failure | ACCEPTED. Metric 3 is computed only with ≥2 non-flagged verifiers; n≤1 → instrument failure. |
| B1 shared corpus ≠ independence | ACCEPTED. Stated in §2, and held-out control items added — true claims about passages the anchor pages never discuss. |
| B2 leading-question format | ACCEPTED, with a correction to the proposed fix. An open-elicitation arm is added, but as a separate call, not "earlier in the same call": a single forward pass sees the whole prompt, so an in-call ordering blinds nothing. The elicitation arm ships the texts with no claims and asks open questions; it is the only genuinely blind arm and it is scored first. |
| B3 external non-AI reference is cheaper | ACCEPTED in part. Reference-settleable items are looked up externally (free, WebFetch) and their answers recorded as the adjudicating authority, with the panel verdict reported beside them as a cross-check rather than as the evidence. Not adopted as a replacement for the panel arm on those items, because the cross-check is itself informative about the instrument. |
| C1 no competence screen | ACCEPTED. A per-model, per-language screen (RU, FR, JA) runs before the main arms and gates inclusion. Scored and published in the run record. |
| C2 P5 order-instability | ACCEPTED in modified form. P5 is retained (the S014 finding concerns pairwise preference under display order, a different task), but metric 3 is reported with and without P5 as a pre-committed sensitivity, and the 3-US/1-CN composition is stated as a limitation. |
| D1 descriptive route bypasses charter §4 | ACCEPTED. §2 now states that descriptive convergence is equally AI-only convergence, QA-grade; only the deterministic audit or an external non-AI reference can upgrade an item. |
| D2 two-way tag too coarse | ACCEPTED. Four-way settled_by tag: audit / external-reference / native-competence / taste. |
| D3 compressed animacy claim is false as stated | ACCEPTED — and it was a real defect. The claim is restated at the granularity the anchor argues (third-person clitic reference to the two cats), and every item carries an explicit scope field. |
| D4 bundled assertions | ACCEPTED. Conjunctive items split. |
| D5 disagreement given power agreement lacks | ACCEPTED. P-C's trigger now requires a lead re-read against the stored text; the amendment follows the re-read, not the vote. |
| E1 stop rule trips on a matcher bug | ACCEPTED. Stop rule made item-scoped; NOT-FOUND-raw vs NOT-FOUND-after-review distinguished; global halt only below 0.70 or on a load-bearing quotation. The dry run the critic asks for had already been performed (see below). |
| E2 NOT-DETERMINABLE unmetered | ACCEPTED. Abstention is excluded from metric 3's denominator and reported as a first-class metric. |
| E3 P-C's first conjunct decides nothing | ACCEPTED. Pre-committed reading written into P-C. |
| E4 P-D cannot fail (process regression) | ACCEPTED. Evaluative items now carry a falsifiable prediction against the A3 discrimination baseline. |
| E5 confidence unused | ACCEPTED. Reported, plus a low-confidence-excluded sensitivity re-analysis. |
| E6 claims file unfrozen | ACCEPTED. claims.json is committed before any API call (commit 81a2fe1, pre-dating every call), its SHA-256 is recorded in every run file, and the analyzer refuses to score a run whose recorded hash differs from the file on disk. |
| E7 stale idempotence | ACCEPTED. Payload hash recorded and reuse refused on mismatch. |
| F1 apostrophe folding | ALREADY HANDLED — confirmed. The implementation folds U+2019/U+2018/U+201C/U+201D/«/» before matching; all four apostrophe-bearing French quotations returned ATTESTED in the dry run. Spec updated to enumerate the codepoints as asked. |
| F2 markdown emphasis in quotations | ALREADY HANDLED by construction — spec updated. Claim strings were transcribed with ** stripped; the spec now says so and the matcher strips them defensively. |
| F3 ellipsis elision | ALREADY HANDLED by construction — spec updated. Elided quotations were entered as separate fragments (e.g. the vivid-present passage as VQ40/41/42), never as concatenations. The spec now forbids elision inside a claim string. |
| F4 editorial closing quote | ALREADY HANDLED — confirmed. The load-bearing Garnett quotation (VQ50) was transcribed with the internal quotation structure Garnett's text actually has and returned ATTESTED. Outer furniture stripping added defensively. |
| F5 lemma vs inflected | ACCEPTED — independently hit and fixed. The dry run reproduced exactly this on Ванюшка→Ванюшкой. Lemma claims now carry lemma: true, are matched by a prefix rule (≥4 chars and ≥60% of the lemma), report ATTESTED-INFLECTED with the surface form, and every such verdict is hand-reviewed in verification.md. The flat truncation the first implementation used produced a false ATTESTED (табачку→табакерку) and a false NOT-FOUND (мамка→мамку); the 60% rule separates them. |
| F6 substring vs word boundary | ACCEPTED — real remaining defect. Single-token claims now match on Unicode word boundaries. али and образ were ATTESTED under the old matcher and needed re-checking. |
| F7 count scope/tokenisation | ACCEPTED in part; measurement partly superseded. The implementation already strips the provenance header in stage 1 (a single body() used by both stages) and already excludes the title line, so the observed counts were 2 full-capital words, 19 italic spans, 3 Gentlemen — i.e. the anchors' figures, not the critic's file-level 13/20/1. Accepted in full: every count claim now carries an explicit executable counting rule in claims.json (scope, regex, case-sensitivity), committed before the run. |
| F8 NBSP by name | ACCEPTED. U+00A0/U+202F/U+2009 named in the normalisation list. |
| F9 witness artefact | ACCEPTED. WITNESS-ARTEFACT disposition added; the known corrupt locus is pre-listed. |
| G1 scope vs the trigger it claims to dispose of | ACCEPTED. A-shaw-spider-thread is added as a third anchor (its translation is stored in the anchor directory, its PD Japanese source in workshop/canon/kumonoito/source.txt), so the run covers all three readings the trigger names. It is also the cheapest payload and the one pair for which panel competence is documented. |
| G2 stage 1 as a standalone unit | ACCEPTED in practice. Stage 1 was run and landed before any API call; its findings stand on their own whatever stage 2 returns. |
| G3 re-scope stage 2 to elicitation + discrimination only | REJECTED in part, with reason. The elicitation arm and the two-armed control are both adopted (B2, A3). Dropping the adjudication arm is not: the elicitation arm can only probe what the lead thought to ask openly, and the anchors' specific assertions — which are what trigger #5 is about — are testable only by putting them. Both arms run; the elicitation arm is scored first and independently, so it cannot be contaminated by the adjudication result. |
| G4 budget check after the spend | ACCEPTED. Worst-case next-call cost is computed from max_tokens × the model's output price and checked before dispatch. |
| G5 senses front matter | ACCEPTED. Trimmed. |
| G6 output flags not pre-committed | ACCEPTED. §8 now pre-commits provisional: true and internal-judgment-only on the analysis and result pages. |
Re-freeze: design.md v2, 2026-07-25. Stage 1 was re-run against the v2 matcher after these changes; its v1 numbers are superseded and both are reported in verification.md.