Repository path: workshop/experiments/E-20260731g-provenance/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260731g-provenance |
| status | frozen |
| created | 2026-07-31 |
| updated | 2026-07-31 |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-figure-audit.md, wiki/method-notes.md, workshop/regimes/R10-anchored-register.md, workshop/experiments/E-20260731f-catalogue-reach/design.md, wiki/findings/results/RS-20260731f-catalogue-reach.md, wiki/base/anchors/A-mchugh-presence/A-mchugh-presence.md, wiki/base/anchors/A-mansfield-garden-party/A-mansfield-garden-party.md, config/models.md |
Design — what a stored figure's provenance fixes, and what it does not
Frozen before any API call and before a word of the translation limb was written.
ARM-figure-audit step 2, T3. Session S075.
0. The unit and its wire
Study limb. ARM-figure-audit's condition 2 — note (bdt), the <tag>.raw verifier
defect — checked across every frozen verifier and every stored response body in the archive.
Translation limb. A matched R10 pair on a new source, replicating the single
observation RS-20260731f-catalogue-reach returned as a by-product: that aiming a translation
at a period register roughly doubled its measured overlap with a period published
translation the translator had not read.
The wire, in one sentence. The study limb asks whether a published figure was computed against the body the run actually accepted; the translation limb takes the project's newest published figure — S074's register/contamination crossover — and asks whether it survives a length control and a second source; so one session tests the project's numbers both for "was it computed on the right text?" and for "does it survive being computed again?"
The translation limb is not material for the study limb and is not presented as such. It is the
same audit — ARM-figure-audit's own direction, in the direction that can only cost the project
something — applied to a different class of published figure. The arm's constraint that "a
figure that survives the check is not thereby validated" is what the translation limb makes
concrete: a contamination figure can have perfect body-provenance and still not be a well-defined
quantity, because the condition that determines it is recorded nowhere.
PART A — the study limb
A1. Question
Note (bdt), S057: "A verifier that opens <tag>.raw reads the REJECTED body wherever the
declared reserve chain fired." It fired for real at S062 and changed a number (note (beo)). The
arm's condition 2 is that this be checked on every frozen verifier, and that any figure that
moves be corrected in place.
Three questions, in order:
- A-Q1 (the bodies). For every stored response body in
workshop/, is the body that is the run's answer of record complete —finish_reason == "stop", non-null content? Every body that is not must be classifiable as (a) superseded by a stored successor that is complete, (b) declared as dropped on the experiment's own pages, or (c) silently the answer of record, which is the live defect. - A-Q2 (the verifiers). For every frozen verifier that opens a
.raw, does it resolve the accepted body by walking the reserve/attempt chain and assertingfinish_reason, or by filename alone? - A-Q3 (the power). If a rejected body were substituted for its accepted counterpart, would
the verifier notice? This is the question stored outputs cannot answer and it is answered by
mutation: for each verifier that reads
.raw, substitute a truncated body and re-run.
A-Q3 is the point of the step. A-Q1 and A-Q2 are a census; without A-Q3 a clean census means only that the defect did not fire, not that the check has any power. S070 hit exactly this on condition 1, where the condition's own premise turned out to be arithmetically impossible.
A2. Materials
Every file under workshop/ matching *.raw, *.json written by a call.py, and every
verify*.py. Nothing is fetched and nothing is re-dispatched. No API call is made by Part A.
A3. Procedure
analysis/sweep.pyenumerates every.rawinworkshop/, parses it, and recordsfinish_reason, content nullity and length, reasoning length, model, provider and cost.- For each defective body it looks for a successor by the two naming conventions in use
(
<tag>-reserveN.raw,<tag>.attemptN.raw) and classifies the run cell as SUPERSEDED, DROPPED-DECLARED, or ACCEPTED-DEFECTIVE. ACCEPTED-DEFECTIVEcells are checked by hand against the experiment's own design, result page and ledger row. A cell whose defect is declared on those pages is reclassifiedDROPPED-DECLAREDwith the citation; a cell whose defect is not declared anywhere is a finding, and the figure that rests on it is recomputed or retracted.analysis/verifier_census.pyclassifies everyverify*.pyby whether it opens a.raw, whether it walks a chain, and whether it assertsfinish_reason.- The mutation. For each verifier that opens a
.rawand whose experiment has both an accepted and a rejected body on disk, the rejected body is copied over the accepted one in a scratch copy of the experiment directory and the verifier is re-run. A verifier that still reports zero failures is insensitive to body identity, which is a stronger defect than (bdt) and is reported as one. Where no rejected body exists, one is synthesised by truncating the accepted body's content to 40% and settingfinish_reasontolength— the shape (bdt) names as the dangerous one, which the archive does not contain.
A4. Predictions, frozen
- A-P1. At least one stored body is
ACCEPTED-DEFECTIVEand undeclared. Fails if every defective body in the archive is either superseded or declared. - A-P2. At least one frozen verifier opens
<tag>.rawby filename alone in an experiment where a chain fired. Fails if none — which would mean condition 2 is not live anywhere and the arm should say so. - A-P3. Every verifier that reads a
.rawcatches the substituted body. Fails for any verifier that reports zero failures on a mutated body. This is the prediction the step exists for and it is registered in the direction that costs the project something: a failure here means the project's verification discipline has been checking that a file exists rather than that a number is right.
A5. Failure criterion for Part A
If the mutation cannot be run on a majority of the .raw-reading verifiers — because they import
paths that no longer resolve, or because copying a body into a scratch directory changes what they
read — Part A reports that condition 2 is not runnable without rebuilding the runs it audits,
which ARM-figure-audit names in advance as a retired ending and as itself the finding.
PART B — the translation limb
B1. Question
RS-20260731f-catalogue-reach reports, as a by-product it did not set out to measure:
the period-targeted arm shares 11 tokens and 8 seven-grams with the published English, the centre-targeted arm 5 and 0. Same translator, same source, neither having read it.
Three things about that observation are untested, and all three are free to test.
- B-Q1 (replication). Does it reproduce on a different source, a different source language, a different author and a different comparator translator?
- B-Q2 (length). S074's two arms are 863 and 968 tokens — the period arm is 12.2% longer, and a longer text has more opportunity to contain any given run. The published figure is not length-controlled. Does it survive one?
- B-Q3 (the floor). What does a non-translating arm score? The pre-translation recall probe run for this unit's material gate (§B3.1) is a declared non-recollection, and it already scores 8 tokens and 5 shared 7-grams against the same comparator. A crossover between two arms is only interpretable against that floor.
B2. Materials
- Source. Gabriele d'Annunzio, «La fine di Candia», San Pantaleone (Firenze, Barbera, 1886),
paragraphs 1–24 of 118, 800 words — matched to
E-20260731f's Kielling unit of 799 to within one word. Italian Wikisource proofread transcription, fetched 2026-07-31 via the MediaWiki API, frozen atmaterial/source-unit.txt. Public domain (d. 1938; first published 1886). - Comparator. "The End of Candia", anonymous English, Short Story Classics (Foreign), vol. 2, Italian and Scandinavian, ed. William Patten, P. F. Collier & Son, 1907, Gutenberg #70578 — the same volume, the same year and the same publisher as S074's «Karen» comparator, by a different translator. 3,426 words, the whole story. Extracted to a file by script by line number and not read by the translator; the extraction printed only line and word counts.
- Targets. The two frozen proposition sets of
E-20260731f/claims.json, reused verbatim and unaltered:A-mchugh-presenceC01–C19 (the centre arm) andA-mansfield-garden-partyM01–M11 (the period arm). Reuse is what makes this a replication rather than a new design. - The S074 pair,
workshop/translations/karen/, for the length control of B-Q2.
B3. Procedure
B3.1 The material gate — done, and reported here as executed
CLAUDE.md's standing rule: measure the lead's contamination on candidate material before
designing an experiment on it. A recall probe was written before the translation limb began
(material/recall-probe.txt), declaring no recall of any English rendering of this story and
offering a good-faith reconstruction as the probe's positive-effort arm. Against the stored
comparator it scores max run 8 tokens ("in a row upon the floor the silver"), 5 shared
7-grams, 0 shared 12-grams — clean on the frozen rule. Gate PASSED; the unit is
selected.
Two confounds on the probe are declared rather than argued away. (i) The lead had read
roughly the first 2,200 characters of the Italian source while choosing material, so the probe
is not a pure memory probe: it is partly a translation from short-term memory of the Italian, and
its 8-token run is therefore best read as a forced-run floor, not as evidence of recall.
(ii) RS-20260730f-recall-floor established that free recall does not work as a control on
canonical material. Neither confound bears on the gate's use here, which is a ceiling check —
does the lead reproduce long runs? — and the answer is no.
B3.2 The two renderings
R10 as frozen at v1.0, both arms, in this order and with no return pass:
- The opportunity list — which of C01–C19 and which of M01–M11 the source unit supplies any material for at all — is written and committed before the source is read for translation.
- Centre arm (
A-mchugh-presence, C01–C19) translated straight through, log written as the draft is written, every logged decision carrying a target code A / X / N and the proposition ids. Committed. - Period arm (
A-mansfield-garden-party, M01–M11) translated straight through from the same source unit, same procedure. Committed. - Only then is anything measured.
The order is fixed and it is a confound, exactly as it was at S074: the second arm is written by a translator who has just written the first. It is declared, not controlled, and B-P4 below is the check on it.
B3.3 The measurement
tools/dependence_check.py, unmodified, against the whole 3,426-word published English —
identical in method to workshop/translations/karen/gate/, which is what makes the two comparable.
Four cells: centre~published, period~published, centre~period, recall-probe~published.
And the length control that S074 did not run, applied to both pairs:
- L1 — rate. Shared 7-grams per 1,000 lead tokens, both arms, both sources.
- L2 — truncation. Truncate the longer arm to the shorter arm's token count and recompute the longest run and the 7-gram count. Done at the token level, both directions reported.
- L3 — re-run on S074's stored pair. The same two controls applied to
workshop/translations/karen/. If S074's crossover does not survive L2, the published figure is corrected in place on the artifact pages and onNEXT.md's successor, which is whatARM-figure-auditis for.
B4. Predictions, frozen
- B-P1 (the replication). The period arm's longest shared run with the 1907 comparator exceeds the centre arm's. Fails if it is equal or shorter.
- B-P2. The period arm's shared 7-gram count with the comparator exceeds the centre arm's.
- B-P3 (the floor). Both arms exceed the recall probe's 8-token / 5-heptagram floor. If an arm does not, the crossover is smaller than the floor and B-P1 is uninterpretable whichever way it goes.
- B-P4 (translator identity dominates). The two arms share more with each other than either shares with the published translation, on both statistics. Fails if either arm's overlap with the published exceeds the two arms' overlap with each other — which would mean register targeting is a larger effect than being the same translator, and would be a much stronger claim than S074's.
- B-P5 (length). S074's crossover survives the L2 truncation control. Fails if truncating the period arm to the centre arm's token count removes the difference — in which case the published figure was a length artifact and is corrected.
- B-P6. Neither new arm reaches 12 tokens against the comparator, the project's frozen
cleanthreshold. Fails if either does, and that arm is declared contaminated on its artifact.
B5. Failure criteria for Part B
- If the two new arms' token counts differ by more than 20%, the raw comparison is reported as length-confounded and only the L1/L2 controlled figures are reported as the result.
- If the opportunity list finds fewer than four propositions in either target set with any material in this source unit, the source unit cannot separate the targets and Part B reports that instead of a crossover.
- No quality judgement is made or elicited about either rendering, by anyone. Tier D has not passed. Both are labeled subjects; the lead never judges its own translation (charter §5).
5. Budget
Part A: $0.00. Part B translation, gate and measurement: $0.00 — lead translation is free and is never ledgered (charter §3, A4).
The only API call in this session is the pre-run critic.
| call | max_tokens | worst case |
|---|---|---|
pre-run critic — qwen/qwen3.7-max, probed-but-not-selected |
12,000 | $0.18 |
declared reserve if note (b) fires — openai/gpt-5.6-terra |
12,000 | $0.09 |
| total worst case | $0.27 |
Note (b) has fired twenty times and NEXT.md's reading of it is that the remedy is the fall-through, not the cap. A reserve seat is therefore declared before dispatch (note (bfc)), at a different lab, and the cap is set at the level that has worked (12,000) rather than raised again.
UTC day 2026-07-31 opens for this session at $1.859570961 of $5.00, headroom $3.140429039. The worst case is 8.6% of headroom.
6. Verification
analysis/verify.py, importing nothing from analysis/sweep.py or analysis/analyse.py,
recomputes: every body classification from the stored .raw bytes; every verifier classification
from the stored source files; every mutation outcome; every contamination cell by an independent
n-gram implementation that does not call tools/dependence_check.py; both length controls; and the
S074 figures it re-audits. At least five mutation tests that must be caught.
The critic's raw body, its cost and its provider are stored and re-summed by verify.py.
Amendment A1 — 2026-07-31, after the pre-run critic. All seven findings accepted.
Critic: qwen/qwen3.7-max, one call, in 4,829 / out 8,738 (of which 31,547 characters of
reasoning), stop, 158.9 s, provider Alibaba, $0.040804105. Verdict NEEDS-AMENDMENT,
seven findings: one BLOCKING, four MANDATORY, two ADVISORY. All seven accepted; none declined.
Raw at runs/critic.raw, prompt at runs/critic.prompt.txt. The seat is probed-but-not-selected
and is a subject in nothing here — S053 role-collision fix, eighteenth session running. This
session dispatches no other API call, so there is no collision to make.
Two of the seven are answered with a stronger remedy than the one proposed, and both of those are marked. Nothing is answered with a weaker one.
A1.1 — BLOCKING, B3.1. "The lead grades its own homework on the material gate."
Accepted; the remedy is strengthened. The finding is exactly right and its direction is the dangerous one: a probe the lead writes itself can pass the gate by avoiding the published wording, and avoidance is invisible in the statistic. The critic proposes an external baseline. A better one is free and is already in the materials.
AMENDED. A deterministic non-lead floor is added and becomes the reference against which
every figure in Part B is read. Short Story Classics (Foreign) vol. 2 contains sixteen stories
Englished for one publisher in one year. Measuring the "The End of Candia" comparator against the
other fifteen gives the distribution of longest-run and shared-7-gram between two unrelated
period English translations, written by different hands, of different stories, from different source
languages, at matched length. No lead text is involved and nothing about it can be gamed. This
is the null that Part B's whole question needs and that neither S074 nor this design had: how many
tokens does 1907 English share with 1907 English by construction?
The recall probe is demoted: it is reported as declared, its 8-token run is reported, and it is no longer offered as the gate's evidence. The gate's evidence is the anthology null.
A1.2 — MANDATORY, B3.2. The ordering confound. "Wipe context between arms; randomise order."
Accepted; the remedy is strengthened, and one half of it is refused as unavailable rather than declined.
- A context wipe is not available to this architecture and saying so is the honest answer. The lead is one continuous agent within a session; there is no isolated draw. This is a standing limit on every matched pair this project has produced, it was not created by this design, and it is now written down rather than assumed away.
- Randomisation is replaced by CROSSING, which is stronger.
E-20260731fwrote the centre arm first (T-karen-R10c-v1: "The centre arm was written first"). This design therefore writes the period arm first and the centre arm second — declared here, before either exists. If the crossover reproduces with the order reversed, order did not produce it; if it inverts, order is a candidate and the two sessions together say so. One randomised draw could not have established either. - B-P4 is reworded as the critic asks: it measures translator identity, not the ordering confound, and it is not a control on the latter.
A1.3 — MANDATORY, B3.3. "L1 and L2 do not control what B-Q2 says they control."
Accepted in full. The longest common run is an extreme-value statistic whose expectation rises with length even at a constant rate, so a rate control does not control it.
AMENDED. - L1 applies to shared 7-gram counts only and is stated as such. - L2 truncates from the TAIL — the longer arm is cut at the token index equal to the shorter arm's token count — and this is the only length control reported for the longest run. - The anthology null of A1.1 supplies the length-matched reference distribution for the longest run that L1 could not.
A1.4 — MANDATORY, A3.5 / A-P3. "The mutation cannot distinguish caught from crashed."
Accepted in full, and this one changes the instrument rather than the prose. analysis/mutate.py
now classifies every mutation outcome into three, not two:
- CAUGHT — the verifier ran to completion and reported at least one check failure.
- CRASHED — the verifier raised. This is not sensitivity: it is a broken script meeting a malformed file, and counting it as a catch is the false positive the critic names.
- SILENT — the verifier exited 0 reporting zero failures against a body it should have rejected.
Only CAUGHT counts toward A-P3. CRASHED is reported separately and is a finding of its own kind: a verifier that crashes on a truncated body would, on a plausibly truncated body, not crash.
A1.5 — MANDATORY, B3.3 / B5. "The tokenizer is never named."
Accepted in full. The tokenizer is tools/ngram_overlap.tokenise() plus the standalone-digit
deletion, exactly as tools/dependence_check.py applies it and unmodified — named here so that
every token count, rate denominator and truncation index in Part B is reproducible. It is
deliberately not a model tokenizer: S074's figures were computed with this one, and a
replication that changed the tokenizer would not be measuring the same quantity.
A1.6 — ADVISORY, B-P3. "A prediction whose failure is unreportable."
Accepted. AMENDED: if B-P3 fails, Part B reports, as a standalone finding, that the floor exceeds the effect — that the project's contamination instrument cannot resolve a register difference of the size S074 reported, which is a reportable result about the instrument and not a missing result about the translations.
A1.7 — ADVISORY, front matter. "The design asserts senses it does not measure."
Accepted. The senses: list is removed. No quality judgement is made or elicited anywhere
in this design, so it is not an evaluative page and the front-matter rule does not require the field.
The critic's proposed replacements (contamination, provenance) are not used, because
senses: takes ids from wiki/goodness-senses.md and neither is one; inventing an id to satisfy a
lint would be worse than omitting the field.
A1.8 — what the amendment does NOT change
The predictions A-P1, A-P2, B-P1, B-P2, B-P5 and B-P6 are untouched, and no failure criterion is
loosened. A-P1 was already at risk when this amendment was written — the hand check of the four
ACCEPTED-DEFECTIVE cells was under way — and it is left standing exactly as frozen.