Repository path: workshop/experiments/E-20260726d-tierD-heldout/design.md · rendered 2026-09-09
Page metadata (front matter)
Frozen design (v2) — Tier D on accuracy, with the held-out arm built
v1 was rejected by the independent critic pass (critic.md, verdict NEEDS-REDESIGN, 22 findings). This is the amended version: 17 implemented, 5 accepted as stated limitations, 1 rebutted in part. Dispositions are tabulated in critic.md and dated in §13. A run may proceed only against this version.
No senses: field: this page designs an evaluation of the jury, not of a translation, and asserts no evaluative claim about any translation.
ARM-tierD step 2. The arm this design builds is the one E-20260725-tierD-ladder §0 declared unsatisfiable from stored materials and dropped, pre-committing to claim no Tier D pass. The materials condition has since been discharged (D-20260725-07, S025) and the qualifying pair verified independent (RS-20260726b-baseline-dependence, S030). Nothing about the control has been relaxed. What changed is the materials.
0. What this run can and cannot be — read first
This run can report Tier D on accuracy and on nothing else. Condition (i) — comparative reception evidence — is ratified for the Garnett/Hapgood pair for the Memoirs of a Sportsman cycle only and for accuracy and cultural-mediation only (D-20260725-07). naturalness, literary-quality, style-correspondence and voice are inadmissible on this pair, because two contemporaneous reviews agree on a directional Garnett advantage on English prose, and a held-out control on a pair with a documented gap on the measured sense imports exactly the bias the control exists to exclude. All six senses are still scored — cross-sense specificity is the headline Tier D metric and needs the untargeted senses — but only accuracy can be reported as a Tier D outcome. cultural-mediation is admissible on the pair and is not targeted by this design's operator; it is reported as a specificity comparator, not as a second gate result.
Three things this run cannot do, stated before it runs:
- It cannot establish that the held-out pair is of equal quality, and it does not try.
D-20260725-06ratified that parity must come from independent external grounds and may not be created by setting a tolerance after seeing a damage effect. Parity here is a premise, discharged externally by conditions (i) and (ii). The arm tests whether the jury behaves as that premise predicts. - It cannot separate quality from canonicity. Garnett is the default English Turgenev and Hapgood is not (
RS-20260725-heldout-pair, S021). The pre-registered reading of a separating held-out arm is written in §7 and claims no more than "the arm did not behave at chance; canonicity is not excluded." - A non-firing held-out arm is weak evidence, and v1 drew the wrong conclusion from that. The rule has power 0.56 against a jury that prefers one text in 80% of units and 0.10 against 50% (§6.4, computed exactly). v1 argued that this "makes a pass easier, not harder" and treated the concession as sufficient. The critic's correction (A2) is adopted: that is mechanically true only because non-firing is a pass condition, and it is not evidentially sound. A non-firing held-out arm licenses exactly one sentence — "this 5-of-6 rule did not detect separation" — and licenses none of: chance behaviour, parity, equivalence, absence of a quality difference, or absence of the canonicity, OCR, orthography and length confounds of §12. The control can kill a detection claim; it cannot discharge itself. Every Tier D statement this run can make is therefore conditional, and §7 writes the condition into the claim rather than into a caveat below it.
1. Question
Does the jury detect deliberate accuracy damage, at two doses, while failing to separate two independent published translations of the same passages — with the sham, held-out and cross-sense controls all measured on the same materials, by the same jurors, under pre-specified directional-consistency rules of the same kind?
2. Why the held-out arm must sit on the same materials as the detection arm
S020's detection arms are on Japanese sources (Akutagawa/Shaw, Genji/lead). Attaching a Russian held-out arm to them would give a null on materials where detection has never been demonstrated, and a null of that kind is uninterpretable: a jury sitting at chance on Garnett-vs-Hapgood might be reading two equally good translations, or might be unable to score Russian→English accuracy at all.
A control needs a case it should pass as well as a case it should reject (standing critic disposition 7). So every arm in this design runs on Russian→English prose, judged by the same three jurors, under analogous pre-specified directional-consistency rules, calibrated separately for 9 units and for 6 (§6.3 — v1 called these "the same rule at the same threshold" and the critic showed that is false in four respects at once). The held-out arm is read against the targeted arm's behaviour on the same materials, not against an abstract chance level.
3. Materials
Everything is public domain or lead-authored and already in the repository. Nothing new is fetched at run time.
3.1 The held-out passages
Turgenev, «Свидание» (Записки охотника), opening, 468 Russian words — workshop/translations/svidanie/R04-v1/source-ru.txt. Two published English translations, stored whole from the S024 audit:
- Garnett 1897, 607 words —
E-20260725-svidanie-audit/runs/T1-garnett.txt - Hapgood 1904, 666 words —
E-20260725-svidanie-audit/runs/T2-hapgood-1903.txt
Split into two items at a landmark named here so the build cannot choose it. The boundary falls immediately before the Russian sentence beginning «Листва на березах была еще почти вся зелена» — in Garnett, "The leaves on the birches were still almost all green"; in Hapgood, "The foliage on the trees was still almost entirely green". Measured word counts either side:
| item | Russian | Garnett | Hapgood |
|---|---|---|---|
| HA (sentences 1–8) | 225 | 299 | 324 |
| HB (sentences 9–14) | 243 | 308 | 342 |
Both halves fall inside the 100–500-word passage unit that D-20260725-06's ratification named as the route forward, and inside a 250–400-word band that keeps the §5 dose comparable across items.
Why this pair qualifies, with each condition cited rather than asserted:
- Condition (i), comparative reception evidence — ratified
D-20260725-07. Two independent contemporaneous venues (The Nation, 4 Feb 1904; The Athenaeum, 20 Jan 1906) agree per dimension on this cycle: "almost exactly the same number of errors", Hapgood "decidedly the more accurate", Garnett better on English prose. Scoped toaccuracyandcultural-mediation. - Condition (ii), pre-run factual-damage audit against the source —
E-20260725-svidanie-audit(S024), 17 typed landmarks: Hapgood 17/17, Garnett 15/17 raw, both flags adjudicated to non-damage (one an unbounded regex in the project's own tool, one a hypernym). - Independence, measured — Garnett 1897 and Hapgood 1904 share zero ≥12-token runs on this text (
RS-20260726b-baseline-dependence, S030). A held-out control requires the two texts not to be copies of each other, and this is the first pair in the project for which that is measured rather than assumed.
3.2 The lead-provenance passage
Charter §5.1: references of both provenances, "so the result is not an artefact of one source."
T-son-makara-R04-v1 — Korolenko, «Сон Макара» §I opening, 270 Russian words → 394 English words, translated by the lead this session and frozen before any measurement of it existed.
Why not a Turgenev or Chekhov lead translation. The arm page rules them out and the measurement is why: the lead reproduces 11–21 consecutive words of Garnett from the Russian alone at 8 of 8 loci (RS-20260726c-forced-or-borrowed-ru), and T-svidanie-R04-v1 shares a 19-token run with Hapgood. A "lead-provenance" reference that is a remembered published one is not a second provenance.
The selection gate, run before this design selected the material and after the translation was frozen — the standing rule from ARM-overlap-dependence (2026-07-26): measure contamination before choosing, never as a diagnostic inside the experiment. contamination/result.json, tools/dependence_check.py, lead against the whole 10,738-word Fell 1916 story rather than the aligned passage, which is the more conservative direction because it gives every possible run a chance to match:
| pair | shared 7-grams | shared 12-grams | shared 15-grams | longest run | verdict |
|---|---|---|---|---|---|
| Fell 1916 ~ lead | 3 | 0 | 0 | 8 tokens | clean |
The longest run is "to die he was very proud of his" — eight tokens of which one is a content word. Against the project's measured range (Ovid 0, Turgenev 21) this sits at the clean end. Fell's English was downloaded before the translation and deliberately not read; only its length and hash were printed. The blind is what makes this a measurement.
Stated limitation: Fell 1916 is the only freely reachable English Makar's Dream, so the gate is a one-comparator test. A clean result against one comparator is weaker than a clean result against three, and the alternative was no measurement at all.
3.3 What the jury is shown, and what is stripped
Each item is the Russian source + two unattributed English texts. Removed by materials/build.py: translator names, dates, licence blocks, title lines, footnote markers, section numbering. Prose unwrapped to one paragraph per line so hard-wrap differences cannot fingerprint a text.
The typographic-neutrality check is mandatory, measured, and reported — it is not an assumption (standing disposition 6). Hapgood's stored text is an OCR'd scan and Garnett's is not, and a scan artefact is a one-glance discriminator that would let the jury separate the held-out pair without reading it. Already found by inspection and confirmed by analysis/split_check.py: Hapgood contains trav- ersed (a line-break hyphenation artefact) and one em-dash where Garnett's span has none.
v1 stopped at hyphen-space artefacts. The critic's B1 is adopted: repairing one artefact does not remove an OCR signature. The build must therefore:
- repair
trav- ersed→traversed, logging the repair verbatim; - normalise both texts through one pipeline for curly/straight quotes and apostrophes, dash forms, ligatures and whitespace — both texts, so the normalisation cannot itself become the discriminator — and log every character class it changed, with counts, per text;
- print a full character-inventory diff between the two texts of each item: every character present in one and absent from the other;
- print, per text, a style profile — sentence count, mean and SD of sentence length, punctuation counts by mark, paragraph count (critic B3: these can trigger a stylistic preference that has nothing to do with accuracy);
- print, per text, an orthographic profile —
-ise/-ize,-our/-or,-re/-er, single/double-ll-, and the count of tokens absent from the other text's spelling variant set (critic B4: 1897 British against 1904 American is a live discriminator, and thenaturalnessregister frame does not neutralise it for overall preference or for the other five senses); - fail loudly if any built text still contains a hyphen-space artefact.
Punctuation differences that are the translators' own — including the em-dash — are left alone and reported, because normalising them would edit the texts the control is about. All of 3–5 are reported with the result, so that a separation in the held-out arm can be checked against them instead of being attributed to quality by default.
4. The operator
One operator: O4, accuracy, reused verbatim from E-20260725-tierD-ladder §3 — semantic errors: wrong referent, wrong word sense, invented detail, dropped negation. Style untouched. Reuse is deliberate: the arm page asks for S020's operator set, and re-authoring an operator would make this a new instrument rather than the missing control on the old one.
Its catalogue basis, restated because charter §5.2 forbids lead-invented operators. O4 does not come from Berman — the frozen mapping at T-berman-tendances-R04-v1 records that none of the twelve tendances déformantes targets accuracy, Berman's quarrel being with the translator who writes well. O4's catalogue is the project's own documented record of accuracy failure in published translations: the Shaw Christianisation (A-shaw-spider-thread), Swann (1974) on Turney's lapses, and A-chekhov-pari on Koteliansky & Murry 1915 — «за пять часов» → "five minutes", «в 12 часов дня» (noon) → "twelve o'clock midnight", «Евангелие» → "the New Testament". Every edit made under O4 must be an instance of a failure type on that list, and materials/perturbations.md names the type per site.
Composition constraint, fixed before any edit exists. Each 8-site set contains at least two of each of the four failure types, so no single type carries the dose.
5. Dose — the axis S020 deferred
Charter §5.2 asks for graded doses; S020 ran one dose and deferred the axis explicitly. This design runs two.
- Heavy — 8 sites, S020's dose, so the targeted cell is comparable to the run that cleared both legs.
- Light — 3 sites, a nested subset of the heavy set.
The nesting rule is pre-registered here, before any edit is written, and this is the whole point. S020's critic (D24) killed v1's light dose because its light sets "were nested subsets that retained the loudest edit, so they could not have measured a dose effect" — the defect was selection after seeing the edits. This design fixes the selection rule first: the 8 sites are numbered in textual order, and the light dose is sites {2, 5, 7}. The rule is stated in this frozen page, so it cannot be tuned to include or exclude any particular edit. Which failure types land in {2, 5, 7} is reported, not controlled — the composition constraint of §4 makes a monoculture unlikely but does not forbid it.
What pre-registration does and does not buy (critic D1). It prevents post-edit cherry-picking of conspicuous edits — the defect S020's critic actually named — and it does not make {2, 5, 7} a representative 3-site dose: positional selection can still correlate with narrative salience, sentence position or edit type. The light dose is one deterministic subset, not a sample. An independently drawn light set would double the perturbed materials and does not fit the budget; nesting is also the tighter control for dose per se, since the sites are held constant.
Dose is a count of sites, not a rate. Reference lengths run 292–394 words, so 8 sites is 2.0–2.7 sites per 100 words. The rate is reported per item. Any per-item difference in detection is confounded with this spread and no claim of the form "detection is stronger on item X" is licensed.
6. Arms, jury, and the rule
6.1 Ten items, 20 payloads, 60 calls
| # | item | arm | reference | second text | units |
|---|---|---|---|---|---|
| 1 | H-HA |
held-out | Garnett HA | Hapgood HA (unmodified) | — |
| 2 | H-HB |
held-out | Garnett HB | Hapgood HB (unmodified) | — |
| 3 | S-GA |
sham | Garnett HA | Garnett HA + 8 neutral substitutions | — |
| 4 | S-K |
sham | lead K | lead K + 8 neutral substitutions | — |
| 5 | T8-GA |
targeted, heavy | Garnett HA | + O4 × 8 | — |
| 6 | T8-HB |
targeted, heavy | Hapgood HB | + O4 × 8 | — |
| 7 | T8-K |
targeted, heavy | lead K | + O4 × 8 | — |
| 8 | T3-GA |
targeted, light | Garnett HA | + O4 × {2,5,7} | — |
| 9 | T3-HB |
targeted, light | Hapgood HB | + O4 × {2,5,7} | — |
| 10 | T3-K |
targeted, light | lead K | + O4 × {2,5,7} | — |
The damaged reference alternates between the two published translators by design — Garnett carries HA, Hapgood carries HB — so no result can be read as "the damage was only ever applied to Garnett". Translator and passage are thereby confounded with each other; that is accepted and stated, and it buys the more important protection.
Order swap is mandatory on every item: each pair is dispatched twice with slots swapped. Slot preference in this project has been measured at 0.50–0.85; on this exact item format S020 measured P1 0.500, P2 0.600, P5 0.550. An unswapped run is uninterpretable (standing disposition 4).
6.2 Jury
P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P5 deepseek/deepseek-v4-pro, resolved from config/models.md at run time and logged as provenance.
P3 and P4 are excluded, and this is a power limitation, not a finding about those models — the same exclusion S020 made, kept deliberately so that this run is a controlled extension of S020 rather than a new instrument: these three jurors' scale usage and slot preference on this exact item format are measured, and a fourth would add units at the cost of that comparability. P4 alone would roughly double the run's cost (measured $0.09089/call in S014). Consequence, stated: the US-taste-correlation revisit trigger in config/models.md cannot fire on this run, and the panel is 2 US labs and 1 Chinese lab rather than panel v1's five.
Six senses, scored 1–7 for each text — accuracy, naturalness, voice, style-correspondence, literary-quality, cultural-mediation — plus a forced overall preference, no ties. Strict JSON. naturalness is put to the jury with its register frame named, verbatim from S020 §5, because S014 measured this sense as register-cued and without the frame it measures period preference.
Judgment is not parallelized (charter §6): strictly sequential dispatch.
6.3 The unit and the firing rule
The unit of analysis is (juror × item), with the two orderings averaged within the unit — not the individual vote. Pooling votes as if independent is pseudo-replication (standing disposition 1); it has been raised, accepted, and then reintroduced once already in this project.
A unit scores +1 if the reference is preferred in both orderings, −1 if the second text is preferred in both, 0 if split. Under a null of independent coin-flip preferences, P(+1) = P(−1) = 0.25 and P(0) = 0.5.
| cell | items | units | rule | direction | exact P under the null |
|---|---|---|---|---|---|
| targeted, heavy | 3 | 9 | ≥7 of 9 are +1 and none is −1 | one-directional | 0.000622 |
| targeted, light | 3 | 9 | same | one-directional | 0.000622 |
| held-out | 2 | 6 | ≥5 of 6 consistent for either text and none opposite | either-directional | 0.006348 |
| sham | 2 | 6 | three-way band, §6.5 | — | — |
Family-wise rate across the two targeted cells: 0.0012. All figures are computed exactly from the trinomial null by analysis/rules.py, not looked up, and the script is committed with this design.
These are analogous rules calibrated separately, and v1's claim that they are "the same rule at the same threshold" was false in four respects at once (critic A1): 7/9 = 0.778 against 5/6 = 0.833, one-directional against either-directional, and null probabilities an order of magnitude apart — 0.000622 against 0.006348 — hence different power. Nothing in this design rests on the symmetry; what it rests on is that both rules are pre-specified directional-consistency rules applied to the same jurors on the same materials, so the question remains "does a rule of this kind fire where nothing was damaged?" and not "is the held-out pair equal?". The held-out rule is the more permissive of the two, which is stated here rather than buried, because it is the direction that flatters a pass.
The units are not independent, and the exact P values are therefore anti-conservative (critic A4, accepted in part). A cell's 9 units are 3 jurors × 3 items; a juror with a fixed taste contributes correlated units, so the true false-positive rate is at least the tabulated figure and may exceed it. The tabulated P is a lower bound and is reported as one.
What is rebutted: the critic also proposed that "a juror can recognize the passage, infer the repeated reference, learn its properties from an earlier payload". It cannot. Each payload is an independent stateless API call with no conversation state and no cross-call memory, and dispatch order is fixed by §10 rather than by anything a model chooses. The dependence here is fixed juror taste, not within-run learning, and the distinction matters because taste is addressable and learning would not be.
The per-juror robustness condition, added in v2 because taste is addressable. A targeted cell may not be reported as firing unless, in addition to the rule above, it fires within at least 2 of the 3 jurors taken separately — i.e. at least 2 jurors are +1 on all 3 of their items with none −1. This is insensitive to juror-level correlation, since it asks each juror's units to agree with each other rather than pooling them. Under the independence null a single juror is +1 on all 3 items with P = 0.25³ = 0.0156, so the condition is easily reachable by a real effect (S020's accuracy cell was 6/6 units) and is not reachable by one loud juror alone. **Per-juror breakdowns are reported for every cell whether or not the condition fires."
6.4 The held-out arm's power, computed rather than hoped
Against a jury that genuinely prefers one text, with P(−1) held at 0.05, the 5-of-6 rule fires with probability:
| true P(+1) | 0.5 | 0.6 | 0.7 | 0.8 | 0.9 |
|---|---|---|---|---|---|
| power | 0.10 | 0.21 | 0.37 | 0.56 | 0.71 |
So a non-firing held-out arm is compatible with a substantial real preference, and §0.3's asymmetry follows arithmetically. What would fix it is more jurors and more items; the audited span is 468 Russian words and the budget is §8. Reported alongside the result, never omitted.
6.5 The sham's three-way rule, verbatim from S020 §6
On its own 6 units: ≥5 of 6 units +1 → the jury prefers unedited text as such; every detection result is confounded with edit-presence and no detection claim is licensed from this run. ≤1 of 6 units +1 → the jury prefers edited text; same conclusion, opposite sign. Otherwise → the false-alarm rate is acceptable at this resolution and detection stands on its own terms.
The sham is an upper bound on the false-alarm rate, not a neutral edit. The reference is a considered text and any eight substitutions move it off a local optimum; a perfectly neutral sham is not constructible. Constraints, carried from S020 v2 and all measured in materials/perturbations.md, not asserted: free variation only; word-, sentence- and paragraph-count neutral; no archaising; no sense shift. S020's measured sham was 2/6 units — small but non-zero.
6.6 Cross-sense specificity — the headline Tier D metric
Charter §5.4. Specificity fires for a targeted cell iff both:
- the mean drop on
accuracyexceeds the mean drop across the non-target senses excludingnaturalnessby ≥ 0.75 scale points; and - the largest per-sense drop, excluding
naturalness, isaccuracy; and drop(naturalness)≤ 0.75 scale points.
Condition 3 is a condition of the rule, not a veto applied after it — v1 wrote it as a separate consequence and thereby contradicted its own "iff" (critic A3): accuracy could clear conditions 1 and 2 while naturalness moved, so specificity formally fired and was narratively voided, with no reading assigned. There is now one rule and no gap.
naturalness is excluded from the baseline of condition 1 for the reason S020's critic gave (B9) — the design predicts it may rise for style-targeting operators, which would mechanically inflate the margin it is meant to be independent of. Both the excluding and the including figures are computed and reported; only the excluding one enters condition 1.
What a naturalness movement licenses, corrected. O4 leaves style untouched, so naturalness should be roughly flat, and condition 3's threshold is looser than S020's 0.25 for O2/O3 because accuracy damage can incidentally read as odd English, which O2/O3 damage cannot. If condition 3 fails, the licensed statement is "specificity is not established for this cell: the target sense's margin cannot be separated from a general downward movement that also reached naturalness" — not v1's "the jury lowers every sense together", which does not follow from a naturalness drop alone, since the other non-target senses may be flat. Which senses moved is reported per cell.
7. Pre-registered readings, written before the run
Every outcome is assigned its reading here. Nothing is decided after seeing numbers.
Gate status, per outcome:
| targeted (heavy) | sham | held-out | recorded in config/models.md |
|---|---|---|---|
| detection + specificity fire, and the per-juror condition of §6.3 holds | in band | does not fire | Tier D reported on accuracy at 8 sites, phrased "the jury detects accuracy-damage at 8 edit sites in 292–394-word Russian→English passages, on a held-out control that did not separate an independent same-quality pair — a control with power 0.56 against an 80% preference, which therefore does not exclude a real separation". The power clause is part of the claim, not a footnote to it (critic A2), and nothing broader may be recorded (charter §5.5) |
| detection + specificity fire, per-juror condition fails | in band | does not fire | NOT PASSED. "the cell fired on pooled units but not within jurors taken separately; the pooled result is not separable from one juror's taste." |
| detection + specificity fire | in band | fires | NOT PASSED. Reading: "the rule that licenses detection also separates two independent same-quality translations; the arm did not behave at chance. Canonicity is not excluded as the cause and this design cannot exclude it." |
| detection fires, specificity fails | in band | does not fire | NOT PASSED. "detects damage, not sense-calibrated on accuracy." |
| detection fails | any | any | NOT PASSED. "no detection at this dose and this power." |
| any | out of band | any | NOT PASSED. No detection claim licensed from this run; the sham result is itself reported. |
The light dose is reported but does not gate, and "fire" below means detection; specificity and the naturalness condition are reported for the light cell separately and never enter the dose statement.
| heavy detection | light detection | light specificity | licensed dose statement |
|---|---|---|---|
| fires | does not fire | — | "this 8-edit composite met the rule; this nested 3-edit composite did not." Not a dose-response, not "eight sites are stronger than three", not "the extra five sites caused the difference", not "3-site damage generally escapes detection" (critic D2) |
| fires | fires | fires | the licensed sense statement names 3 sites, the smaller dose being the stronger claim |
| fires | fires | fails or condition 3 fails | detection is reported at 3 sites; specificity is reported at 8 sites only, and the sense statement names 8 |
| does not fire | any | — | §7's last row governs; the light cell is reported and licenses nothing |
If the held-out arm fires, the arm still closes resolved, not retired. That is a licensed finding about the jury and is the arm's stated completion criterion.
8. Predictions, written before the run
- Detection fires at 8 sites. S020 got 6/6 units on
accuracyon Japanese sources; this is the same operator on an easier source language for the panel (S015 screened Russian competence at ≥5/6 for every panel member used). - Specificity fires at 8 sites. S020's margin was +3.10 against a 0.75 criterion, with accuracy falling 6.83 → 1.75.
- Detection does not fire at 3 sites.
internal-judgment-only, and the least confident of these. If it does fire, the licensed dose statement strengthens. - The held-out arm does not fire. Predicted, but see §6.4 — the prediction is weakly tested by construction.
- The sham lands in the middle band, as S020's did at 2/6.
drop(naturalness)stays under 0.75 in both targeted cells.- The lead-provenance item behaves like the published ones. v1 wrote this as "detection is at least as strong … it is reported", which every possible outcome satisfies — a prediction with no failure condition, the fourth occurrence of that defect in this project (critic C1, standing disposition 3). Stated so it can fail: on the heavy cell,
T8-K's three units contain at least 2 of 3 at +1, and its meanaccuracydrop is within 1.50 scale points of the mean ofT8-GAandT8-HB. Either failing is the falsification, and the reading is that the both-provenances requirement has found something: detection on this jury depends on whether the reference is published or lead-authored. Both figures are reported whichever way they land.
9. Failure criteria — what voids or qualifies the run
- >10% of calls failing or unparseable after one retry → reported as a failed run, not patched.
- Sham outside the middle band → §6.5's rule applies; detection claims are withdrawn.
- Post-run verification failing to reproduce a reported number → the number is withdrawn, not corrected in place.
- Any built text still containing a hyphen-space artefact → the build fails and the run does not dispatch (§3.3).
- Reconstruction check failing →
verifyrebuilds each variant from its reference plus the logged edits and asserts byte equality. This catches edits that are present but unlogged, which is the direction S020's v1 check missed and its critic found. - No order-flip threshold. Replaced by a pre-committed P5-excluded sensitivity analysis reported alongside the primary, per the disposition accepted in
E-20260725-anchor-verification(C2). P5's order-flip rate was measured at 0.44 in S014.
10. Budget
Built from max_tokens, not from an assumed output length — note (abc), the defect that produced this project's only estimate overrun. Per-model max_tokens is sized from S020's measured per-call maxima on this exact task, with headroom:
| juror | S020 measured max output | max_tokens here |
headroom |
|---|---|---|---|
P1 gpt-5.6-terra |
1,052 | 3,000 | 2.9× |
P2 gemini-3.6-flash |
2,517 | 6,000 | 2.4× |
P5 deepseek-v4-pro |
9,798 | 14,000 | 1.4× |
P5 came within 2% of S020's 10,000 cap and every one of the 60 calls returned finish_reason: stop; raising its cap is the cheap insurance and its output price is the lowest of the three.
Worst-case input is taken at 3,500 tokens (S020 measured 2,074 max on Japanese sources; Cyrillic tokenises less efficiently and this design ships the source plus two texts).
| juror | price in/out per M | worst case per call |
|---|---|---|
| P1 | $2.50 / $15.00 | $0.0538 |
| P2 | $1.50 / $7.50 | $0.0503 |
| P5 | $1.65 / $3.30 — the worst plausible provider, not the list price | $0.0520 |
A payload is one item in one ordering; a call is one dispatch of a payload to one juror. 10 items × 2 orderings = 20 payloads = 60 calls.
Worst case for the 60 calls: $0.1561 per payload × 20 = $3.12 — and v1 stopped there, which was wrong (critic F1). §9 permits one retry per call and voids the run only above a 10% failure rate, so the retry allowance is budgeted at 10% of each stage's calls, rounded up, priced at the dearest single call ($0.0538): 8 retries = $0.43. True worst case: $3.55. Central estimate, from S020's measured per-call means on the identical run size: $0.755.
The 4.1× spread between central and worst case is what building from max_tokens actually costs, and it is the honest number. P5's price is deliberately not the list price: S020 routed the same slug across eight different providers (Baidu, DeepInfra, Fireworks, GMICloud, SiliconFlow, StreamLake, Together, Wafer) inside one 20-call arm, and S022 measured a Venice route at 3.8× list.
The abort rule is a full-stage reservation, not a running margin (critic F2: "worst call seen so far" is not a predeclared worst possible next call, and an abort on that basis can land mid-stage despite the promise of stage-boundary truncation — leaving a half-dispatched held-out cell, which is the one outcome the ordering exists to prevent).
Before entering any stage, the runner must reserve that stage's full worst case including its permitted retries, computed from the §10 per-call table, against the day's remaining headroom. If the reservation does not fit, the stage is not entered and the deferral is written into NEXT.md. Per-stage worst case, retries included at one per payload:
| stage | payloads | calls | + retry allowance | worst case |
|---|---|---|---|---|
| 1 sham | 4 | 12 | 2 | $0.73 |
| 2 held-out | 4 | 12 | 2 | $0.73 |
| 3 targeted heavy | 6 | 18 | 2 | $1.04 |
| 4 targeted light | 6 | 18 | 2 | $1.04 |
| total | 20 | 60 | 8 | $3.55 |
So the run requires $0.73 of headroom to enter stage 1, $2.51 to be certain of reaching the end of stage 3 (which is the last stage the gate needs), and $3.55 to finish including the dose axis. A day with less than $3.55 runs as far as its headroom reserves and defers the rest, which is exactly what the §10 ordering exists to make survivable. The $5.00 cap accommodates the full run on an otherwise empty day; a day already carrying spend may not.
Dispatch order, so that what truncates first is what can be afforded (note (v)):
- sham — dispatched first and standalone, because it gates everything else and an abort must never be able to drop the control that licenses the rest (S020 critic E28).
- held-out — second, because it is this arm's entire purpose.
- targeted, heavy — third; the positive control that makes the held-out result interpretable.
- targeted, light — last, because the dose axis is the one part of this design that can be lost without losing the gate.
Stage 1 gates stage 2 on the scale-usage condition S020 pre-registered: pooled score standard deviation ≥ 0.75 and at least 4 distinct integers used across the six senses. S020 measured this and passed; if the jurors compress into two or three adjacent values, a 0.75-point specificity margin is not measurable and the run stops with the scale finding alone, which is a legitimate and cheap outcome.
11. Verification
A verifier that recomputes every reported number from runs/ and the items manifest only, re-parsing each model's own text rather than trusting the runner's cached parse. It must:
- import nothing from
tools/(amendment A11 ofE-20260726c-forced-or-borrowed-ru, and the reason isRS-20260726c-name-tokens-repair: a 218-check verification pass found nothing while a shared function was broken in three ways, because both implementations called it); - rebuild each variant from reference + logged edits and assert byte equality (§9);
- recompute the null probabilities and the power table of §§6.3–6.4 by an independent path;
- print the resolved provider and billed cost per call, read off each response (note (x)).
Four checks added in v2, because the critic showed the list above does not catch them (E1–E4). Byte reconstruction from "reference + logged edits" verifies that the logging is consistent; it says nothing about whether the reference, the split or the build is right, and a manifest can faithfully record an already-corrupted text.
- Reconstruct the split independently from the landmark named in §3.1, by exact substring match in all three languages, and assert §3.1's word-count table. Already implemented and passing as
analysis/split_check.py, run before this design was frozen — it is what turned §3.1's counts from assertions into measurements. - Assert that every built reference text is a contiguous substring of its stored source file after the §3.3 normalisation is applied to both. A stripping regex that deletes prose cannot satisfy this, which is the failure E2 names.
- Recompute the contamination gate of §3.2 by the verifier's own tokenisation and matcher, not by calling
tools/dependence_check.py. §11 in v1 did not re-run the gate at all, so a broken matcher there would have passed unseen — precisely the failure mode ofRS-20260726c-name-tokens-repair. - State plainly what the cost check is. Reading
providerandusage.costoff each response is provenance, not invoice-level verification; the key-usage delta is the independent check and is recorded inconfig/budget.mdalongside the per-request sum.
12. Known threats, stated in advance
- Canonicity is not controlled and cannot be (§0.2). The largest single threat to the held-out arm.
- The held-out control is under-powered (§6.4). The second largest, and it points the opposite way: it makes a pass easier, not harder.
- The perturbations and the sham are lead-authored, as in S020. They control for edit-presence, not for lead-shapedness. Both come from the same hand that wrote this design.
- Hapgood's text is an OCR'd scan and Garnett's is not. Mitigated and measured in §3.3; not eliminated. A residual scan signature that no artefact counter catches would let the jury separate the held-out pair without reading it, and would look exactly like a quality judgment.
- The source's presence changes the confound rather than removing it (critic B5, and the sharpest thing in the pass). With the Russian in the payload the jury can reward literal alignment where a freer rendering is equally accurate; without it, it would reward English fluency instead. There is no version of this design without the source, because
accuracywould be unscoreable, so this is not a defect that could have been designed out — it is the price of scoring the sense at all. - The two texts differ in length by 8–11% within each item (Garnett HA 299 / Hapgood HA 324; Garnett HB 308 / Hapgood HB 342). That permits preference for compression, expansiveness or apparent completeness. Not fixable without editing the texts the control is about; measured and reported per item (critic B2).
- Style and orthography are live discriminators independent of quality. Sentence-length and punctuation profiles, and 1897 British against 1904 American spelling and idiom, can drive an overall preference on their own. The
naturalnessregister frame addressesnaturalnessand does not neutralise them for the forced preference or for the other five senses (critic B3, B4). Profiled and counted in §3.3. - A model may recognise memorised wording, edition-specific phrasing or a translator's signature — which is not the same thing as canonicity (critic B7). Canonicity is about which translation is the default; this is about whether a specific string is in training data. v1 folded the second into the first.
- The project holds one witness per translation. Editorial intervention or silent modernisation in either stored edition is not attributable to either translator and cannot be detected without a second witness, which is not freely available for Hapgood (critic B8).
- Translator and passage are confounded in the targeted arm (§6.1), accepted for a stated reason.
- The dose contrast is nested (§5).
- The contamination gate has one comparator (§3.2).
- Three jurors, nine units per targeted cell. This distinguishes near-ceiling from near-chance and nothing finer. No claim of the form "detection is stronger on X than on Y" is licensed unless the gap is the full width of the scale.
- Detection is not judgment. Firing on
accuracyhere would still not license ranking two good translations. That is Tier P, which ran once and failed.
13. Amendment record (v1 → v2, 2026-07-26, S033)
Every change below is a disposition of a finding in critic.md, applied before any call was dispatched and before any perturbation existed. Nothing was amended for a reason internal to this page.
| § | amendment | finding |
|---|---|---|
| 0.3 | the "makes a pass easier, not harder" argument withdrawn; what a non-firing held-out arm licenses stated as one sentence and what it does not licence enumerated | A2 |
| 2, 6.3 | "the same rule at the same threshold" withdrawn as false in four respects; replaced with analogous rules calibrated separately, with both P values printed adjacent and the more permissive one named | A1 |
| 3.3 | typographic check widened from hyphen-space artefacts to a six-part build requirement: repair, one-pipeline normalisation of both texts, character-inventory diff, style profile, orthographic profile, fail-loud | B1, B3, B4, B6 |
| 5 | what pre-registering {2,5,7} buys and does not buy, stated | D1 |
| 6.3 | unit dependence conceded, the exact P declared a lower bound, the "juror learns across payloads" mechanism rebutted on the stateless-call ground, and a per-juror robustness condition added to the reporting rule | A4 |
| 6.6 | the naturalness threshold made condition 3 of the iff rather than a veto applied after it, removing a flat contradiction; the unsupported gloss "the jury lowers every sense together" deleted | A3, C2 |
| 7 | a row added for pooled-fires/per-juror-fails; the power clause moved into the recorded Tier D claim; a light-dose sub-table added and "both fire" defined; the dose statement narrowed to what it licenses | A2, C3, D2 |
| 8 | prediction 7 given a measurable failure condition on named statistics | C1 |
| 10 | worst case rebuilt on 60 calls plus a retry allowance; the running-margin abort replaced with a full-stage reservation rule and a per-stage table | F1, F2 |
| 11 | four verifier requirements added: independent split reconstruction, contiguous-substring assertion against the stored source, independent recomputation of the contamination gate, and a plain statement of what the cost check is | E1–E4 |
| 12 | six threats added or separated: source-presence, length difference, style/orthography, memorised phrasing as distinct from canonicity, single witness | B2, B5, B7, B8 |
Accepted as stated limitations rather than fixed (B2, B5, B8, D1, and the under-power of §6.4). Each is in §12 with its reason. Rebutted in part: A4's within-run-learning mechanism.