Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260801f-tierD-run/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260801f-tierD-run
statusfrozen
created2026-08-01
updated2026-08-01
linkswiki/arms/ARM-tierD-run.md, framework/tierD-repaired-rules.md, framework/control-arm-spec.md, framework/closure.md, config/models.md, config/budget.md, wiki/goodness-senses.md, workshop/experiments/E-20260726d-tierD-heldout/design.md, wiki/findings/results/RS-20260726d-tierD-heldout.md, wiki/findings/results/RS-20260727b-tierD-rules.md, wiki/findings/results/RS-20260728g-scale-usage.md, wiki/findings/results/RS-20260729c-neutral-summary.md, wiki/decisions/resolved/D-20260725-06-heldout-arm-operationalisation.md, wiki/decisions/resolved/D-20260725-07-athenaeum-1906-condition-ii.md, workshop/translations/les-shumit/R04-v1/translation.md, workshop/translations/bargamot/R04-v1/variant-F1.md, PROJECT.md

Frozen design (v2) — Tier D under the repaired rules

v1 was frozen, then put through the independent pre-run critic pass (critic.md, verdict NEEDS-AMENDMENT, seven findings, two BLOCKING). This is the amended version: all seven accepted, five in full and two in part with the refused remedy reasoned. Dispositions are tabulated in critic.md and dated in §14. A run may proceed only against this version, and no stage had been dispatched when the amendments were made.

ARM-tierD-run step 1. This design imports framework/tierD-repaired-rules.md whole and does not re-derive it. That page exists because three sessions of repair work (ARM-tierD-repair, S040–S055) closed with the instruction "import this block; do not re-derive it" and no design imported it for twenty-six sessions.

No senses: field: this page designs an evaluation of the jury, not of a translation, and asserts no evaluative claim about any translation.

0. What this run can and cannot be — read first

Tier D is NOT PASSED and has been since S034. Nothing on this page changes that; only the run can, and only on the terms below.

  1. It can report Tier D on accuracy and on nothing else. Condition (i) is ratified for the Garnett/Hapgood pair for the Memoirs of a Sportsman cycle only and for accuracy and cultural-mediation only (D-20260725-07). The other four senses are inadmissible on this pair: two contemporaneous reviews agree on a directional Garnett advantage on English prose, and a held-out control on a pair with a documented gap on the measured sense imports the bias the control exists to exclude. All six senses are still scored — cross-sense specificity is the headline metric and needs the untargeted senses.
  2. It carries R4 clause 4 forward explicitly. The evidence base licensing clauses 1–2 of the held-out materials contains no printed style exhibit from the Memoirs cycle at all: all eighteen paired extracts the 1904 Nation review prints are cited to A Nobleman's Nest, the work clause 3 refuses to license. This design does not need clause 3 — it takes the licensed senses and nothing else — and it states clause 4 rather than inheriting clause 2 as though its evidence were work-matched.
  3. A non-firing held-out arm licenses exactly one sentence — "this 5-of-6 rule did not detect separation" — and licenses none of: chance behaviour, parity, equivalence, or absence of the canonicity, OCR, orthography and length confounds of §12. Its power against an 80% preference is 0.6554 (§6.4). The control can kill a detection claim; it cannot discharge itself.
  4. The prior positive control is a gate, not evidence. It uses the same operator as the targeted arm, so it is a rehearsal of that arm on different materials, not independent support for it. Its firing licenses nothing whatever about the targeted arm; its failing stops the run. §6.7 states this in the rule itself rather than in a caveat.
  5. A FAIL is completion. ARM-tierD-run §Done when: a FAIL closes the arm resolved with a statement of what a redesign would change or that the approach is exhausted, and may not spawn a repair arm without Tom (wiki/reassessment-2026-08-01.md §6).

1. Question

On which goodness senses, and at what dose, does the panel detect deliberate, sense-targeted degradation — under the repaired rules, with all four controls (prior positive, sham, held-out, cross-sense) measured on the same materials by the same jurors?

2. What is imported, what is carried, and what is new

Imported whole from framework/tierD-repaired-rules.md R1 the sham's decision rule and its power requirement · R2 edit-kind matching · R3 the stage-1 gate's replacement · R4 the held-out materials' licence and its four clauses · R5's five surviving constraints · the required pre-check
Carried byte-identical from E-20260726d-tierD-heldout (S034) the five built reference texts, their sources, the OCR repair, the one-pipeline typographic normalisation, and the three O4 edit tables. materials/build.py asserts byte equality against that frozen manifest (SHA-256 recorded in materials/items.json) rather than re-deriving a build whose verifier already passed 71 checks
New this session a rebuilt sham: 5 items / 15 units instead of 2 items / 6, and matched to the operator on edit kind (R1, R2) · a second lead-provenance reference (T-les-shumit-R04-v1, translated in session, contamination measured clean) and its O4 table · the prior positive control as stage 1 (R3) · a drawn light dose instead of one fixed positional subset (R5)

What the S034 verdict is not. R5: the repaired rules apply to future runs only; nothing here reopens S034, and no figure of S034's is restated as though it had been re-earned.

3. Materials

Everything is public domain or lead-authored and already in the repository. Nothing is fetched at run time.

3.1 The held-out passages (carried)

Turgenev, «Свидание» (Записки охотника), opening, 468 Russian words, split at the landmark named in the S034 design and independently reconstructed there (analysis/split_check.py). Garnett 1897 and Hapgood 1904, stored whole from the S024 audit. Item H-HA 299 / 324 words, H-HB 308 / 341. Conditions (i), (ii), (iii) discharged at D-20260725-07, E-20260725-svidanie-audit and RS-20260726b-baseline-dependence (zero shared ≥12-token runs) respectively.

3.2 The two lead-provenance references

Charter §5.1: references of both provenances. S034 had one lead item and therefore three units on the whole both-provenances question; this design has two.

Why not a Turgenev or Chekhov lead translation, restated because it still binds. The lead reproduces 11–21 consecutive words of Garnett from the Russian alone at 8 of 8 loci (RS-20260726c), and T-svidanie-R04-v1 shares a 19-token run with Hapgood. A lead reference that is a remembered published one is not a second provenance.

3.3 The positive-control pair (R3's named materials)

T-bargamot-R04-v1 Unit B (Andreyev, «Баргамот и Гараська», 441 English words) against variant-F1 — 8 sites, two of each of O4's four documented failure types, matched to the operator at +4.54% length and 4 of 8 word-count-changing sites. Built at S050 for exactly this purpose. The Russian for Unit B (16 paragraphs, 297 words) was not stored with the artifact and is fetched once into materials/bargamot-unitB-ru.txt at build time, from the same wikisource text the artifact cites, so that the item's format is identical to every other item's.

3.4 What the jury is shown

Each item is the Russian source + two unattributed English texts. Translator names, dates, licence blocks, titles, footnote markers and section numbering are stripped; prose is unwrapped to one paragraph per line. The six-part typographic audit of the S034 build (OCR repair of trav- ersed, one-pipeline normalisation of both texts, character-inventory diff, style profile, orthographic profile, fail-loud on hyphen-space artefacts) is inherited with the texts and its log is carried in items.json. lowgrowing in Hapgood HB remains an unrepaired OCR artefact and a named confound on H-HB; this design does not authorise its repair, which R5 permits only by name before the freeze.

3.5 The sham, rebuilt (R1, R2)

Five items — Garnett HA, Garnett HB, Hapgood HA, lead K, lead L — eight sites each, free variation only, no sense shift, no archaising, sentence- and paragraph-count neutral. The provenance mix is 3 published : 2 lead, matching the targeted arm's 2 : 2 more closely than S034's 1 : 1 on two items.

The edit-kind match, declared before the sham existed and measured by the build. O4's own profile across S034's 24 targeted sites: 10 of 24 change the word count, 0 of 24 change no lexeme, +4.26% net. Criteria: each sham item ≥ 4 of 8 length-changing sites, 0 of 8 lexeme-neutral sites, net in [+2.0%, +6.0%]. Measured (materials/perturbations.md):

item net % length-changing lexeme-neutral
S-GA +11 +3.68% 7 of 8 0
S-GB +8 +2.60% 5 of 8 0
S-HA +11 +3.40% 6 of 8 0
S-K +8 +2.03% 4 of 8 0
S-L +11 +2.77% 8 of 8 0

S034's sham, for contrast: 0 of 16 length-changing, 6 of 16 lexeme-neutral. That is the defect R2 named — a floor measured on edits that cannot change length does not bound the false-alarm rate of edits that do — and it is repaired here.

The required pre-check (a), run before the freeze (analysis/metric_a_precheck.py; free, no API call). Metric A, five cues, on every non-held-out arm:

cue sham (n=5) targeted-8 (n=4) targeted-3 (n=4)
words 1.000 (larger) 1.000 (larger) 0.875
chars 1.000 1.000 0.750
paras 0.500 0.500 0.500
sents 0.500 0.500 0.500
commas 0.600 0.875 0.875

The length cue is present, at ceiling, in the sham and in the targeted arm, in the same direction. The cue cannot be removed from an operator whose job includes adding invented detail. What the repair buys is narrower than v1 claimed, and the critic's BLOCKING finding 3 is why the claim is narrowed: on S034's materials this statistic was 0.500 on the sham against 1.000 on the targeted arm, so the false-alarm floor was measured on edits that could not carry a cue the treatment did carry. It now is. That makes the sham no longer blind to the cue; it does not make the sham a measurement of the cue.

The failure mode this leaves, named because the critic named it. The sham matches the operator on edit kind (length- and lexeme-changing site rates) and deliberately not on edit type (free variation against sense damage — matching type would make it a second treatment arm, not a sham). So a sham firing may be the perceptible awkwardness of free variation rather than any length response, and the two cannot be separated by this run. What a firing licenses is therefore exactly what §6.5 says and no more: the false-alarm rate is above the floor, whatever the cue. Attributing it to length would need a length-only sham — that is pre-check (b) of framework/tierD-repaired-rules.md, it costs a paid cell, and it is not run here. Metric A is computable separability, not perceptibility and not use (note (bck)).

framework/control-arm-spec.md R5 (Metric A and a correlation at freeze time) is satisfied by the table above. v2 asserted that R1–R4 did not bind this design because it reports no stratified estimate; the ratification vote routed under §13 returned that reading INCORRECT, and §6.8 is the correction.

3.6 An instrument gate discharged before it could bite

tools/ngram_overlap.extract() — through which the contamination figure of §3.2 is computed — has a live defect: it requires the exact heading ## The translation and, having found it, skips later ## headings instead of stopping, so a translator's log can be pooled into a translation body. The standing gate is an audit of which published figures were computed through it, before any repair (wiki/backlog.md, S081). Audited this session over all 86 stored translation artifacts: 60 raise SystemExit (loud, so no silent figure), 16 extract cleanly, and 10 would pool a later section — of which three are unit headings inside the translation itself and seven are logs. Every published figure computed on any of those seven was computed by a local slicer, not by extract() (E-20260729e, E-20260731g, E-20260731f, E-20260731c, E-20260801e — the last records the reason in its own docstring), and the one cells manifest that does pass .md artifacts to extract() (E-20260725c-contamination-sweep) passes three files that all carry a --- rule before their logs. No published figure is false. The repair is not made: the defect cannot be fixed generically, because ## I inside a translation and ## Translator's log after one are indistinguishable to the function, and the honest contract is that a --- rule must terminate the body. This design's own use is asserted rather than assumed — build.py fails if the extracted body contains front matter or a log.

4. The operator

One operator: O4, accuracy — semantic errors: wrong referent (WR), wrong word sense (WS), invented detail (ID), dropped negation (DN). Style untouched. Reused verbatim from E-20260725-tierD-ladder §3 so that this run extends the instrument rather than replacing it.

Its catalogue basis (charter §5.2 forbids lead-invented operators): the project's own documented record of accuracy failure in published translations — the Shaw Christianisation (A-shaw-spider-thread), Swann (1974) on Turney, and A-chekhov-pari on Koteliansky & Murry 1915 («за пять часов» → "five minutes"; «в 12 часов дня» → "twelve o'clock midnight"; «Евангелие» → "the New Testament"). Every site names its type and its Russian basis in materials/perturbations.md. Composition constraint: at least two of each of the four types per 8-site set, which with 8 sites and 4 types forces exactly two of each. The new set (T8-L) is authored to that constraint and its sites are numbered in textual order by the build, which asserts that the written order is already textual — so the numbering cannot be chosen.

5. Dose — and the light set is drawn, not chosen

The draw rule. The three sites are combination number int(sha256("E-20260801f-tierD-run|<item id>"), 16) mod 56 in the lexicographically sorted list of 3-subsets of {1..8}. It depends on the design id and the item id only, so it is fixed before any edit can be inspected, cannot be re-rolled, and is reproducible by anyone. Drawn: T3-GA {1,4,8} · T3-HB {6,7,8} · T3-K {3,5,6} · T3-L {1,3,7}.

What this buys over S034 and what it still does not buy. S034 used the single fixed subset {2,5,7} for all three items, pre-registered before the edits existed; its own design conceded that this makes the light cell one deterministic subset, not a sample, because positional selection can correlate with narrative salience. Four independent draws over four items is a sample from the 3-of-8 subsets — R5 named this and priced it at ~$0.35 — and it is still four draws, so no dose-response curve follows and the composition of types inside each draw is reported, not controlled. T3-K's draw happens to be length-neutral (0 words) and T3-HB's falls entirely in the last third of its text; both are consequences of the draw and both are reported.

Dose is a count of sites, not a rate. References run 299–441 words, so 8 sites is 1.8–2.7 per 100 words. No claim of the form "detection is stronger on item X" is licensed.

6. Arms, jury, and the rules

6.1 Sixteen items, 32 payloads, 96 calls

stage # items arm units
1 1 PC-B prior positive control — lead Bargamot Unit B vs variant-F1 3
2 5 S-GA S-GB S-HA S-K S-L sham 15
3 2 H-HA H-HB held-out — Garnett vs Hapgood, unmodified 6
4 4 T8-GA T8-HB T8-K T8-L targeted, heavy — O4 × 8 12
5 4 T3-GA T3-HB T3-K T3-L targeted, light — O4 × drawn 3 12

The damaged reference alternates between the two published translators (Garnett carries HA, Hapgood HB) so that no result reads as "the damage was only ever applied to Garnett"; translator and passage are thereby confounded with each other, which is accepted and stated.

Order swap is mandatory on every item: each pair is dispatched twice with slots swapped. Slot preference on this exact item format was measured at P1 0.500, P2 0.600, P5 0.550 (S020).

6.2 Jury

P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P5 deepseek/deepseek-v4-pro, resolved from config/models.md at run time and logged as provenance. P3 and P4 are excluded, as in S020 and S034, so that this run is a controlled extension rather than a new instrument; this is a power limitation, not a finding about those models, and it means the US-taste-correlation revisit trigger cannot fire on this run.

Six senses, scored 1–7 for each text — accuracy, naturalness, voice, style-correspondence, literary-quality, cultural-mediation — plus a forced overall preference, no ties. Strict JSON. naturalness is put to the jury with its register frame named, verbatim from S020 §5, because S014 measured this sense as register-cued.

Judgment is not parallelized (charter §6): strictly sequential dispatch.

6.3 The unit and the firing rules

The unit is (juror × item), with the two orderings averaged within it — not the individual vote; pooling votes as if independent is pseudo-replication. A unit is +1 if the reference is preferred in both orderings, −1 if the second text is, 0 if split. Under a null of independent coin-flip preferences P(+1) = P(−1) = 0.25, P(0) = 0.5.

All probabilities below are computed by exhaustive enumeration of the trinomial in analysis/rules.py, which imports nothing from tools/.

cell units rule exact P under the null
prior positive control 3 ≥ 2 of 3 at +1 and none −1 0.109375
sham 15 three-way band, §6.5 branch 0.004193
held-out 6 ≥ 5 of 6 consistent for either text, none opposite 0.006348
targeted, heavy 12 ≥ 8 of 12 at +1 and none −1 0.000594
targeted, light 12 same 0.000594

Why 8 of 12 and not 9. The threshold is chosen by a criterion stated before the numbers were looked at: the k whose exact null probability is closest to S034's targeted rule (0.000622). That is k = 8 (0.000594); k = 9 would be 0.000122. What this holds constant is the false-positive rate, and nothing else — v1 said it made the cell "neither easier nor harder to fire than the run it extends", which the critic showed is false in both directions: at fixed α, dropping the proportion from 7/9 = 0.778 to 8/12 = 0.667 raises power against moderate effects and lowers it against overwhelming ones. The lower proportion is accepted deliberately for the first of those, and it is the direction that flatters a pass, which is why it is stated here rather than buried. The realised counts are reported so that any reader can apply any threshold.

Family-wise rate across the two targeted cells: 0.00119.

The units are not independent and the exact P values are therefore lower bounds. A cell's 12 units are 3 jurors × 4 items; a juror with a fixed taste contributes correlated units. The tabulated P is a lower bound and is reported as one. (The mechanism is fixed juror taste, not within-run learning: each payload is an independent stateless call with no conversation state.)

The per-juror robustness condition. A targeted cell may not be reported as firing unless, in addition, it fires within at least 2 of the 3 jurors taken separately — at least 2 jurors +1 on all 4 of their items with none −1. Under the null a single juror does that with P = 0.25⁴ = 0.003906. Per-juror breakdowns are reported for every cell whether or not the condition fires.

6.4 The held-out arm's power, computed rather than hoped

Against a jury that genuinely prefers one text, P(−1) held at 0.05:

true P(+1) 0.5 0.6 0.7 0.8 0.9
power 0.109 0.233 0.420 0.655 0.886

Unchanged from S034 — the arm has the same two items and the same six units, because the licensed materials are the only ones that exist. §0.3 follows arithmetically and is reported with the result.

6.5 The sham's band, repaired (R1)

On its 15 units: N₊ ≥ 9 → the jury prefers unedited text as such; every detection result is confounded with edit-presence and no detection claim is licensed from this run. N₋ ≥ 9 → the jury prefers edited text; the sham materials were not neutral and the false-alarm rate is unmeasured. Otherwise → the false-alarm rate is acceptable at this resolution.

The two branches are matched by construction and are not collapsed into one conclusion — R1's repair. Each branch has exact null probability 0.004193 (S034's design carried an unrepaired band whose branches ran 0.004639 against 0.533936, a 115-fold asymmetry, and whose lower branch duly fired).

And it is now sized for power, which nobody had computed before S040. Against a genuine 70% edit-presence bias the rule fires with probability 0.8689; the 5-of-6 rule on 6 units fired with probability 0.4202. That is what "≈15 units, not 6" buys, and it is why the sham is five items rather than two.

The sham is an upper bound on the false-alarm rate, not a neutral edit. The reference is a considered text and any eight substitutions move it off a local optimum; a perfectly neutral sham is not constructible.

6.6 Cross-sense specificity — the headline metric

Charter §5.4. Specificity fires for a targeted cell iff all three:

  1. the mean drop on accuracy exceeds the mean drop across the non-target senses excluding naturalness by ≥ 0.75 scale points; and
  2. the largest per-sense drop, excluding naturalness, is accuracy; and
  3. drop(naturalness) ≤ 0.75 scale points.

Condition 3 is a condition of the rule, not a veto applied after it. naturalness is excluded from condition 1's baseline because the design predicts it may move for reasons that would mechanically inflate the margin; both the excluding and the including figures are computed and reported, and only the excluding one enters condition 1. If condition 3 fails, the licensed statement is "specificity is not established for this cell: the target sense's margin cannot be separated from a general downward movement that also reached naturalness".

6.7 The prior positive control — R3's replacement for the unrepairable gate

§10's scale-usage gate is not imported, with or without a number. S050 established that no threshold on within-juror dispersion discriminates on this panel: all six (run × juror) cells clear the 0.75 criterion margin on their targeted arm (1.500–4.083) while the lowest observed sham dispersion is 0.143, so any threshold above it fails a demonstrably capable juror and any threshold at or below it passes everybody. R3 settles the replacement as a prior positive control on a known-difference pair, on the ground that being prior was the gate's whole function.

The rule: stage 2 is entered only if PC-B returns ≥ 2 of 3 units at +1 with none at −1.

Why ≥2 and not 3 of 3, decided before the run and on the gate's error costs. 3 of 3 has null probability 0.0156 and power 0.512 against a jury whose unit is +1 with probability 0.8 — a gate that would block a working run half the time. ≥2 of 3 has null probability 0.109375 and power 0.896 at 0.8, 0.972 at 0.9. A gate's expensive error is the false block, which costs the run; its false pass costs little, because every control downstream still applies and the sham independently licenses or withdraws the detection claim. The 0.109 false-pass rate is stated as part of the rule, not hidden in it.

What three units can and cannot resolve, after the critic's finding 1. Error rates in full: false pass 0.109375; false block 0.104 against a jury whose unit is +1 with probability 0.8, 0.028 at 0.9, 0.216 at 0.7. The critic is right that this is coarse, and the remedy it prescribed — a second item — is refused because the materials do not exist: variant-F1 is the only known-difference pair the project has, and R3 names it by name. What is accepted is the claim about resolution. This control excludes a dead instrument, not a weak one: a jury at chance, or one preferring the damaged text, or an item format that fails to elicit usable scores. A firing licenses no statement about the jury's sensitivity, and none is made from it.

Three things the positive control is not. It is not evidence for the targeted arm (§0.4). It is not a Tier D result of any kind. And "the gate would have blocked a run that then worked" would show conservatism, not invalidity — R5's constraint, which ARM-tierD-repair violated once at S050 and had refuted before running.

Its per-sense margins are reported alongside the units, because a gate that fires on the wrong sense is a fact about the instrument that the run should not discard.

6.8 The R1 stratification — added by the ratification vote, before dispatch

framework/control-arm-spec.md R1 as amended: "R1 binds every paired comparison in scope, including designs whose primary outputs are threshold counts, pass/fail tallies, or mean score drops against a pre-registered criterion." The vote (§13) returned this design's contrary reading INCORRECT in terms worth quoting — "not a rule that switches off when the outputs are threshold counts rather than CI-backed preference proportions" — and every arm here is a paired comparison of a reference against a variant of it. So:

  1. Every headline count is reported split by length-sign stratum, per stratum item count and raw aggregate, alongside the pooled figure. No pooled count is a headline on its own.
  2. The strata of this design are degenerate, and that is the disclosure R1 exists to force. Fifteen of sixteen items have Δwords > 0; the sixteenth (T3-K) has Δ = 0; there is no negative stratum at all. The failure R1 protects against — a pooled figure near chance that is the exact average of two opposite-signed halves — cannot occur here, because there is no opposite-signed half. That is a stronger statement than "we stratified", and it is only available because the split was computed. The held-out arm is the one place where the second text is longer than the reference for a reason the design did not create (Hapgood over Garnett, +8.4% and +10.7%), and it is reported as its own stratum.
  3. R4 disclosure: this design does not have R4's per-stratum item counts and does not claim a stratified estimate. R4 puts the requirement at 22–30 items per stratum; the sham has 5 items and the targeted cells 4 each. Every quantity reported is a count against a threshold fixed before the run, and none of them is an estimate of anything.
  4. F3's within-stratum magnitude check is run post-hoc on the positive stratum — the association between |Δwords| and the unit outcome, across the items of each cell — and reported as descriptive, or the result page states that it was not computed. With 4–5 items per cell it can detect only a very large association (RS-20260728c's floor at n = 12 was |ρ| ≈ 0.58), and that limitation is reported with the number.

7. Pre-registered readings, written before the run

positive control targeted (heavy) sham held-out recorded in config/models.md
does not fire — — — NOT PASSED. Run stops at stage 1. "The instrument did not separate a known-difference pair at this power; no downstream cell was dispatched." Cost of the finding: one stage
fires detection + specificity fire, and the per-juror condition holds in band does not fire Tier D reported on accuracy at 8 sites, phrased "the jury detects accuracy-damage at 8 edit sites in 299–441-word Russian→English passages, on a held-out control that did not separate an independent same-quality pair — a control with power 0.655 against an 80% preference, which therefore does not exclude a real separation". The power clause is part of the claim
fires detection + specificity fire, per-juror condition fails in band does not fire NOT PASSED. "the cell fired on pooled units but not within jurors taken separately."
fires detection + specificity fire in band fires NOT PASSED. "the rule that licenses detection also separates two independent unmodified translations; the arm did not behave at chance. Canonicity or unmeasured quality differences are not excluded." (critic finding 6: the phrase v1 used here, "same-quality", asserted inside a reading what the design holds only as an externally discharged premise)
fires detection fires, specificity fails in band does not fire NOT PASSED. "detects damage, not sense-calibrated on accuracy."
fires detection fails any any NOT PASSED. "no detection at this dose and this power."
fires any N₊ ≥ 9 any NOT PASSED. No detection claim licensed: detection is confounded with edit-presence. The cue is not identified — length, free-variation awkwardness and anything else the sham carries are not separable here (§3.5)
fires any N₋ ≥ 9 any NOT PASSED. The sham materials were not neutral; the false-alarm rate is unmeasured. Same non-identification of the cue

The light dose is reported but does not gate.

heavy detection light detection light specificity licensed dose statement
fires does not fire — "this 8-edit composite met the rule; these four drawn 3-edit composites did not." Not a dose-response, not "eight sites are stronger than three"
fires fires fires the licensed sense statement names 3 sites
fires fires fails detection is reported at 3 sites; specificity at 8 sites only
fires fires on pooled units but fails the §6.3 per-juror condition — "detection is reported at 8 sites only; the light cell fired on pooled units but not within jurors taken separately." (critic finding 4: v1's light table had no row for this, though the heavy table did)
does not fire any — the last row of the table above governs

8. Predictions, written before the run

  1. The positive control fires. Both prior Tier D runs detected O4 damage at ceiling.
  2. Detection fires at 8 sites (S020: 6/6 units on accuracy; S034: 9/9 at both doses).
  3. Specificity fires at 8 sites.
  4. Detection does not fire at 3 sites. internal-judgment-only, and the least confident. S034's nested {2,5,7} did fire at 3 sites; these are four independent draws, so a failure here would be informative about the draw rather than about the dose.
  5. The held-out arm does not fire, and no more than 3 of its 6 units are consistent for either text. The first clause alone is 0.9936 under the null and would have been a truism — the critic's finding 5. The count clause gives it a reachable failure condition (P(≥ 4 consistent) = 0.075195 under the null) without treating the arm as an estimator, which §0.3 and D-20260725-06 both forbid; the critic's prescribed remedy, predicting a preference proportion, is refused for exactly that reason.
  6. The sham lands in the middle band. S020 measured 2 of 6; S034 measured 0 of 6 at the lower branch. This is the prediction the whole rebuild is about, and it can now fail informatively in either direction, which it could not before.
  7. drop(naturalness) stays under 0.75 in both targeted cells.
  8. The two lead-provenance items behave like the published ones. Stated so it can fail: on the heavy cell, T8-K and T8-L together contain at least 4 of their 6 units at +1, and their mean accuracy drop is within 1.50 scale points of the mean of T8-GA and T8-HB. Either failing is the falsification, and the reading is that detection on this jury depends on whether the reference is published or lead-authored — a fact about provenance, and not evidence about accuracy damage in general, for the reason in §12. Both figures are reported whichever way they land.

9. Failure criteria

10. Budget

Built from max_tokens, not from an assumed output length — note (abc). Caps are sized from the measured per-call maxima of the S034 run on this exact task, which is the same instrument:

juror S034 measured max completion max_tokens here headroom
P1 886 3,500 3.9×
P2 2,839 6,000 2.1×
P5 5,893 10,000 1.7×

Worst-case input 3,500 tokens (S034 measured 2,227 max; PC-B's payload is the longest at 441 + 461 English words plus 297 Russian).

juror price in / out per M worst case per call
P1 $1.25 / $7.50 (corrected S061) $0.030625
P2 $1.50 / $7.50 $0.050250
P5 $1.65 / $3.30 — the worst plausible provider, not list $0.038775

A payload is one item in one ordering; a call is one dispatch of a payload to one juror. Per payload $0.11965; per item $0.23930.

stage items calls retry allowance worst case, reserved
1 positive control 1 6 1 $0.290
2 sham 5 30 3 $1.347
3 held-out 2 12 2 $0.579
4 targeted heavy 4 24 3 $1.108
5 targeted light 4 24 3 $1.108
total 16 96 12 $4.432

Central estimate from S034's measured per-call means on the same task: $0.0121/call × 96 = $1.16.

The abort rule is a full-stage reservation, not a running margin. Before entering any stage the runner reserves that stage's whole worst case, retries included, against the day's remaining headroom; if it does not fit, the stage is not entered and the deferral is written into NEXT.md. Dispatch order is stages 1→5 exactly as numbered, so that what truncates first is what can be lost without losing the gate: the positive control gates everything; the sham licenses or withdraws every detection claim; the held-out arm is the arm's purpose; the heavy targeted cell is the positive result that makes the held-out null interpretable; the dose axis is last.

A day with less than $4.432 runs as far as its headroom reserves and defers the rest, which is what the ordering exists to make survivable. This is the split the arm page prescribes — "split by design (fewer stages per dispatch day), never by weakening a control."

11. Verification

A verifier that recomputes every reported number from runs/ and the items manifest only, re-parsing each model's own text rather than trusting the runner's cached parse. It must:

What the cost check is: reading provider and usage.cost off each response is provenance, not invoice-level verification; the key-usage delta is the independent check and is recorded in config/budget.md alongside the per-request sum.

12. Known threats, stated in advance

13. The control-arm-spec ratification, routed in this session

framework/control-arm-spec.md carries a standing gate: "the next design that cites any rule on this page as binding must route the ratification vote in the same session, before its own dispatch, and record the outcome there." This design cites R5 as binding (§3.5) and R1–R4 as non-binding-with-a-reason. Routed before any stage was dispatched. Seat P3 x-ai/grok-4.5 — a panel member, so a non-Anthropic panel vote as charter §8 requires, and deliberately not one of this design's jurors, so the voter is not ratifying a spec it is about to be measured under. Provider xAI, stop, 73.1 s, $0.0240956. Record in ratify/.

Verdict RATIFY-WITH-AMENDMENT, amendment A1 applied verbatim to framework/control-arm-spec.md — and the vote returned this design's own reading of R1–R4 INCORRECT, which is the amendment recorded at §6.8. The design was changed before it dispatched anything. The vote's own strongest counter-argument is recorded on the spec page: that a provisional page inheriting an uncalibrated jury should have been REJECTED until re-measured, or ratified only in part.

14. Amendment record (v1 → v2, 2026-08-01)

Every change is a disposition of a finding in critic.md, applied before any call was dispatched. Nothing was amended for a reason internal to this page.

§ amendment finding
3.5 the claim that the repair lets the sham catch a length response withdrawn; replaced with what it actually buys (the sham is no longer blind to a cue the treatment carries) plus the named non-identification failure mode 3, BLOCKING
6.3 "neither easier nor harder to fire" withdrawn as false in both directions; replaced with what matching α does and does not hold constant 2
6.7 both error rates printed; the resolution claim narrowed to excludes a dead instrument, not a weak one; the prescribed second item refused with the reason (the materials do not exist) 1
7 "same-quality" removed from the held-out firing reading, in the critic's own words; the two sham rows given the cue-non-identification clause; a per-juror-fails row added to the light table 6, BLOCKING · 3 · 4
8 prediction 5 given a reachable failure condition that does not treat the held-out arm as an estimator; prediction 8's reading narrowed to provenance 5 · 7
12 the author-wrote-reference-damage-and-design threat added, with the strongest inference it invalidates named 7

Accepted in part, with the refused remedy reasoned: findings 1 (no second known-difference pair exists) and 5 (a preference-proportion prediction is forbidden by §0.3 and D-20260725-06). Nothing was rebutted.

v2 → v2.1, the ratification amendment (same session, still before dispatch)

§ amendment source
3.5, 6.8, 13 the claim that control-arm-spec R1–R4 do not bind this design withdrawn; §6.8 added with the length-sign split, the degenerate-strata disclosure, the R4 shortfall statement and F3's within-stratum check the routed ratification vote, RATIFY-WITH-AMENDMENT + reading INCORRECT

Two independent voices changed this design before it cost anything: the critic (seven findings) and the ratification vote (one overturned reading). Neither was the lead.