Repository path: workshop/experiments/E-20260801f-tierD-run/design.md · rendered 2026-09-09
Page metadata (front matter)
Frozen design (v2) — Tier D under the repaired rules
v1 was frozen, then put through the independent pre-run critic pass (critic.md, verdict
NEEDS-AMENDMENT, seven findings, two BLOCKING). This is the amended version: all seven accepted,
five in full and two in part with the refused remedy reasoned. Dispositions are tabulated in
critic.md and dated in §14. A run may proceed only against this version, and no stage had been
dispatched when the amendments were made.
ARM-tierD-run step 1. This design imports framework/tierD-repaired-rules.md whole and does
not re-derive it. That page exists because three sessions of repair work (ARM-tierD-repair,
S040–S055) closed with the instruction "import this block; do not re-derive it" and no design
imported it for twenty-six sessions.
No senses: field: this page designs an evaluation of the jury, not of a translation, and
asserts no evaluative claim about any translation.
0. What this run can and cannot be — read first
Tier D is NOT PASSED and has been since S034. Nothing on this page changes that; only the run can, and only on the terms below.
- It can report Tier D on
accuracyand on nothing else. Condition (i) is ratified for the Garnett/Hapgood pair for the Memoirs of a Sportsman cycle only and foraccuracyandcultural-mediationonly (D-20260725-07). The other four senses are inadmissible on this pair: two contemporaneous reviews agree on a directional Garnett advantage on English prose, and a held-out control on a pair with a documented gap on the measured sense imports the bias the control exists to exclude. All six senses are still scored — cross-sense specificity is the headline metric and needs the untargeted senses. - It carries R4 clause 4 forward explicitly. The evidence base licensing clauses 1–2 of the held-out materials contains no printed style exhibit from the Memoirs cycle at all: all eighteen paired extracts the 1904 Nation review prints are cited to A Nobleman's Nest, the work clause 3 refuses to license. This design does not need clause 3 — it takes the licensed senses and nothing else — and it states clause 4 rather than inheriting clause 2 as though its evidence were work-matched.
- A non-firing held-out arm licenses exactly one sentence — "this 5-of-6 rule did not detect separation" — and licenses none of: chance behaviour, parity, equivalence, or absence of the canonicity, OCR, orthography and length confounds of §12. Its power against an 80% preference is 0.6554 (§6.4). The control can kill a detection claim; it cannot discharge itself.
- The prior positive control is a gate, not evidence. It uses the same operator as the targeted arm, so it is a rehearsal of that arm on different materials, not independent support for it. Its firing licenses nothing whatever about the targeted arm; its failing stops the run. §6.7 states this in the rule itself rather than in a caveat.
- A FAIL is completion.
ARM-tierD-run§Done when: a FAIL closes the armresolvedwith a statement of what a redesign would change or that the approach is exhausted, and may not spawn a repair arm without Tom (wiki/reassessment-2026-08-01.md§6).
1. Question
On which goodness senses, and at what dose, does the panel detect deliberate, sense-targeted degradation — under the repaired rules, with all four controls (prior positive, sham, held-out, cross-sense) measured on the same materials by the same jurors?
2. What is imported, what is carried, and what is new
Imported whole from framework/tierD-repaired-rules.md |
R1 the sham's decision rule and its power requirement · R2 edit-kind matching · R3 the stage-1 gate's replacement · R4 the held-out materials' licence and its four clauses · R5's five surviving constraints · the required pre-check |
Carried byte-identical from E-20260726d-tierD-heldout (S034) |
the five built reference texts, their sources, the OCR repair, the one-pipeline typographic normalisation, and the three O4 edit tables. materials/build.py asserts byte equality against that frozen manifest (SHA-256 recorded in materials/items.json) rather than re-deriving a build whose verifier already passed 71 checks |
| New this session | a rebuilt sham: 5 items / 15 units instead of 2 items / 6, and matched to the operator on edit kind (R1, R2) · a second lead-provenance reference (T-les-shumit-R04-v1, translated in session, contamination measured clean) and its O4 table · the prior positive control as stage 1 (R3) · a drawn light dose instead of one fixed positional subset (R5) |
What the S034 verdict is not. R5: the repaired rules apply to future runs only; nothing here reopens S034, and no figure of S034's is restated as though it had been re-earned.
3. Materials
Everything is public domain or lead-authored and already in the repository. Nothing is fetched at run time.
3.1 The held-out passages (carried)
Turgenev, «Свидание» (Записки охотника), opening, 468 Russian words, split at the landmark named
in the S034 design and independently reconstructed there (analysis/split_check.py). Garnett 1897
and Hapgood 1904, stored whole from the S024 audit. Item H-HA 299 / 324 words, H-HB 308 /
341. Conditions (i), (ii), (iii) discharged at D-20260725-07, E-20260725-svidanie-audit and
RS-20260726b-baseline-dependence (zero shared ≥12-token runs) respectively.
3.2 The two lead-provenance references
Charter §5.1: references of both provenances. S034 had one lead item and therefore three units on the whole both-provenances question; this design has two.
T-son-makara-R04-v1(carried) — Korolenko, «Сон Макара» §I, 270 Russian → 394 English. Contamination gate at S034: 3 shared 7-grams, 0 twelve-grams, longest run 8 against the whole Fell 1916 story.T-les-shumit-R04-v1(new, this session) — Korolenko, «Лес шумит», opening, 276 Russian → 397 English. Same narrator-with-a-gun register as the Sportsman's Sketches passages. Contamination gate run before this design selected it and before the comparator was opened: 0 shared 7-grams, 0 twelve-grams, longest run 6 tokens against the whole 9,344-word Fell 1916 The Murmuring Forest — the cleanest measurement in the project's record against a full-length comparator. One priming event is declared on the artifact: the story's English title as Fell gives it was seen while locating the comparator.
Why not a Turgenev or Chekhov lead translation, restated because it still binds. The lead
reproduces 11–21 consecutive words of Garnett from the Russian alone at 8 of 8 loci
(RS-20260726c), and T-svidanie-R04-v1 shares a 19-token run with Hapgood. A lead reference
that is a remembered published one is not a second provenance.
3.3 The positive-control pair (R3's named materials)
T-bargamot-R04-v1 Unit B (Andreyev, «Баргамот и Гараська», 441 English words) against
variant-F1 — 8 sites, two of each of O4's four documented failure types, matched to the
operator at +4.54% length and 4 of 8 word-count-changing sites. Built at S050 for exactly this
purpose. The Russian for Unit B (16 paragraphs, 297 words) was not stored with the artifact and is
fetched once into materials/bargamot-unitB-ru.txt at build time, from the same wikisource text
the artifact cites, so that the item's format is identical to every other item's.
3.4 What the jury is shown
Each item is the Russian source + two unattributed English texts. Translator names, dates,
licence blocks, titles, footnote markers and section numbering are stripped; prose is unwrapped to
one paragraph per line. The six-part typographic audit of the S034 build (OCR repair of
trav- ersed, one-pipeline normalisation of both texts, character-inventory diff, style profile,
orthographic profile, fail-loud on hyphen-space artefacts) is inherited with the texts and its log
is carried in items.json. lowgrowing in Hapgood HB remains an unrepaired OCR artefact and a
named confound on H-HB; this design does not authorise its repair, which R5 permits only by
name before the freeze.
3.5 The sham, rebuilt (R1, R2)
Five items — Garnett HA, Garnett HB, Hapgood HA, lead K, lead L — eight sites each, free variation only, no sense shift, no archaising, sentence- and paragraph-count neutral. The provenance mix is 3 published : 2 lead, matching the targeted arm's 2 : 2 more closely than S034's 1 : 1 on two items.
The edit-kind match, declared before the sham existed and measured by the build. O4's own
profile across S034's 24 targeted sites: 10 of 24 change the word count, 0 of 24 change no lexeme,
+4.26% net. Criteria: each sham item ≥ 4 of 8 length-changing sites, 0 of 8
lexeme-neutral sites, net in [+2.0%, +6.0%]. Measured (materials/perturbations.md):
| item | net | % | length-changing | lexeme-neutral |
|---|---|---|---|---|
| S-GA | +11 | +3.68% | 7 of 8 | 0 |
| S-GB | +8 | +2.60% | 5 of 8 | 0 |
| S-HA | +11 | +3.40% | 6 of 8 | 0 |
| S-K | +8 | +2.03% | 4 of 8 | 0 |
| S-L | +11 | +2.77% | 8 of 8 | 0 |
S034's sham, for contrast: 0 of 16 length-changing, 6 of 16 lexeme-neutral. That is the defect R2 named — a floor measured on edits that cannot change length does not bound the false-alarm rate of edits that do — and it is repaired here.
The required pre-check (a), run before the freeze (analysis/metric_a_precheck.py; free, no
API call). Metric A, five cues, on every non-held-out arm:
| cue | sham (n=5) | targeted-8 (n=4) | targeted-3 (n=4) |
|---|---|---|---|
| words | 1.000 (larger) | 1.000 (larger) | 0.875 |
| chars | 1.000 | 1.000 | 0.750 |
| paras | 0.500 | 0.500 | 0.500 |
| sents | 0.500 | 0.500 | 0.500 |
| commas | 0.600 | 0.875 | 0.875 |
The length cue is present, at ceiling, in the sham and in the targeted arm, in the same direction. The cue cannot be removed from an operator whose job includes adding invented detail. What the repair buys is narrower than v1 claimed, and the critic's BLOCKING finding 3 is why the claim is narrowed: on S034's materials this statistic was 0.500 on the sham against 1.000 on the targeted arm, so the false-alarm floor was measured on edits that could not carry a cue the treatment did carry. It now is. That makes the sham no longer blind to the cue; it does not make the sham a measurement of the cue.
The failure mode this leaves, named because the critic named it. The sham matches the operator
on edit kind (length- and lexeme-changing site rates) and deliberately not on edit type (free
variation against sense damage — matching type would make it a second treatment arm, not a sham).
So a sham firing may be the perceptible awkwardness of free variation rather than any length
response, and the two cannot be separated by this run. What a firing licenses is therefore
exactly what §6.5 says and no more: the false-alarm rate is above the floor, whatever the cue.
Attributing it to length would need a length-only sham — that is pre-check (b) of
framework/tierD-repaired-rules.md, it costs a paid cell, and it is not run here.
Metric A is computable separability, not perceptibility and not use (note (bck)).
framework/control-arm-spec.md R5 (Metric A and a correlation at freeze time) is satisfied by the
table above. v2 asserted that R1–R4 did not bind this design because it reports no stratified
estimate; the ratification vote routed under §13 returned that reading INCORRECT, and §6.8 is the
correction.
3.6 An instrument gate discharged before it could bite
tools/ngram_overlap.extract() — through which the contamination figure of §3.2 is computed — has
a live defect: it requires the exact heading ## The translation and, having found it, skips
later ## headings instead of stopping, so a translator's log can be pooled into a translation
body. The standing gate is an audit of which published figures were computed through it, before
any repair (wiki/backlog.md, S081). Audited this session over all 86 stored translation
artifacts: 60 raise SystemExit (loud, so no silent figure), 16 extract cleanly, and 10 would
pool a later section — of which three are unit headings inside the translation itself and seven
are logs. Every published figure computed on any of those seven was computed by a local slicer,
not by extract() (E-20260729e, E-20260731g, E-20260731f, E-20260731c,
E-20260801e — the last records the reason in its own docstring), and the one cells manifest that
does pass .md artifacts to extract() (E-20260725c-contamination-sweep) passes three files
that all carry a --- rule before their logs. No published figure is false. The repair is not
made: the defect cannot be fixed generically, because ## I inside a translation and
## Translator's log after one are indistinguishable to the function, and the honest contract is
that a --- rule must terminate the body. This design's own use is asserted rather than assumed —
build.py fails if the extracted body contains front matter or a log.
4. The operator
One operator: O4, accuracy — semantic errors: wrong referent (WR), wrong word sense (WS),
invented detail (ID), dropped negation (DN). Style untouched. Reused verbatim from
E-20260725-tierD-ladder §3 so that this run extends the instrument rather than replacing it.
Its catalogue basis (charter §5.2 forbids lead-invented operators): the project's own
documented record of accuracy failure in published translations — the Shaw Christianisation
(A-shaw-spider-thread), Swann (1974) on Turney, and A-chekhov-pari on Koteliansky & Murry 1915
(«за пять часов» → "five minutes"; «в 12 часов дня» → "twelve o'clock midnight"; «Евангелие» →
"the New Testament"). Every site names its type and its Russian basis in
materials/perturbations.md. Composition constraint: at least two of each of the four types
per 8-site set, which with 8 sites and 4 types forces exactly two of each. The new set (T8-L) is
authored to that constraint and its sites are numbered in textual order by the build, which
asserts that the written order is already textual — so the numbering cannot be chosen.
5. Dose — and the light set is drawn, not chosen
- Heavy — 8 sites, S020's and S034's dose.
- Light — 3 sites, drawn per item by a rule fixed in this page.
The draw rule. The three sites are combination number
int(sha256("E-20260801f-tierD-run|<item id>"), 16) mod 56 in the lexicographically sorted list of
3-subsets of {1..8}. It depends on the design id and the item id only, so it is fixed before
any edit can be inspected, cannot be re-rolled, and is reproducible by anyone. Drawn:
T3-GA {1,4,8} · T3-HB {6,7,8} · T3-K {3,5,6} · T3-L {1,3,7}.
What this buys over S034 and what it still does not buy. S034 used the single fixed subset {2,5,7} for all three items, pre-registered before the edits existed; its own design conceded that this makes the light cell one deterministic subset, not a sample, because positional selection can correlate with narrative salience. Four independent draws over four items is a sample from the 3-of-8 subsets — R5 named this and priced it at ~$0.35 — and it is still four draws, so no dose-response curve follows and the composition of types inside each draw is reported, not controlled. T3-K's draw happens to be length-neutral (0 words) and T3-HB's falls entirely in the last third of its text; both are consequences of the draw and both are reported.
Dose is a count of sites, not a rate. References run 299–441 words, so 8 sites is 1.8–2.7 per 100 words. No claim of the form "detection is stronger on item X" is licensed.
6. Arms, jury, and the rules
6.1 Sixteen items, 32 payloads, 96 calls
| stage | # | items | arm | units |
|---|---|---|---|---|
| 1 | 1 | PC-B |
prior positive control — lead Bargamot Unit B vs variant-F1 |
3 |
| 2 | 5 | S-GA S-GB S-HA S-K S-L |
sham | 15 |
| 3 | 2 | H-HA H-HB |
held-out — Garnett vs Hapgood, unmodified | 6 |
| 4 | 4 | T8-GA T8-HB T8-K T8-L |
targeted, heavy — O4 × 8 | 12 |
| 5 | 4 | T3-GA T3-HB T3-K T3-L |
targeted, light — O4 × drawn 3 | 12 |
The damaged reference alternates between the two published translators (Garnett carries HA, Hapgood HB) so that no result reads as "the damage was only ever applied to Garnett"; translator and passage are thereby confounded with each other, which is accepted and stated.
Order swap is mandatory on every item: each pair is dispatched twice with slots swapped. Slot preference on this exact item format was measured at P1 0.500, P2 0.600, P5 0.550 (S020).
6.2 Jury
P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P5 deepseek/deepseek-v4-pro,
resolved from config/models.md at run time and logged as provenance. P3 and P4 are excluded, as
in S020 and S034, so that this run is a controlled extension rather than a new instrument; this
is a power limitation, not a finding about those models, and it means the US-taste-correlation
revisit trigger cannot fire on this run.
Six senses, scored 1–7 for each text — accuracy, naturalness, voice,
style-correspondence, literary-quality, cultural-mediation — plus a forced overall
preference, no ties. Strict JSON. naturalness is put to the jury with its register frame
named, verbatim from S020 §5, because S014 measured this sense as register-cued.
Judgment is not parallelized (charter §6): strictly sequential dispatch.
6.3 The unit and the firing rules
The unit is (juror × item), with the two orderings averaged within it — not the individual vote; pooling votes as if independent is pseudo-replication. A unit is +1 if the reference is preferred in both orderings, −1 if the second text is, 0 if split. Under a null of independent coin-flip preferences P(+1) = P(−1) = 0.25, P(0) = 0.5.
All probabilities below are computed by exhaustive enumeration of the trinomial in
analysis/rules.py, which imports nothing from tools/.
| cell | units | rule | exact P under the null |
|---|---|---|---|
| prior positive control | 3 | ≥ 2 of 3 at +1 and none −1 | 0.109375 |
| sham | 15 | three-way band, §6.5 | branch 0.004193 |
| held-out | 6 | ≥ 5 of 6 consistent for either text, none opposite | 0.006348 |
| targeted, heavy | 12 | ≥ 8 of 12 at +1 and none −1 | 0.000594 |
| targeted, light | 12 | same | 0.000594 |
Why 8 of 12 and not 9. The threshold is chosen by a criterion stated before the numbers were looked at: the k whose exact null probability is closest to S034's targeted rule (0.000622). That is k = 8 (0.000594); k = 9 would be 0.000122. What this holds constant is the false-positive rate, and nothing else — v1 said it made the cell "neither easier nor harder to fire than the run it extends", which the critic showed is false in both directions: at fixed α, dropping the proportion from 7/9 = 0.778 to 8/12 = 0.667 raises power against moderate effects and lowers it against overwhelming ones. The lower proportion is accepted deliberately for the first of those, and it is the direction that flatters a pass, which is why it is stated here rather than buried. The realised counts are reported so that any reader can apply any threshold.
Family-wise rate across the two targeted cells: 0.00119.
The units are not independent and the exact P values are therefore lower bounds. A cell's 12 units are 3 jurors × 4 items; a juror with a fixed taste contributes correlated units. The tabulated P is a lower bound and is reported as one. (The mechanism is fixed juror taste, not within-run learning: each payload is an independent stateless call with no conversation state.)
The per-juror robustness condition. A targeted cell may not be reported as firing unless, in addition, it fires within at least 2 of the 3 jurors taken separately — at least 2 jurors +1 on all 4 of their items with none −1. Under the null a single juror does that with P = 0.25⁴ = 0.003906. Per-juror breakdowns are reported for every cell whether or not the condition fires.
6.4 The held-out arm's power, computed rather than hoped
Against a jury that genuinely prefers one text, P(−1) held at 0.05:
| true P(+1) | 0.5 | 0.6 | 0.7 | 0.8 | 0.9 |
|---|---|---|---|---|---|
| power | 0.109 | 0.233 | 0.420 | 0.655 | 0.886 |
Unchanged from S034 — the arm has the same two items and the same six units, because the licensed materials are the only ones that exist. §0.3 follows arithmetically and is reported with the result.
6.5 The sham's band, repaired (R1)
On its 15 units: N₊ ≥ 9 → the jury prefers unedited text as such; every detection result is confounded with edit-presence and no detection claim is licensed from this run. N₋ ≥ 9 → the jury prefers edited text; the sham materials were not neutral and the false-alarm rate is unmeasured. Otherwise → the false-alarm rate is acceptable at this resolution.
The two branches are matched by construction and are not collapsed into one conclusion — R1's repair. Each branch has exact null probability 0.004193 (S034's design carried an unrepaired band whose branches ran 0.004639 against 0.533936, a 115-fold asymmetry, and whose lower branch duly fired).
And it is now sized for power, which nobody had computed before S040. Against a genuine 70% edit-presence bias the rule fires with probability 0.8689; the 5-of-6 rule on 6 units fired with probability 0.4202. That is what "≈15 units, not 6" buys, and it is why the sham is five items rather than two.
The sham is an upper bound on the false-alarm rate, not a neutral edit. The reference is a considered text and any eight substitutions move it off a local optimum; a perfectly neutral sham is not constructible.
6.6 Cross-sense specificity — the headline metric
Charter §5.4. Specificity fires for a targeted cell iff all three:
- the mean drop on
accuracyexceeds the mean drop across the non-target senses excludingnaturalnessby ≥ 0.75 scale points; and - the largest per-sense drop, excluding
naturalness, isaccuracy; and drop(naturalness)≤ 0.75 scale points.
Condition 3 is a condition of the rule, not a veto applied after it. naturalness is excluded from
condition 1's baseline because the design predicts it may move for reasons that would mechanically
inflate the margin; both the excluding and the including figures are computed and reported, and
only the excluding one enters condition 1. If condition 3 fails, the licensed statement is
"specificity is not established for this cell: the target sense's margin cannot be separated from
a general downward movement that also reached naturalness".
6.7 The prior positive control — R3's replacement for the unrepairable gate
§10's scale-usage gate is not imported, with or without a number. S050 established that no threshold on within-juror dispersion discriminates on this panel: all six (run × juror) cells clear the 0.75 criterion margin on their targeted arm (1.500–4.083) while the lowest observed sham dispersion is 0.143, so any threshold above it fails a demonstrably capable juror and any threshold at or below it passes everybody. R3 settles the replacement as a prior positive control on a known-difference pair, on the ground that being prior was the gate's whole function.
The rule: stage 2 is entered only if PC-B returns ≥ 2 of 3 units at +1 with none at −1.
Why ≥2 and not 3 of 3, decided before the run and on the gate's error costs. 3 of 3 has null probability 0.0156 and power 0.512 against a jury whose unit is +1 with probability 0.8 — a gate that would block a working run half the time. ≥2 of 3 has null probability 0.109375 and power 0.896 at 0.8, 0.972 at 0.9. A gate's expensive error is the false block, which costs the run; its false pass costs little, because every control downstream still applies and the sham independently licenses or withdraws the detection claim. The 0.109 false-pass rate is stated as part of the rule, not hidden in it.
What three units can and cannot resolve, after the critic's finding 1. Error rates in full:
false pass 0.109375; false block 0.104 against a jury whose unit is +1 with probability 0.8,
0.028 at 0.9, 0.216 at 0.7. The critic is right that this is coarse, and the remedy it
prescribed — a second item — is refused because the materials do not exist: variant-F1 is the
only known-difference pair the project has, and R3 names it by name. What is accepted is the claim
about resolution. This control excludes a dead instrument, not a weak one: a jury at chance, or
one preferring the damaged text, or an item format that fails to elicit usable scores. A firing
licenses no statement about the jury's sensitivity, and none is made from it.
Three things the positive control is not. It is not evidence for the targeted arm (§0.4). It is
not a Tier D result of any kind. And "the gate would have blocked a run that then worked" would
show conservatism, not invalidity — R5's constraint, which ARM-tierD-repair violated once at
S050 and had refuted before running.
Its per-sense margins are reported alongside the units, because a gate that fires on the wrong sense is a fact about the instrument that the run should not discard.
6.8 The R1 stratification — added by the ratification vote, before dispatch
framework/control-arm-spec.md R1 as amended: "R1 binds every paired comparison in scope,
including designs whose primary outputs are threshold counts, pass/fail tallies, or mean score
drops against a pre-registered criterion." The vote (§13) returned this design's contrary reading
INCORRECT in terms worth quoting — "not a rule that switches off when the outputs are
threshold counts rather than CI-backed preference proportions" — and every arm here is a paired
comparison of a reference against a variant of it. So:
- Every headline count is reported split by length-sign stratum, per stratum item count and raw aggregate, alongside the pooled figure. No pooled count is a headline on its own.
- The strata of this design are degenerate, and that is the disclosure R1 exists to force.
Fifteen of sixteen items have Δwords > 0; the sixteenth (
T3-K) has Δ = 0; there is no negative stratum at all. The failure R1 protects against — a pooled figure near chance that is the exact average of two opposite-signed halves — cannot occur here, because there is no opposite-signed half. That is a stronger statement than "we stratified", and it is only available because the split was computed. The held-out arm is the one place where the second text is longer than the reference for a reason the design did not create (Hapgood over Garnett, +8.4% and +10.7%), and it is reported as its own stratum. - R4 disclosure: this design does not have R4's per-stratum item counts and does not claim a stratified estimate. R4 puts the requirement at 22–30 items per stratum; the sham has 5 items and the targeted cells 4 each. Every quantity reported is a count against a threshold fixed before the run, and none of them is an estimate of anything.
- F3's within-stratum magnitude check is run post-hoc on the positive stratum — the
association between |Δwords| and the unit outcome, across the items of each cell — and reported
as descriptive, or the result page states that it was not computed. With 4–5 items per cell it
can detect only a very large association (
RS-20260728c's floor at n = 12 was |ρ| ≈ 0.58), and that limitation is reported with the number.
7. Pre-registered readings, written before the run
| positive control | targeted (heavy) | sham | held-out | recorded in config/models.md |
|---|---|---|---|---|
| does not fire | — | — | — | NOT PASSED. Run stops at stage 1. "The instrument did not separate a known-difference pair at this power; no downstream cell was dispatched." Cost of the finding: one stage |
| fires | detection + specificity fire, and the per-juror condition holds | in band | does not fire | Tier D reported on accuracy at 8 sites, phrased "the jury detects accuracy-damage at 8 edit sites in 299–441-word Russian→English passages, on a held-out control that did not separate an independent same-quality pair — a control with power 0.655 against an 80% preference, which therefore does not exclude a real separation". The power clause is part of the claim |
| fires | detection + specificity fire, per-juror condition fails | in band | does not fire | NOT PASSED. "the cell fired on pooled units but not within jurors taken separately." |
| fires | detection + specificity fire | in band | fires | NOT PASSED. "the rule that licenses detection also separates two independent unmodified translations; the arm did not behave at chance. Canonicity or unmeasured quality differences are not excluded." (critic finding 6: the phrase v1 used here, "same-quality", asserted inside a reading what the design holds only as an externally discharged premise) |
| fires | detection fires, specificity fails | in band | does not fire | NOT PASSED. "detects damage, not sense-calibrated on accuracy." |
| fires | detection fails | any | any | NOT PASSED. "no detection at this dose and this power." |
| fires | any | N₊ ≥ 9 | any | NOT PASSED. No detection claim licensed: detection is confounded with edit-presence. The cue is not identified — length, free-variation awkwardness and anything else the sham carries are not separable here (§3.5) |
| fires | any | N₋ ≥ 9 | any | NOT PASSED. The sham materials were not neutral; the false-alarm rate is unmeasured. Same non-identification of the cue |
The light dose is reported but does not gate.
| heavy detection | light detection | light specificity | licensed dose statement |
|---|---|---|---|
| fires | does not fire | — | "this 8-edit composite met the rule; these four drawn 3-edit composites did not." Not a dose-response, not "eight sites are stronger than three" |
| fires | fires | fires | the licensed sense statement names 3 sites |
| fires | fires | fails | detection is reported at 3 sites; specificity at 8 sites only |
| fires | fires on pooled units but fails the §6.3 per-juror condition | — | "detection is reported at 8 sites only; the light cell fired on pooled units but not within jurors taken separately." (critic finding 4: v1's light table had no row for this, though the heavy table did) |
| does not fire | any | — | the last row of the table above governs |
8. Predictions, written before the run
- The positive control fires. Both prior Tier D runs detected O4 damage at ceiling.
- Detection fires at 8 sites (S020: 6/6 units on
accuracy; S034: 9/9 at both doses). - Specificity fires at 8 sites.
- Detection does not fire at 3 sites.
internal-judgment-only, and the least confident. S034's nested {2,5,7} did fire at 3 sites; these are four independent draws, so a failure here would be informative about the draw rather than about the dose. - The held-out arm does not fire, and no more than 3 of its 6 units are consistent for either
text. The first clause alone is 0.9936 under the null and would have been a truism — the
critic's finding 5. The count clause gives it a reachable failure condition (P(≥ 4 consistent) =
0.075195 under the null) without treating the arm as an estimator, which §0.3 and
D-20260725-06both forbid; the critic's prescribed remedy, predicting a preference proportion, is refused for exactly that reason. - The sham lands in the middle band. S020 measured 2 of 6; S034 measured 0 of 6 at the lower branch. This is the prediction the whole rebuild is about, and it can now fail informatively in either direction, which it could not before.
drop(naturalness)stays under 0.75 in both targeted cells.- The two lead-provenance items behave like the published ones. Stated so it can fail: on the
heavy cell,
T8-KandT8-Ltogether contain at least 4 of their 6 units at +1, and their meanaccuracydrop is within 1.50 scale points of the mean ofT8-GAandT8-HB. Either failing is the falsification, and the reading is that detection on this jury depends on whether the reference is published or lead-authored — a fact about provenance, and not evidence about accuracy damage in general, for the reason in §12. Both figures are reported whichever way they land.
9. Failure criteria
- > 10% of calls failing or unparseable after one retry → reported as a failed run, not patched.
- The positive control not firing → the run stops at stage 1 and reports that.
- Sham out of band → §6.5 applies; detection claims are withdrawn.
- Post-run verification failing to reproduce a reported number → the number is withdrawn.
- Any built text containing a hyphen-space artefact → the build fails and nothing dispatches.
- Reconstruction check failing → the verifier rebuilds each variant from its reference plus the logged rows and asserts byte equality; an edit present but unlogged fails the run.
- No order-flip threshold. Replaced by a pre-committed P5-excluded sensitivity analysis reported alongside the primary (P5's order-flip rate was 0.44 at S014).
- A raised
max_tokensis an amendment, declared before the call, with the reason, inconfig/budget.md— note (bgt)/(b), which fired four times in the last five sessions.
10. Budget
Built from max_tokens, not from an assumed output length — note (abc). Caps are sized from
the measured per-call maxima of the S034 run on this exact task, which is the same instrument:
| juror | S034 measured max completion | max_tokens here |
headroom |
|---|---|---|---|
| P1 | 886 | 3,500 | 3.9× |
| P2 | 2,839 | 6,000 | 2.1× |
| P5 | 5,893 | 10,000 | 1.7× |
Worst-case input 3,500 tokens (S034 measured 2,227 max; PC-B's payload is the longest at 441 + 461 English words plus 297 Russian).
| juror | price in / out per M | worst case per call |
|---|---|---|
| P1 | $1.25 / $7.50 (corrected S061) | $0.030625 |
| P2 | $1.50 / $7.50 | $0.050250 |
| P5 | $1.65 / $3.30 — the worst plausible provider, not list | $0.038775 |
A payload is one item in one ordering; a call is one dispatch of a payload to one juror. Per payload $0.11965; per item $0.23930.
| stage | items | calls | retry allowance | worst case, reserved |
|---|---|---|---|---|
| 1 positive control | 1 | 6 | 1 | $0.290 |
| 2 sham | 5 | 30 | 3 | $1.347 |
| 3 held-out | 2 | 12 | 2 | $0.579 |
| 4 targeted heavy | 4 | 24 | 3 | $1.108 |
| 5 targeted light | 4 | 24 | 3 | $1.108 |
| total | 16 | 96 | 12 | $4.432 |
Central estimate from S034's measured per-call means on the same task: $0.0121/call × 96 = $1.16.
The abort rule is a full-stage reservation, not a running margin. Before entering any stage the
runner reserves that stage's whole worst case, retries included, against the day's remaining
headroom; if it does not fit, the stage is not entered and the deferral is written into
NEXT.md. Dispatch order is stages 1→5 exactly as numbered, so that what truncates first is what
can be lost without losing the gate: the positive control gates everything; the sham licenses or
withdraws every detection claim; the held-out arm is the arm's purpose; the heavy targeted cell is
the positive result that makes the held-out null interpretable; the dose axis is last.
A day with less than $4.432 runs as far as its headroom reserves and defers the rest, which is what the ordering exists to make survivable. This is the split the arm page prescribes — "split by design (fewer stages per dispatch day), never by weakening a control."
11. Verification
A verifier that recomputes every reported number from runs/ and the items manifest only,
re-parsing each model's own text rather than trusting the runner's cached parse. It must:
- import nothing from
tools/(the A11 rule: a 218-check pass once found nothing while a shared function was broken in three ways, because both implementations called it); - rebuild every variant from reference + logged rows and assert byte equality;
- recompute all null probabilities and the power tables of §§6.3–6.5, 6.7 by an independent path;
- assert each carried reference text is byte-identical to the S034 manifest at its recorded SHA-256, and each new reference byte-identical to its translation artifact;
- recompute the §3.5 edit-kind measurements and the §3.6 Metric A table;
- recompute the light draw from the rule and assert it matches the manifest — the one place a post-hoc choice could hide;
- print the resolved provider and billed cost per call, read off each response (note (x));
- carry mutation tests: at least six deliberate corruptions of stored bodies, each asserted to change bytes on disk and each caught (note (bgu), standing since S079).
What the cost check is: reading provider and usage.cost off each response is provenance,
not invoice-level verification; the key-usage delta is the independent check and is recorded in
config/budget.md alongside the per-request sum.
12. Known threats, stated in advance
- Canonicity is not controlled and cannot be. Garnett is the default English Turgenev; Hapgood is not. The largest single threat to the held-out arm.
- The held-out control is under-powered (§6.4), and it points the other way: it makes a pass easier, not harder.
- The length cue is at ceiling in both the sham and the targeted arm (§3.5). The repair makes the sham able to catch its use; it does not establish that the jury does or does not use it, and pre-check (b) is not run.
- The perturbations and the sham are lead-authored. They control for edit-presence, not for lead-shapedness, and they come from the same hand that wrote this design.
- On the two lead-provenance items the design author wrote the reference, the damage and the
design (critic finding 7). The strongest inference that fact invalidates: detection on
T8-K/T8-Lis not evidence that the jury detects accuracy damage as a category, because it cannot be separated from the jury detecting this author's deliberate deviations from this author's own baseline. Generalisation rests on the published-provenance items, where the reference is Garnett's or Hapgood's and only the damage is the lead's — and there the lead still wrote the damage. Prediction 8 is therefore about provenance, not about the operator. - Hapgood's text is an OCR'd scan and Garnett's is not; mitigated and measured, not eliminated.
lowgrowingremains a named confound onH-HB. - The source's presence changes the confound rather than removing it: with the Russian in the
payload the jury can reward literal alignment where a freer rendering is equally accurate;
without it, it would reward English fluency.
accuracyis unscoreable without it. - The held-out texts differ in length by 8–11% within each item, permitting preference for compression or apparent completeness. Not fixable without editing the texts the control is about.
- Style and orthography are live discriminators independent of quality — 1897 British against
1904 American. The
naturalnessregister frame does not neutralise them for the forced preference or for the other five senses. - The positive control is one item and three units (§6.7), and it is the same operator as the arm it gates (§0.4).
- R4 clause 4: the exclusion of the four prose senses rests, in printed exhibits, entirely on the work clause 3 refuses to license (§0.2).
- Three jurors. This distinguishes near-ceiling from near-chance and nothing finer.
- Detection is not judgment. Firing on
accuracywould still not license ranking two good translations. That is Tier P, which ran once and failed.
13. The control-arm-spec ratification, routed in this session
framework/control-arm-spec.md carries a standing gate: "the next design that cites any rule on
this page as binding must route the ratification vote in the same session, before its own dispatch,
and record the outcome there." This design cites R5 as binding (§3.5) and R1–R4 as
non-binding-with-a-reason. Routed before any stage was dispatched. Seat P3 x-ai/grok-4.5 — a panel member, so a
non-Anthropic panel vote as charter §8 requires, and deliberately not one of this design's jurors,
so the voter is not ratifying a spec it is about to be measured under. Provider xAI, stop,
73.1 s, $0.0240956. Record in ratify/.
Verdict RATIFY-WITH-AMENDMENT, amendment A1 applied verbatim to
framework/control-arm-spec.md — and the vote returned this design's own reading of R1–R4
INCORRECT, which is the amendment recorded at §6.8. The design was changed before it dispatched
anything. The vote's own strongest counter-argument is recorded on the spec page: that a
provisional page inheriting an uncalibrated jury should have been REJECTED until re-measured, or
ratified only in part.
14. Amendment record (v1 → v2, 2026-08-01)
Every change is a disposition of a finding in critic.md, applied before any call was
dispatched. Nothing was amended for a reason internal to this page.
| § | amendment | finding |
|---|---|---|
| 3.5 | the claim that the repair lets the sham catch a length response withdrawn; replaced with what it actually buys (the sham is no longer blind to a cue the treatment carries) plus the named non-identification failure mode | 3, BLOCKING |
| 6.3 | "neither easier nor harder to fire" withdrawn as false in both directions; replaced with what matching α does and does not hold constant | 2 |
| 6.7 | both error rates printed; the resolution claim narrowed to excludes a dead instrument, not a weak one; the prescribed second item refused with the reason (the materials do not exist) | 1 |
| 7 | "same-quality" removed from the held-out firing reading, in the critic's own words; the two sham rows given the cue-non-identification clause; a per-juror-fails row added to the light table | 6, BLOCKING · 3 · 4 |
| 8 | prediction 5 given a reachable failure condition that does not treat the held-out arm as an estimator; prediction 8's reading narrowed to provenance | 5 · 7 |
| 12 | the author-wrote-reference-damage-and-design threat added, with the strongest inference it invalidates named | 7 |
Accepted in part, with the refused remedy reasoned: findings 1 (no second known-difference pair
exists) and 5 (a preference-proportion prediction is forbidden by §0.3 and D-20260725-06).
Nothing was rebutted.
v2 → v2.1, the ratification amendment (same session, still before dispatch)
| § | amendment | source |
|---|---|---|
| 3.5, 6.8, 13 | the claim that control-arm-spec R1–R4 do not bind this design withdrawn; §6.8 added with the length-sign split, the degenerate-strata disclosure, the R4 shortfall statement and F3's within-stratum check |
the routed ratification vote, RATIFY-WITH-AMENDMENT + reading INCORRECT |
Two independent voices changed this design before it cost anything: the critic (seven findings) and the ratification vote (one overturned reading). Neither was the lead.