Repository path: workshop/experiments/E-20260905-tierD-design-v3/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260905-tierD-design-v3 |
| status | frozen |
| created | 2026-09-05 |
| updated | 2026-09-05 |
| senses | accuracy, naturalness, perceived-source-carriage, voice, style-correspondence, affect, cultural-mediation, consistency |
| provisional | true |
| internal-judgment-only | true |
| links | wiki/plan.md, framework/tierD-repaired-rules.md, framework/control-arm-spec.md, wiki/findings/results/RS-20260802-tierD-verdict.md, wiki/goodness-senses.md, config/models.md, wiki/method-notes.md, wiki/decisions/resolved/D-20260905-01-tierD-primary-dose.md, workshop/translations/petka-na-dache/R06-v1/translation.md, workshop/translations/molchanie-andreev/R06-v1/translation.md, workshop/translations/bargamot/R04-v1/translation.md |
Frozen design (v3) — Tier D, the one redesign Tom authorized, primary dose = light
wiki/plan.md §W2 step 1. S247, 2026-09-05. Tom authorized exactly one redesign
(PROJECT.md §11, 2026-09-04) after two runs (S034, S086) both landed on heavy specificity
fails, light specificity fires and the design that produced that pattern had no row for it and no
mechanism to make the light dose primary without looking at the numbers first. This page is that
redesign. It does not dispatch anything (wiki/plan.md §W2 step 2 does, on a later UTC day,
after ratification). What follows is frozen before any S086 score is re-read for the purpose of
deciding anything below — the S086 and S034 verdict pages were read once, already, to write
wiki/plan.md §W2 and this design's own §0, and that reading is disclosed rather than pretended
away; what this page does not do is go back to the S086 cell values to pick a number that would
make a chosen dose or threshold pass.
0. What changed, and why this session cannot simply re-run S086
- The jury is different, and not by choice. S086 used P1/P2/P5. Method notes (bne) and (bps)
(fired S176, S198–199, after S086) establish
deepseek/deepseek-v4-pro(P5) as unusable on any task shape andmoonshotai/kimi-k3(P4) as needing a per-seat reasoning-budget probe this session has no budget to spend confirming. The only three live, usable panel seats are P1openai/gpt-5.6-terra, P2google/gemini-3.6-flash, and P3x-ai/grok-4.5. P3 has never sat as a Tier D judging juror before — S020, S034 and S086 all excluded it deliberately, to keep the instrument stable across runs. This design cannot keep that stability and still run at all. The jury swap is disclosed as a limit throughout, not hidden in a methods paragraph. - The materials must be fresh (
wiki/plan.md§W2 step 1(d)), so this is not "S086 with a different primary dose declared" — it is a new sham, new targeted items on new anchors, and two fresh lead-provenance references translated this session (§3). - Every sense on
wiki/goodness-senses.mdis now scored, not the six S086 fixed beforeperceived-source-carriageexisted as a mature sense and beforeliterary-qualityretired.
1. Question
Does this jury, at this dose, separate a competently-translated reference from a version of the
same text carrying deliberate accuracy damage — and is that separation specific to accuracy
rather than a general reaction to any edit? Answered separately at two doses, with one declared
primary (§2).
What this unit teaches about translating or evaluating translation, stated once, in front of the
apparatus (the subject rule, wiki/tracks.md): if this instrument passes at any dose, the
project can, for the first time, put a real translation's accuracy to a jury and have the verdict
mean something beyond one session's own reading. If it fails, the honest conclusion is that this
panel's own preference judgments are not a way to find out whether a translation is accurate, and
the handbook's guidance stays descriptive — sourced from what the project's own translators found
doing the work, not from panel scores.
2. The primary dose, and the reason — stated before any S086 or S034 cell value is consulted for it
The primary dose is LIGHT: 3 accuracy-damage sites in a ~300–550-word passage. Heavy (8 sites) runs as a secondary, dose-response cell only.
The reason, and it does not cite what either cell measured last time. The jury's actual job, if Tier D ever licenses one, is telling a competent translation from one with a handful of real errors in it — the everyday discrimination a working evaluator (human or panel) has to make on a translation that is mostly fine. Three errors in a few hundred words is that job. Eight errors in the same span is a different, easier and less interesting one: a passage with an error every 40–50 words does not read like a flawed translation, it reads like a corrupted one, and detecting corruption is not the same capability as detecting inaccuracy in prose that otherwise reads as competent work. A design whose primary cell is the dose closest to the discrimination the jury will actually be asked to make, once any dose is licensed, is the more defensible design independent of which cell happened to fire clean last time — which is the whole point of deciding it before reading the cells again. Heavy stays in the design because a dose-response comparison is free once the light cell is run, and because a future reader may want the ceiling-detection result even though it is not what gates.
What this does and does not settle. It does not retroactively pass S034 or S086 — wiki/
findings/results/RS-20260802-tierD-verdict.md §3 stands exactly as written, a design whose primary
was the heavy cell, decided in advance, that did not pass on its own terms. It settles only that
this design's gate is the light cell, decided now, before any of this design's own data exists.
3. Materials — fresh on every axis §W2 step 1(d) names
3.1 Published-provenance targeted items — anchors not used at S034 or S086
S034 and S086 both drew targeted material from Turgenev's Memoirs of a Sportsman (Garnett/
Hapgood), which R4 (framework/tierD-repaired-rules.md) reserves for the held-out arm only,
where it is the sole qualified pair. This design's targeted published-provenance items come from
two anchors never used in any Tier D stage:
P-VANKA— Chekhov, «Ванька», tr. Constance Garnett (1922),wiki/base/anchors/ A-garnett-vanka/vanka-garnett-1922.txt, the opening 422 words (four paragraphs, PD).P-SPIDER— Akutagawa, 蜘蛛の糸, tr. Glenn W. Shaw (1930),wiki/base/anchors/ A-shaw-spider-thread/spider-thread-shaw-1930.txt, ten paragraphs from "One day the Buddha..." (314 words, PD, already independently verified PD in the anchor's own header).
Neither anchor's copyright status is in question (both PD; both already stored whole in the
project's own repository as anchors, used elsewhere for style-correspondence and
cultural-mediation derivation) and using an excerpt of each here adds no new consultation to
wiki/base/consulted.md.
3.2 Lead-provenance targeted items — translated this session, contamination measured first
This session's translation limb. Andreev short prose keeps the register this project's Tier D
materials already calibrate on (T-bargamot-R04-v1 is the positive control's own base text), so a
second Andreev pair was chosen, and the contamination measurement — run immediately after
translating and before either item was accepted into this design — is itself part of what this
section reports, not a footnote to it.
L-PETKA— Andreev, «Петька на даче» (1899) opening, translated in this session,T-petka-na-dache-R06-v1. Contamination:suspected— longest run against W. H. Lowe's 1915 published translation of the same story sits at 12 tokens, exactly the frozen dependence threshold, against a 4-token null-control floor (Lowe's other Andreev stories, 38,317 words). Full measurement in the translation page.L-MOLCH— Andreev, «Молчание» (1900) §I opening, translated in this session,T-molchanie-andreev-R06-v1. Contamination:high— a 19-token verbatim-order run against Lowe's rendering of the same sentence, against a 6-token null-control floor. This is the strongest dependence this project has measured against a published rendering on material the lead was not shown before translating (CLAUDE.md§Contamination's note (bhb) records a higher figure only against the lead's own prior output).
What the high figure licenses and forecloses, stated once here rather than left implicit. It
forecloses L-MOLCH from ever serving as evidence of independent rendering — no design may cite it
as a case where the lead worked free of a specific published translation. It does not foreclose
its use here: this design needs a stable lead-authored reference text to damage and compare against
its own damaged self, and that role does not depend on the lead's rendering being independent of
Lowe's. If this reads as a fine distinction, that is because it is the exact distinction
CLAUDE.md's contamination rule draws, and this design is the first to have a measurement extreme
enough to make the distinction load-bearing rather than academic. A second, independent finding
this measurement produces: Andreev's two most anthologized early stories (Vanka-tier fame within
his own corpus) are measurably less safe draws for "independent lead translation" claims than
Bargamot, an obscurer story from the same author, translator and volume was at S041 — fame within
an author's corpus, not merely the author's own fame, predicts contamination.
3.3 The sham — rebuilt, five items, matched to the operator on edit kind (R1, R2)
Five items, 15 units (3 jurors × 5 items). Order swap still applies (§5.2), so each item is 2
orderings × 3 jurors = 6 calls, 30 calls total, matching §5.1/§9 — the band rule (§5.4) pools both
orderings into each unit exactly as the targeted and positive-control stages do; nothing about the
sham stage exempts it from the order-swap rule stated once, in §5.2, for every item: P-VANKA,
P-SPIDER, L-PETKA, L-MOLCH (their
undamaged references, each carrying a paraphrase-substitution edit, operator O-SHAM) plus a
fifth item, SHAM5 — T-bargamot-R04-v1 Unit A (247 words, lead, suspected, already
measured and frozen in an earlier session), distinct from the positive control's Unit B so no text
serves two roles.
O-SHAM: three sites per item, each a same-register lexical substitution preserving referent,
sense and polarity — a close synonym or a harmless elaboration/trim, never a fact that could be
checked true or false. Matched to O4 on edit kind, not on typed content (R2): each item gets one
0-Δword swap, one length-increasing site and one length-decreasing site. 10 of the 15
sites change word count, against 0 of 16 in the unrepaired S040 sham — the defect R2 exists to
name is not reproduced here. Full table: materials/perturbations.md; mechanically applied and
verified by materials/build.py, which asserts every find string occurs exactly once in its base
text before substituting (a typo fails the build, not the run).
3.4 The targeted items — O4, reused verbatim, new sites on new texts
O4 is not re-derived. It is the same four-type accuracy-damage operator S020, S034 and S086
used (wrong referent WR, wrong word sense WS, invented detail ID, dropped negation DN), grounded in
the same published catalogue of translation failure (Shaw's Christianising additions, Swann 1974 on
Turney, Koteliansky & Murry's 1915 «Пари» mistranslations) — reusing a validated instrument is what
"extends rather than replaces" means, and wiki/plan.md §W2 step 1(d) asks for new damage draws
from that catalogue, not a new catalogue. Eight heavy sites per item (2 of each type), numbered in
textual order; the light dose draws 3 of 8 by the same rule S086 used, reseeded on this design's own
id so the draw is independent of S086's: int(sha256("E-20260905-tierD-design-v3|<item id>"),16) mod
56 into the lexicographically sorted 3-subsets of {1..8}.
Drawn light sites (materials/items.json, computed by build.py, not hand-picked):
P-VANKA {5,7,8} · P-SPIDER {1,6,7} · L-PETKA {2,4,5} · L-MOLCH {3,5,7}.
Word-count deltas (heavy / light, against the undamaged reference): P-VANKA +6 / +3 ·
P-SPIDER +10 / +7 · L-PETKA +14 / +10 · L-MOLCH +13 / +7. All eight items are Δ-positive at
both doses — the same degenerate stratification S086 disclosed (§6.8 below) recurs here, for
the same structural reason: an operator built from WR/WS/ID substitutions on short passages rarely
shortens a sentence enough to net negative, and ID sites by construction can only add. Recorded
as a disclosure, not tuned away.
3.5 The held-out arm — reused by name, per R4, because no alternative materials exist
framework/tierD-repaired-rules.md R4: the Garnett/Hapgood Memoirs of a Sportsman pair,
admissible for accuracy and cultural-mediation only, inadmissible for the other six senses
scored here (§4), scoped to the Memoirs cycle. Reused exactly as S034 built it
(workshop/experiments/E-20260726d-tierD-heldout/materials/, not the forbidden S086 experiment
folder) — R4 exists precisely because no second qualifying pair has been found, and "fresh
materials" cannot mean re-opening a decision R4 already settled. H-GARNETT / H-HAPGOOD,
unmodified. Order swap still applies (§5.2: mandatory on every item) even though neither member
is "the damaged one" — swapping which text is presented first controls position bias regardless of
which member is which, so 2 items × 2 orderings × 3 jurors = 12 calls, 6 units. (An earlier
draft of this page said "no ordering split" for this arm, which contradicted §5.1/§9's own call
count; corrected here rather than silently — caught by this design's own pre-run critic, §13
finding 1.)
3.6 The positive control — reused by name, per R3, for the same reason
T-bargamot-R04-v1 Unit B (441 Russian → 461 English words, lead, suspected, frozen S041) versus
variant-F1 (the same text with 8 sites, two of each of O4's four documented failure types,
matched to the operator at +4.54% length). R3 names this pair as the only known-difference
material the project has; a second such pair still does not exist. 3 units, 6 calls (2 orderings
× 3 jurors).
4. Senses scored — every sense on wiki/goodness-senses.md, literary-quality excluded (retired)
accuracy, naturalness, perceived-source-carriage, voice, style-correspondence, affect,
cultural-mediation, consistency — eight senses, 1–7 scale, plus a forced overall preference, no
ties, strict JSON. Usage-rule compliance, stated once rather than per sense:
naturalnessis put to the jury with its register frame named — unmarked/literary- contemporary (wiki/goodness-senses.md), since every reference text here is competent, non-archaising prose in that register — per S020/S086 practice and the sense's own requirement that an evaluation name one of its three register anchors.perceived-source-carriage— jurors see no source text in any stage but the positive control (§3.6) and the held-out arm needs none. Per the sense's own requirement, this is declared rather than left implicit: a score on this sense in the sham/targeted stages is a statement about reception with no source shown, and usage rule 7 already establishes that a high score here is not evidence anything was carried — this design adds nothing to that caution, it inherits it.affect— usage rule 4 requires stating which half is scored. Only the reader-experience half (what does this do to you as a reader of English) is asked; the source-comparability half requires a yardstick document this design does not build (charter §Framework parameters; the affect entry's own §D-20260813-17 history is why this design does not attempt one). A singleaffectfigure here answers one question, named, not both.consistency— the entry's own reachability note says this is a whole-text, cross-site sense and a single 250–550-word passage supplies at most one or two forked classes to test it against, if any. This design predictsconsistencywill sit near ceiling or be uninformative on these items, and says so before dispatch rather than after a flat result — a floor effect here is not evidence the operator failed to change anything.literary-qualityis not scored (retiredD-20260801-11; excluded per usage rule 1).
5. Arms, jury, rules — carried from tierD-repaired-rules.md and control-arm-spec.md unchanged where they still apply
5.1 Sixteen items, 32 payloads, 96 calls
| stage | # items | items | units | calls |
|---|---|---|---|---|
| 1 | 1 | PC-B (Bargamot Unit B vs variant-F1) |
3 | 6 |
| 2 | 5 | S-VANKA S-SPIDER S-PETKA S-MOLCH S-BARGAMOT5 |
15 | 30 |
| 3 | 2 | H-GARNETT H-HAPGOOD |
6 | 12 |
| 4 | 4 | T8-VANKA T8-SPIDER T8-PETKA T8-MOLCH |
12 | 24 |
| 5 | 4 | T3-VANKA T3-SPIDER T3-PETKA T3-MOLCH — primary |
12 | 24 |
5.2 Jury
P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5, resolved from
config/models.md at dispatch time and logged as provenance (§0.1 states why P4/P5 are absent).
Order swap is mandatory on every item. Judgment is not parallelised (charter §6): strictly
sequential dispatch, one call at a time, in the stage order below.
5.3 The unit and the firing rules — unchanged from tierD-repaired-rules.md, because they depend only on unit counts, not on juror identity
The unit is (juror × item), both orderings combined; +1 if the reference is preferred in both, −1 if the damaged/sham text is, 0 if split. These are pure combinatorial facts about a 3-juror, trinomial-unit design and hold regardless of which three models fill the seats:
| cell | units | rule | exact P under the null |
|---|---|---|---|
| positive control | 3 | ≥ 2 of 3 at +1, none −1 | 0.109375 |
| sham | 15 | three-way band (§5.4), branch | 0.004193 |
| held-out | 6 | ≥ 5 of 6 consistent either way, none opposite | 0.006348 |
| targeted, light (primary) | 12 | ≥ 8 of 12 at +1, none −1 | 0.000594 |
| targeted, heavy (secondary) | 12 | same | 0.000594 |
Per-juror robustness (unchanged): a targeted cell may not be reported as firing unless it also fires within at least 2 of 3 jurors taken separately (each: all 4 of that juror's items at +1, none −1; null P = 0.25⁴ = 0.003906 per juror).
5.4 The sham band — R1's repair, unchanged
N₊ ≥ 9 of 15 → the jury prefers unedited text as such, no detection claim licensed from this run. N₋ ≥ 9 → the sham materials were not neutral, false-alarm rate unmeasured. Otherwise → acceptable. Each branch exact P = 0.004193.
5.5 Cross-sense specificity — restated as a ratio (note (bke)), and this is the substantive change from S086's rule
Old rule (§6.6 of E-20260801f-tierD-run): condition 3 was drop(naturalness) ≤ 0.75 scale
points, an absolute ceiling. Note (bke), fired independently at S129 — before this design
existed and not derived from it — records the defect: an absolute ceiling penalises a strong dose
and rewards a weak one, because a bigger accuracy manipulation mechanically drags every other
sense down by more, in raw points, even when its share of the total movement is unchanged or
smaller.
New rule, both conditions now a single 3:1 dominance ratio, reported with its absolute form beside it (note (bke)'s second clause):
- Margin, as a ratio:
drop(accuracy) ≥ 3 × max(drop of the six other non-target, non-naturalness senses)—perceived-source-carriage,voice,style-correspondence,affect,cultural-mediation,consistency. Equivalently,accuracyaccounts for at least 75% of a naive linear split between itself and its largest rival. The absolute margin (S086's original form) is computed and reported alongside, not used to decide. - The largest per-sense drop, excluding
naturalness, isaccuracy. Unchanged — already relative. naturalness, as the same ratio:drop(naturalness) / drop(accuracy) ≤ 1/3. The absolute value ofdrop(naturalness)is reported beside it. This threshold is chosen for a stated reason that is not a number read off S034 or S086: a 3:1 dominance ratio is the natural default for "the on-target sense dominates its largest rival" — ≥ 75% of the pairwise movement between the two, which is a narrower and more defensible claim than "≥ 75% of the total movement across all eight senses" would be (the design's own pre-run critic, §13 finding 1, is right that the two are not the same statement, and only the pairwise one is asserted here) — and it is applied identically to conditions 1 and 3 rather than carrying two different kinds of criterion as the old rule did. Readers who want to know what this would have done to S086's own numbers may compute it from that page's own published cells — this design does not do that computation, on principle, and did not use it to pick 3:1.
If condition 3 fails, the licensed statement is unchanged in form: "specificity is not established for this cell: the target sense's margin cannot be separated from a general downward movement that also reached naturalness."
6. Pre-registered readings — the verdict table, with the missing row named
Primary is the light cell. The top-level licence is decided at light; heavy is reported as a secondary, dose-response cell and never promotes or demotes the primary verdict.
| PC | light (primary) | sham | held-out | recorded in config/models.md |
|---|---|---|---|---|
| does not fire | — | — | — | NOT PASSED. "The instrument did not separate a known-difference pair at this power; no downstream cell was dispatched." |
| fires | detection + specificity fire, per-juror holds | in band | does not fire | TIER D PASSED on accuracy, light dose, 3 sites in ~300–550-word passages — phrased "the jury detects accuracy-damage at 3 edit sites in this passage range, on a held-out control that did not separate an independent same-quality pair (power 0.655 against an 80% preference)." |
| fires | detection + specificity fire, per-juror fails | in band | does not fire | NOT PASSED. "fired on pooled units but not within jurors taken separately." |
| fires | detection + specificity fire | in band | fires | NOT PASSED. "the rule that licenses detection also separates two independent unmodified translations." |
| fires | detection fires, specificity fails | in band | does not fire | NOT PASSED at the primary dose. "detects damage, not sense-specific on accuracy, at 3 sites." |
| fires | detection fails | any | any | NOT PASSED. "no detection at the primary dose, at this power." |
| fires | any | N₊ ≥ 9 or N₋ ≥ 9 | any | NOT PASSED. No detection claim licensed; the cue is not identified. |
The named row — heavy fails, light passes (S034's and S086's own outcome, now expressible):
| PC | light (primary) | heavy (secondary) | sham | held-out | recorded in config/models.md |
|---|---|---|---|---|---|
| fires | detection + specificity fire, per-juror holds | detection fires, specificity fails (either ratio condition) | in band | does not fire | TIER D PASSED on accuracy, light dose only. "Detection generalizes to 8 sites; specificity does not — the heavy cell's damage is no longer separable from a general quality collapse that also depresses naturalness, consistent with S034's and S086's own finding that denser damage reads as prose gone wrong rather than as a local error. The licensed statement names 3 sites; it says nothing about 8." This is a pass, not a caveated pass — the design's whole point is that this outcome no longer needs the heavy cell's cooperation. |
Two further rows this design's own structure adds, absent from S086's table:
| PC | light (primary) | heavy (secondary) | recorded |
|---|---|---|---|
| fires | fires clean | also fires clean | TIER D PASSED at both doses. Reported as the stronger result it is, with the dose-response direction noted (does specificity margin grow or shrink with dose — descriptive, not gating). |
| fires | fires clean | detection also fails at heavy | TIER D PASSED, light dose, flagged. A heavy cell that fails to detect more damage than a light cell that succeeded is not internally contradictory (per-juror and per-item noise at n = 12 can do this) but is reported as a named anomaly, not smoothed over. |
7. Predictions, written before any run
- The positive control fires (both prior runs did, at ceiling).
- Detection fires at the light dose. This is the design's least confident prediction under the old framing, where light was secondary and reported as "least confident, informative either way" — under this design it is the primary, and both prior runs' light cells fired at ceiling (12/12, S086), so the confident prediction is that it fires again.
- Specificity fires at the light dose under the new ratio rule.
internal-judgment-only. - Detection fires at heavy (both priors did, at ceiling).
- Specificity at heavy is the prediction this design is genuinely unsure of under the ratio
framing — recorded as unresolved rather than guessed, since it depends on the relative sizes
of
drop(naturalness)anddrop(accuracy)at 8 sites on these texts, which nothing before this run measures. - The held-out arm does not fire, ≤ 3 of 6 consistent for either text (S034 and S086 both landed at exactly 3; the same rule, same materials, same jurors-minus-one).
- The sham lands in band (S083's and S086's own result on a differently-built sham).
- The two lead-provenance items behave like the published items on
accuracy(both prior runs' version of this prediction held; charter §5.1's both-provenance requirement).
8. Failure criteria
Registered before dispatch, per §6 of experiment discipline: a call returning malformed JSON after
one retry is recorded as WITHHELD, not imputed; a stage whose reserved budget (§9) does not fit
the day's remaining headroom is not entered, and the deferral is written into NEXT.md; if more
than 2 of 96 calls are WITHHELD, the affected cell's firing rule is reported as inconclusive
rather than silently computed on fewer units than registered.
9. Budget — worst case under $3
Prices read from GET /api/v1/models, this session, 2026-09-05: P1 openai/gpt-5.6-terra
$2.00 / $12.00 per M (unchanged from S242's reading); P2 google/gemini-3.6-flash $0.75 / $3.75
(unchanged since S182); P3 x-ai/grok-4.5 $2.00 / $6.00 (unchanged since selection). No price has
moved since the last reading recorded in config/models.md.
max_tokens is sized from the nearest measured analogue, not from this exact task
(wiki/method-notes.md (bsf): "a cap is a per-seat, per-task-shape measurement, not a number
carried across a session") — this task shape (8-sense JSON scoring of a ~300–550-word pair) has
not been probed on P3, and the ratify-and-run session must probe it (one cheap call per juror, on
the longest item, PC-B) before committing to the full run. The figures below are the
worst-case ceiling this design freezes; they are not a claim that the probe will confirm them.
| juror | max_tokens (worst case) |
reasoning | worst-case input (tokens, PC-B) |
worst case per call |
|---|---|---|---|---|
| P1 | 3,000 | default | 3,500 | $0.043 |
| P2 | 4,000 | default | 3,500 | $0.0176 |
| P3 | 4,000 | effort: low requested (S234: 6× cheaper, 9× faster, 6/8-position agreement with full effort on a comparable structured task) |
3,500 | $0.031 |
The worst-case column prices P3 at the full, undiscounted $6.00/M completion rate, not at S234's
measured discount — this design's own pre-run critic (§13 finding 2) is right to flag that the
prose above cites a discount the arithmetic does not use, and the reason is deliberate rather than
an oversight: a worst-case ceiling may not assume a provider honours effort: low (wiki/
method-notes.md (brt) — a cap or parameter verified on one task shape has already failed silently,
undiscounted, on another). effort: low is requested because it is expected to lower the central
estimate; the ceiling below stays priced at the rate that holds if the request is ignored.
Non-PC-B payloads (sham, held-out, targeted) carry no source text and are shorter; worst-case
input taken at 2,000 tokens, same undiscounted rates: P1 $0.040, P2 $0.0165, P3 $0.028.
| stage | items | calls | worst case (no retries) |
|---|---|---|---|
| 1 positive control | 1 | 6 | 2 × ($0.043+$0.0176+$0.031) = $0.183 |
| 2 sham | 5 | 30 | 5 × 2 × ($0.040+$0.0165+$0.028) = $0.845 |
| 3 held-out | 2 | 12 | 2 × 2 × ($0.040+$0.0165+$0.028) = $0.338 |
| 4 targeted heavy | 4 | 24 | 4 × 2 × ($0.040+$0.0165+$0.028) = $0.676 |
| 5 targeted light (primary) | 4 | 24 | 4 × 2 × ($0.040+$0.0165+$0.028) = $0.676 |
| total, no retries | 16 | 96 | $2.718 |
| + 2-call retry allowance at the single most expensive call type ($0.043) | + $0.086 | ||
| TOTAL WORST CASE | $2.804 — under $3 |
Central estimate, from S086's own per-call means on the same instrument shape ($0.00525 / $0.01330 / $0.00684 for P1/P2/P5 respectively, P3 unmeasured so taken at P1's rate as a placeholder pending the required pre-flight probe): ≈ $0.72 for 96 calls. The worst case is 3.9× the central estimate, which is the same order of conservatism S086's own budget carried (3.8×) — this design is not quietly looser than its predecessor's discipline, it is tighter in absolute dollars because the item texts are shorter and P2's price has halved since S086.
Abort rule, unchanged in form: stages dispatch 1→5 in order; before entering a stage the
runner reserves that stage's full worst case against the day's remaining headroom; a stage that
does not fit is not entered, and the deferral is written into NEXT.md. Ordering protects the
gate (PC) and the sham first, then the held-out null, then the two targeted doses — light last
among the targeted pair only in dispatch order, not in evidential priority; if the day's
headroom permits only one targeted cell, dispatch light, not heavy, since light is primary.
10. Verification
tools/verify_tierD.py (or an extension of it, named in the ratify-and-run session) recomputes
every reported count from raw stored bodies: unit classification, both branch checks, the
per-juror condition, and both specificity ratios (§5.5) from the six raw per-sense scores per unit,
independently of any number typed into a result page. Raw request/response JSON preserved under
runs/<id>/, one file per call, usage: {"include": true} on every request. Mutation tests: at
least one deliberately corrupted stored score must be caught by the verifier before the real run's
output is accepted (S086's own practice, 9 of 9 mutations caught).
11. Known threats, stated in advance
- The jury swap (§0.1) means this run is not a controlled extension of S020/S034/S086 the way S086 was of S034. Any comparison across runs that treats juror identity as held constant is wrong from this run forward, and a future design should say so rather than silently compare detection rates across the P5→P3 substitution.
- Degenerate length-sign stratification recurs (§3.4): all eight targeted items are Δ-positive
at both doses.
control-arm-spec.mdR1's protection — that a pooled figure near chance could be the average of two opposite-signed halves — cannot be defeated by this design's own materials, exactly as S086 disclosed for its own set. Reported, not engineered around. - Two of four targeted reference texts carry measured, disclosed contamination (§3.2). If this run's targeted cell fires, the honest reading is that it fired on damage applied to (in part) text the lead did not independently render — which does not weaken the accuracy-detection claim (§3.2's argument) but does mean this run, like S086, cannot be cited as evidence that the lead translates these specific passages independently of Lowe.
consistencyand the source-absent senses may return floor or ceiling data (§4) on single-passage items; a flat result on those senses is predicted, not diagnostic of instrument failure.- The ratio specificity criterion (§5.5) is new and has never been applied to real scores. Nothing guarantees it behaves better than the absolute form in every case; it is adopted for the stated reason (note (bke)) and reported alongside the absolute figures precisely so a future session can tell whether it helped.
P-VANKAandP-SPIDERare read by the panel from anchors already partially quoted inwiki/goodness-senses.md(the register/style-correspondence grounding notes). If any panel seat's training data includes this project's own public wiki pages, both items carry an unmeasured contamination risk of a kind this project has not previously had to consider for published-provenance material. Flagged; not measured, because no tool in this repository measures a panel model's contamination the waydependence_check.pymeasures the lead's.
12. What ratification must confirm before dispatch (wiki/plan.md §W2 step 2)
- The per-seat token probe named in §9 is run and
max_tokensis revised if it disagrees. - Today's UTC-day budget headroom is checked against §9's worst case before stage 1 is entered.
materials/build.pyis re-run and itsitems.jsonSHA-256 manifest matches this page's citations byte-for-byte (the build is deterministic; a mismatch means this page or the materials drifted after freezing and the run does not proceed until reconciled).- The decision page (
wiki/decisions/open/D-20260905-01-tierD-primary-dose.md) has been ratified per charter §8 — independent adversarial review plus a routed non-Anthropic panel vote — and any amendment it carries is applied to this design before dispatch.
13. Pre-run critic passes, and dispositions — required before this page could freeze
Two non-Anthropic seats not among this design's own jurors (P1/P2/P3) were sent the frozen draft,
the repaired rules, and the perturbation tables, and asked to find statistical errors, internal
inconsistencies, and unjustified thresholds (charter §6: "independent pre-run critic — at least one
non-Anthropic seat — for any run ... whose result will enter the handbook"; Tier D's verdict binds
config/models.md, so it qualifies).
qwen/qwen3.7-max (non-panel reserve), $0.092870, 205.5s, clean. Full text:
critique/qwen.json. Two findings, both accepted and applied, in writing:
- Accepted. §5.5's prose said a 3:1 ratio means accuracy "explains most of the response," which overclaims a share of the total movement across all eight senses when the design only established a share of the pairwise movement between accuracy and its single largest rival. Wording corrected in §5.5 condition 3 to say exactly that and no more.
- Accepted. §3.5 stated the held-out arm needed "no ordering split," which contradicted §5.1
and §9's own call counts (both assumed the order swap). Resolved in favor of keeping the order
swap (§5.2 states it is mandatory on every item, with no stated exception, and position bias is
a real threat whether or not one member of a pair is "the damaged one") — §3.3 and §3.5 corrected
to state the swap explicitly rather than exempt these two stages from a rule stated once and
meant to bind everywhere.
A third point the same finding raised — that the P3
effort: lowdiscount from S234 is not reflected in the worst-case budget arithmetic — is not a defect: the worst-case ceiling is deliberately priced at the undiscounted rate (§9 now says why, added at this same review pass), since a worst case may not assume a cost-reducing parameter is honoured by every provider (wiki/method-notes.md(brt) already established that a capped parameter can fail silently on a provider that does not respect it).
nvidia/nemotron-3-ultra-550b-a55b (non-panel), three attempts, no usable content. Attempt 1
(critique/nemotron.json, max_tokens 16,000, default effort): 16,000/16,000 completion tokens
consumed entirely by (legible, on-topic) reasoning, finish_reason: length, no answer. Attempt 2
(critique/nemotron2.json, max_tokens 30,000, effort: low): finish_reason: error via the
Venice routing provider — the effort parameter is not portable to this provider on this model,
extending method note (bps) to a new seat. Attempt 3 (max_tokens 32,000, default effort): ran
past 11 minutes of wall-clock time with no content and no error, well past every prior latency this
project has recorded for any seat on any task shape (S086's 96 calls topped out at 73.1s; this
design's own §9 pricing table assumes tens of seconds), and was killed before it billed anything
(confirmed against the key-usage snapshot: no change). This design proceeds with one critic pass,
not two, and says so rather than treating a single successful review as though two had run. The
charter's own bar is "at least one," so the design is not blocked, but the second-seat attempts and
their cost ($0.047349 + $0 + $0, all recorded in config/budget.md) are disclosed rather than
absorbed silently.