Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260905-tierD-design-v3/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260905-tierD-design-v3
statusfrozen
created2026-09-05
updated2026-09-05
sensesaccuracy, naturalness, perceived-source-carriage, voice, style-correspondence, affect, cultural-mediation, consistency
provisionaltrue
internal-judgment-onlytrue
linkswiki/plan.md, framework/tierD-repaired-rules.md, framework/control-arm-spec.md, wiki/findings/results/RS-20260802-tierD-verdict.md, wiki/goodness-senses.md, config/models.md, wiki/method-notes.md, wiki/decisions/resolved/D-20260905-01-tierD-primary-dose.md, workshop/translations/petka-na-dache/R06-v1/translation.md, workshop/translations/molchanie-andreev/R06-v1/translation.md, workshop/translations/bargamot/R04-v1/translation.md

Frozen design (v3) — Tier D, the one redesign Tom authorized, primary dose = light

wiki/plan.md §W2 step 1. S247, 2026-09-05. Tom authorized exactly one redesign (PROJECT.md §11, 2026-09-04) after two runs (S034, S086) both landed on heavy specificity fails, light specificity fires and the design that produced that pattern had no row for it and no mechanism to make the light dose primary without looking at the numbers first. This page is that redesign. It does not dispatch anything (wiki/plan.md §W2 step 2 does, on a later UTC day, after ratification). What follows is frozen before any S086 score is re-read for the purpose of deciding anything below — the S086 and S034 verdict pages were read once, already, to write wiki/plan.md §W2 and this design's own §0, and that reading is disclosed rather than pretended away; what this page does not do is go back to the S086 cell values to pick a number that would make a chosen dose or threshold pass.

0. What changed, and why this session cannot simply re-run S086

  1. The jury is different, and not by choice. S086 used P1/P2/P5. Method notes (bne) and (bps) (fired S176, S198–199, after S086) establish deepseek/deepseek-v4-pro (P5) as unusable on any task shape and moonshotai/kimi-k3 (P4) as needing a per-seat reasoning-budget probe this session has no budget to spend confirming. The only three live, usable panel seats are P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, and P3 x-ai/grok-4.5. P3 has never sat as a Tier D judging juror before — S020, S034 and S086 all excluded it deliberately, to keep the instrument stable across runs. This design cannot keep that stability and still run at all. The jury swap is disclosed as a limit throughout, not hidden in a methods paragraph.
  2. The materials must be fresh (wiki/plan.md §W2 step 1(d)), so this is not "S086 with a different primary dose declared" — it is a new sham, new targeted items on new anchors, and two fresh lead-provenance references translated this session (§3).
  3. Every sense on wiki/goodness-senses.md is now scored, not the six S086 fixed before perceived-source-carriage existed as a mature sense and before literary-quality retired.

1. Question

Does this jury, at this dose, separate a competently-translated reference from a version of the same text carrying deliberate accuracy damage — and is that separation specific to accuracy rather than a general reaction to any edit? Answered separately at two doses, with one declared primary (§2).

What this unit teaches about translating or evaluating translation, stated once, in front of the apparatus (the subject rule, wiki/tracks.md): if this instrument passes at any dose, the project can, for the first time, put a real translation's accuracy to a jury and have the verdict mean something beyond one session's own reading. If it fails, the honest conclusion is that this panel's own preference judgments are not a way to find out whether a translation is accurate, and the handbook's guidance stays descriptive — sourced from what the project's own translators found doing the work, not from panel scores.

2. The primary dose, and the reason — stated before any S086 or S034 cell value is consulted for it

The primary dose is LIGHT: 3 accuracy-damage sites in a ~300–550-word passage. Heavy (8 sites) runs as a secondary, dose-response cell only.

The reason, and it does not cite what either cell measured last time. The jury's actual job, if Tier D ever licenses one, is telling a competent translation from one with a handful of real errors in it — the everyday discrimination a working evaluator (human or panel) has to make on a translation that is mostly fine. Three errors in a few hundred words is that job. Eight errors in the same span is a different, easier and less interesting one: a passage with an error every 40–50 words does not read like a flawed translation, it reads like a corrupted one, and detecting corruption is not the same capability as detecting inaccuracy in prose that otherwise reads as competent work. A design whose primary cell is the dose closest to the discrimination the jury will actually be asked to make, once any dose is licensed, is the more defensible design independent of which cell happened to fire clean last time — which is the whole point of deciding it before reading the cells again. Heavy stays in the design because a dose-response comparison is free once the light cell is run, and because a future reader may want the ceiling-detection result even though it is not what gates.

What this does and does not settle. It does not retroactively pass S034 or S086 — wiki/ findings/results/RS-20260802-tierD-verdict.md §3 stands exactly as written, a design whose primary was the heavy cell, decided in advance, that did not pass on its own terms. It settles only that this design's gate is the light cell, decided now, before any of this design's own data exists.

3. Materials — fresh on every axis §W2 step 1(d) names

3.1 Published-provenance targeted items — anchors not used at S034 or S086

S034 and S086 both drew targeted material from Turgenev's Memoirs of a Sportsman (Garnett/ Hapgood), which R4 (framework/tierD-repaired-rules.md) reserves for the held-out arm only, where it is the sole qualified pair. This design's targeted published-provenance items come from two anchors never used in any Tier D stage:

Neither anchor's copyright status is in question (both PD; both already stored whole in the project's own repository as anchors, used elsewhere for style-correspondence and cultural-mediation derivation) and using an excerpt of each here adds no new consultation to wiki/base/consulted.md.

3.2 Lead-provenance targeted items — translated this session, contamination measured first

This session's translation limb. Andreev short prose keeps the register this project's Tier D materials already calibrate on (T-bargamot-R04-v1 is the positive control's own base text), so a second Andreev pair was chosen, and the contamination measurement — run immediately after translating and before either item was accepted into this design — is itself part of what this section reports, not a footnote to it.

What the high figure licenses and forecloses, stated once here rather than left implicit. It forecloses L-MOLCH from ever serving as evidence of independent rendering — no design may cite it as a case where the lead worked free of a specific published translation. It does not foreclose its use here: this design needs a stable lead-authored reference text to damage and compare against its own damaged self, and that role does not depend on the lead's rendering being independent of Lowe's. If this reads as a fine distinction, that is because it is the exact distinction CLAUDE.md's contamination rule draws, and this design is the first to have a measurement extreme enough to make the distinction load-bearing rather than academic. A second, independent finding this measurement produces: Andreev's two most anthologized early stories (Vanka-tier fame within his own corpus) are measurably less safe draws for "independent lead translation" claims than Bargamot, an obscurer story from the same author, translator and volume was at S041 — fame within an author's corpus, not merely the author's own fame, predicts contamination.

3.3 The sham — rebuilt, five items, matched to the operator on edit kind (R1, R2)

Five items, 15 units (3 jurors × 5 items). Order swap still applies (§5.2), so each item is 2 orderings × 3 jurors = 6 calls, 30 calls total, matching §5.1/§9 — the band rule (§5.4) pools both orderings into each unit exactly as the targeted and positive-control stages do; nothing about the sham stage exempts it from the order-swap rule stated once, in §5.2, for every item: P-VANKA, P-SPIDER, L-PETKA, L-MOLCH (their undamaged references, each carrying a paraphrase-substitution edit, operator O-SHAM) plus a fifth item, SHAM5 — T-bargamot-R04-v1 Unit A (247 words, lead, suspected, already measured and frozen in an earlier session), distinct from the positive control's Unit B so no text serves two roles.

O-SHAM: three sites per item, each a same-register lexical substitution preserving referent, sense and polarity — a close synonym or a harmless elaboration/trim, never a fact that could be checked true or false. Matched to O4 on edit kind, not on typed content (R2): each item gets one 0-Δword swap, one length-increasing site and one length-decreasing site. 10 of the 15 sites change word count, against 0 of 16 in the unrepaired S040 sham — the defect R2 exists to name is not reproduced here. Full table: materials/perturbations.md; mechanically applied and verified by materials/build.py, which asserts every find string occurs exactly once in its base text before substituting (a typo fails the build, not the run).

3.4 The targeted items — O4, reused verbatim, new sites on new texts

O4 is not re-derived. It is the same four-type accuracy-damage operator S020, S034 and S086 used (wrong referent WR, wrong word sense WS, invented detail ID, dropped negation DN), grounded in the same published catalogue of translation failure (Shaw's Christianising additions, Swann 1974 on Turney, Koteliansky & Murry's 1915 «Пари» mistranslations) — reusing a validated instrument is what "extends rather than replaces" means, and wiki/plan.md §W2 step 1(d) asks for new damage draws from that catalogue, not a new catalogue. Eight heavy sites per item (2 of each type), numbered in textual order; the light dose draws 3 of 8 by the same rule S086 used, reseeded on this design's own id so the draw is independent of S086's: int(sha256("E-20260905-tierD-design-v3|<item id>"),16) mod 56 into the lexicographically sorted 3-subsets of {1..8}.

Drawn light sites (materials/items.json, computed by build.py, not hand-picked): P-VANKA {5,7,8} · P-SPIDER {1,6,7} · L-PETKA {2,4,5} · L-MOLCH {3,5,7}.

Word-count deltas (heavy / light, against the undamaged reference): P-VANKA +6 / +3 · P-SPIDER +10 / +7 · L-PETKA +14 / +10 · L-MOLCH +13 / +7. All eight items are Δ-positive at both doses — the same degenerate stratification S086 disclosed (§6.8 below) recurs here, for the same structural reason: an operator built from WR/WS/ID substitutions on short passages rarely shortens a sentence enough to net negative, and ID sites by construction can only add. Recorded as a disclosure, not tuned away.

3.5 The held-out arm — reused by name, per R4, because no alternative materials exist

framework/tierD-repaired-rules.md R4: the Garnett/Hapgood Memoirs of a Sportsman pair, admissible for accuracy and cultural-mediation only, inadmissible for the other six senses scored here (§4), scoped to the Memoirs cycle. Reused exactly as S034 built it (workshop/experiments/E-20260726d-tierD-heldout/materials/, not the forbidden S086 experiment folder) — R4 exists precisely because no second qualifying pair has been found, and "fresh materials" cannot mean re-opening a decision R4 already settled. H-GARNETT / H-HAPGOOD, unmodified. Order swap still applies (§5.2: mandatory on every item) even though neither member is "the damaged one" — swapping which text is presented first controls position bias regardless of which member is which, so 2 items × 2 orderings × 3 jurors = 12 calls, 6 units. (An earlier draft of this page said "no ordering split" for this arm, which contradicted §5.1/§9's own call count; corrected here rather than silently — caught by this design's own pre-run critic, §13 finding 1.)

3.6 The positive control — reused by name, per R3, for the same reason

T-bargamot-R04-v1 Unit B (441 Russian → 461 English words, lead, suspected, frozen S041) versus variant-F1 (the same text with 8 sites, two of each of O4's four documented failure types, matched to the operator at +4.54% length). R3 names this pair as the only known-difference material the project has; a second such pair still does not exist. 3 units, 6 calls (2 orderings × 3 jurors).

4. Senses scored — every sense on wiki/goodness-senses.md, literary-quality excluded (retired)

accuracy, naturalness, perceived-source-carriage, voice, style-correspondence, affect, cultural-mediation, consistency — eight senses, 1–7 scale, plus a forced overall preference, no ties, strict JSON. Usage-rule compliance, stated once rather than per sense:

5. Arms, jury, rules — carried from tierD-repaired-rules.md and control-arm-spec.md unchanged where they still apply

5.1 Sixteen items, 32 payloads, 96 calls

stage # items items units calls
1 1 PC-B (Bargamot Unit B vs variant-F1) 3 6
2 5 S-VANKA S-SPIDER S-PETKA S-MOLCH S-BARGAMOT5 15 30
3 2 H-GARNETT H-HAPGOOD 6 12
4 4 T8-VANKA T8-SPIDER T8-PETKA T8-MOLCH 12 24
5 4 T3-VANKA T3-SPIDER T3-PETKA T3-MOLCH — primary 12 24

5.2 Jury

P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5, resolved from config/models.md at dispatch time and logged as provenance (§0.1 states why P4/P5 are absent). Order swap is mandatory on every item. Judgment is not parallelised (charter §6): strictly sequential dispatch, one call at a time, in the stage order below.

5.3 The unit and the firing rules — unchanged from tierD-repaired-rules.md, because they depend only on unit counts, not on juror identity

The unit is (juror × item), both orderings combined; +1 if the reference is preferred in both, −1 if the damaged/sham text is, 0 if split. These are pure combinatorial facts about a 3-juror, trinomial-unit design and hold regardless of which three models fill the seats:

cell units rule exact P under the null
positive control 3 ≥ 2 of 3 at +1, none −1 0.109375
sham 15 three-way band (§5.4), branch 0.004193
held-out 6 ≥ 5 of 6 consistent either way, none opposite 0.006348
targeted, light (primary) 12 ≥ 8 of 12 at +1, none −1 0.000594
targeted, heavy (secondary) 12 same 0.000594

Per-juror robustness (unchanged): a targeted cell may not be reported as firing unless it also fires within at least 2 of 3 jurors taken separately (each: all 4 of that juror's items at +1, none −1; null P = 0.25⁴ = 0.003906 per juror).

5.4 The sham band — R1's repair, unchanged

N₊ ≥ 9 of 15 → the jury prefers unedited text as such, no detection claim licensed from this run. N₋ ≥ 9 → the sham materials were not neutral, false-alarm rate unmeasured. Otherwise → acceptable. Each branch exact P = 0.004193.

5.5 Cross-sense specificity — restated as a ratio (note (bke)), and this is the substantive change from S086's rule

Old rule (§6.6 of E-20260801f-tierD-run): condition 3 was drop(naturalness) ≤ 0.75 scale points, an absolute ceiling. Note (bke), fired independently at S129 — before this design existed and not derived from it — records the defect: an absolute ceiling penalises a strong dose and rewards a weak one, because a bigger accuracy manipulation mechanically drags every other sense down by more, in raw points, even when its share of the total movement is unchanged or smaller.

New rule, both conditions now a single 3:1 dominance ratio, reported with its absolute form beside it (note (bke)'s second clause):

  1. Margin, as a ratio: drop(accuracy) ≥ 3 × max(drop of the six other non-target, non-naturalness senses) — perceived-source-carriage, voice, style-correspondence, affect, cultural-mediation, consistency. Equivalently, accuracy accounts for at least 75% of a naive linear split between itself and its largest rival. The absolute margin (S086's original form) is computed and reported alongside, not used to decide.
  2. The largest per-sense drop, excluding naturalness, is accuracy. Unchanged — already relative.
  3. naturalness, as the same ratio: drop(naturalness) / drop(accuracy) ≤ 1/3. The absolute value of drop(naturalness) is reported beside it. This threshold is chosen for a stated reason that is not a number read off S034 or S086: a 3:1 dominance ratio is the natural default for "the on-target sense dominates its largest rival" — ≥ 75% of the pairwise movement between the two, which is a narrower and more defensible claim than "≥ 75% of the total movement across all eight senses" would be (the design's own pre-run critic, §13 finding 1, is right that the two are not the same statement, and only the pairwise one is asserted here) — and it is applied identically to conditions 1 and 3 rather than carrying two different kinds of criterion as the old rule did. Readers who want to know what this would have done to S086's own numbers may compute it from that page's own published cells — this design does not do that computation, on principle, and did not use it to pick 3:1.

If condition 3 fails, the licensed statement is unchanged in form: "specificity is not established for this cell: the target sense's margin cannot be separated from a general downward movement that also reached naturalness."

6. Pre-registered readings — the verdict table, with the missing row named

Primary is the light cell. The top-level licence is decided at light; heavy is reported as a secondary, dose-response cell and never promotes or demotes the primary verdict.

PC light (primary) sham held-out recorded in config/models.md
does not fire — — — NOT PASSED. "The instrument did not separate a known-difference pair at this power; no downstream cell was dispatched."
fires detection + specificity fire, per-juror holds in band does not fire TIER D PASSED on accuracy, light dose, 3 sites in ~300–550-word passages — phrased "the jury detects accuracy-damage at 3 edit sites in this passage range, on a held-out control that did not separate an independent same-quality pair (power 0.655 against an 80% preference)."
fires detection + specificity fire, per-juror fails in band does not fire NOT PASSED. "fired on pooled units but not within jurors taken separately."
fires detection + specificity fire in band fires NOT PASSED. "the rule that licenses detection also separates two independent unmodified translations."
fires detection fires, specificity fails in band does not fire NOT PASSED at the primary dose. "detects damage, not sense-specific on accuracy, at 3 sites."
fires detection fails any any NOT PASSED. "no detection at the primary dose, at this power."
fires any N₊ ≥ 9 or N₋ ≥ 9 any NOT PASSED. No detection claim licensed; the cue is not identified.

The named row — heavy fails, light passes (S034's and S086's own outcome, now expressible):

PC light (primary) heavy (secondary) sham held-out recorded in config/models.md
fires detection + specificity fire, per-juror holds detection fires, specificity fails (either ratio condition) in band does not fire TIER D PASSED on accuracy, light dose only. "Detection generalizes to 8 sites; specificity does not — the heavy cell's damage is no longer separable from a general quality collapse that also depresses naturalness, consistent with S034's and S086's own finding that denser damage reads as prose gone wrong rather than as a local error. The licensed statement names 3 sites; it says nothing about 8." This is a pass, not a caveated pass — the design's whole point is that this outcome no longer needs the heavy cell's cooperation.

Two further rows this design's own structure adds, absent from S086's table:

PC light (primary) heavy (secondary) recorded
fires fires clean also fires clean TIER D PASSED at both doses. Reported as the stronger result it is, with the dose-response direction noted (does specificity margin grow or shrink with dose — descriptive, not gating).
fires fires clean detection also fails at heavy TIER D PASSED, light dose, flagged. A heavy cell that fails to detect more damage than a light cell that succeeded is not internally contradictory (per-juror and per-item noise at n = 12 can do this) but is reported as a named anomaly, not smoothed over.

7. Predictions, written before any run

  1. The positive control fires (both prior runs did, at ceiling).
  2. Detection fires at the light dose. This is the design's least confident prediction under the old framing, where light was secondary and reported as "least confident, informative either way" — under this design it is the primary, and both prior runs' light cells fired at ceiling (12/12, S086), so the confident prediction is that it fires again.
  3. Specificity fires at the light dose under the new ratio rule. internal-judgment-only.
  4. Detection fires at heavy (both priors did, at ceiling).
  5. Specificity at heavy is the prediction this design is genuinely unsure of under the ratio framing — recorded as unresolved rather than guessed, since it depends on the relative sizes of drop(naturalness) and drop(accuracy) at 8 sites on these texts, which nothing before this run measures.
  6. The held-out arm does not fire, ≤ 3 of 6 consistent for either text (S034 and S086 both landed at exactly 3; the same rule, same materials, same jurors-minus-one).
  7. The sham lands in band (S083's and S086's own result on a differently-built sham).
  8. The two lead-provenance items behave like the published items on accuracy (both prior runs' version of this prediction held; charter §5.1's both-provenance requirement).

8. Failure criteria

Registered before dispatch, per §6 of experiment discipline: a call returning malformed JSON after one retry is recorded as WITHHELD, not imputed; a stage whose reserved budget (§9) does not fit the day's remaining headroom is not entered, and the deferral is written into NEXT.md; if more than 2 of 96 calls are WITHHELD, the affected cell's firing rule is reported as inconclusive rather than silently computed on fewer units than registered.

9. Budget — worst case under $3

Prices read from GET /api/v1/models, this session, 2026-09-05: P1 openai/gpt-5.6-terra $2.00 / $12.00 per M (unchanged from S242's reading); P2 google/gemini-3.6-flash $0.75 / $3.75 (unchanged since S182); P3 x-ai/grok-4.5 $2.00 / $6.00 (unchanged since selection). No price has moved since the last reading recorded in config/models.md.

max_tokens is sized from the nearest measured analogue, not from this exact task (wiki/method-notes.md (bsf): "a cap is a per-seat, per-task-shape measurement, not a number carried across a session") — this task shape (8-sense JSON scoring of a ~300–550-word pair) has not been probed on P3, and the ratify-and-run session must probe it (one cheap call per juror, on the longest item, PC-B) before committing to the full run. The figures below are the worst-case ceiling this design freezes; they are not a claim that the probe will confirm them.

juror max_tokens (worst case) reasoning worst-case input (tokens, PC-B) worst case per call
P1 3,000 default 3,500 $0.043
P2 4,000 default 3,500 $0.0176
P3 4,000 effort: low requested (S234: 6× cheaper, 9× faster, 6/8-position agreement with full effort on a comparable structured task) 3,500 $0.031

The worst-case column prices P3 at the full, undiscounted $6.00/M completion rate, not at S234's measured discount — this design's own pre-run critic (§13 finding 2) is right to flag that the prose above cites a discount the arithmetic does not use, and the reason is deliberate rather than an oversight: a worst-case ceiling may not assume a provider honours effort: low (wiki/ method-notes.md (brt) — a cap or parameter verified on one task shape has already failed silently, undiscounted, on another). effort: low is requested because it is expected to lower the central estimate; the ceiling below stays priced at the rate that holds if the request is ignored.

Non-PC-B payloads (sham, held-out, targeted) carry no source text and are shorter; worst-case input taken at 2,000 tokens, same undiscounted rates: P1 $0.040, P2 $0.0165, P3 $0.028.

stage items calls worst case (no retries)
1 positive control 1 6 2 × ($0.043+$0.0176+$0.031) = $0.183
2 sham 5 30 5 × 2 × ($0.040+$0.0165+$0.028) = $0.845
3 held-out 2 12 2 × 2 × ($0.040+$0.0165+$0.028) = $0.338
4 targeted heavy 4 24 4 × 2 × ($0.040+$0.0165+$0.028) = $0.676
5 targeted light (primary) 4 24 4 × 2 × ($0.040+$0.0165+$0.028) = $0.676
total, no retries 16 96 $2.718
+ 2-call retry allowance at the single most expensive call type ($0.043) + $0.086
TOTAL WORST CASE $2.804 — under $3

Central estimate, from S086's own per-call means on the same instrument shape ($0.00525 / $0.01330 / $0.00684 for P1/P2/P5 respectively, P3 unmeasured so taken at P1's rate as a placeholder pending the required pre-flight probe): ≈ $0.72 for 96 calls. The worst case is 3.9× the central estimate, which is the same order of conservatism S086's own budget carried (3.8×) — this design is not quietly looser than its predecessor's discipline, it is tighter in absolute dollars because the item texts are shorter and P2's price has halved since S086.

Abort rule, unchanged in form: stages dispatch 1→5 in order; before entering a stage the runner reserves that stage's full worst case against the day's remaining headroom; a stage that does not fit is not entered, and the deferral is written into NEXT.md. Ordering protects the gate (PC) and the sham first, then the held-out null, then the two targeted doses — light last among the targeted pair only in dispatch order, not in evidential priority; if the day's headroom permits only one targeted cell, dispatch light, not heavy, since light is primary.

10. Verification

tools/verify_tierD.py (or an extension of it, named in the ratify-and-run session) recomputes every reported count from raw stored bodies: unit classification, both branch checks, the per-juror condition, and both specificity ratios (§5.5) from the six raw per-sense scores per unit, independently of any number typed into a result page. Raw request/response JSON preserved under runs/<id>/, one file per call, usage: {"include": true} on every request. Mutation tests: at least one deliberately corrupted stored score must be caught by the verifier before the real run's output is accepted (S086's own practice, 9 of 9 mutations caught).

11. Known threats, stated in advance

  1. The jury swap (§0.1) means this run is not a controlled extension of S020/S034/S086 the way S086 was of S034. Any comparison across runs that treats juror identity as held constant is wrong from this run forward, and a future design should say so rather than silently compare detection rates across the P5→P3 substitution.
  2. Degenerate length-sign stratification recurs (§3.4): all eight targeted items are Δ-positive at both doses. control-arm-spec.md R1's protection — that a pooled figure near chance could be the average of two opposite-signed halves — cannot be defeated by this design's own materials, exactly as S086 disclosed for its own set. Reported, not engineered around.
  3. Two of four targeted reference texts carry measured, disclosed contamination (§3.2). If this run's targeted cell fires, the honest reading is that it fired on damage applied to (in part) text the lead did not independently render — which does not weaken the accuracy-detection claim (§3.2's argument) but does mean this run, like S086, cannot be cited as evidence that the lead translates these specific passages independently of Lowe.
  4. consistency and the source-absent senses may return floor or ceiling data (§4) on single-passage items; a flat result on those senses is predicted, not diagnostic of instrument failure.
  5. The ratio specificity criterion (§5.5) is new and has never been applied to real scores. Nothing guarantees it behaves better than the absolute form in every case; it is adopted for the stated reason (note (bke)) and reported alongside the absolute figures precisely so a future session can tell whether it helped.
  6. P-VANKA and P-SPIDER are read by the panel from anchors already partially quoted in wiki/goodness-senses.md (the register/style-correspondence grounding notes). If any panel seat's training data includes this project's own public wiki pages, both items carry an unmeasured contamination risk of a kind this project has not previously had to consider for published-provenance material. Flagged; not measured, because no tool in this repository measures a panel model's contamination the way dependence_check.py measures the lead's.

12. What ratification must confirm before dispatch (wiki/plan.md §W2 step 2)

  1. The per-seat token probe named in §9 is run and max_tokens is revised if it disagrees.
  2. Today's UTC-day budget headroom is checked against §9's worst case before stage 1 is entered.
  3. materials/build.py is re-run and its items.json SHA-256 manifest matches this page's citations byte-for-byte (the build is deterministic; a mismatch means this page or the materials drifted after freezing and the run does not proceed until reconciled).
  4. The decision page (wiki/decisions/open/D-20260905-01-tierD-primary-dose.md) has been ratified per charter §8 — independent adversarial review plus a routed non-Anthropic panel vote — and any amendment it carries is applied to this design before dispatch.

13. Pre-run critic passes, and dispositions — required before this page could freeze

Two non-Anthropic seats not among this design's own jurors (P1/P2/P3) were sent the frozen draft, the repaired rules, and the perturbation tables, and asked to find statistical errors, internal inconsistencies, and unjustified thresholds (charter §6: "independent pre-run critic — at least one non-Anthropic seat — for any run ... whose result will enter the handbook"; Tier D's verdict binds config/models.md, so it qualifies).

qwen/qwen3.7-max (non-panel reserve), $0.092870, 205.5s, clean. Full text: critique/qwen.json. Two findings, both accepted and applied, in writing:

  1. Accepted. §5.5's prose said a 3:1 ratio means accuracy "explains most of the response," which overclaims a share of the total movement across all eight senses when the design only established a share of the pairwise movement between accuracy and its single largest rival. Wording corrected in §5.5 condition 3 to say exactly that and no more.
  2. Accepted. §3.5 stated the held-out arm needed "no ordering split," which contradicted §5.1 and §9's own call counts (both assumed the order swap). Resolved in favor of keeping the order swap (§5.2 states it is mandatory on every item, with no stated exception, and position bias is a real threat whether or not one member of a pair is "the damaged one") — §3.3 and §3.5 corrected to state the swap explicitly rather than exempt these two stages from a rule stated once and meant to bind everywhere. A third point the same finding raised — that the P3 effort: low discount from S234 is not reflected in the worst-case budget arithmetic — is not a defect: the worst-case ceiling is deliberately priced at the undiscounted rate (§9 now says why, added at this same review pass), since a worst case may not assume a cost-reducing parameter is honoured by every provider (wiki/method-notes.md (brt) already established that a capped parameter can fail silently on a provider that does not respect it).

nvidia/nemotron-3-ultra-550b-a55b (non-panel), three attempts, no usable content. Attempt 1 (critique/nemotron.json, max_tokens 16,000, default effort): 16,000/16,000 completion tokens consumed entirely by (legible, on-topic) reasoning, finish_reason: length, no answer. Attempt 2 (critique/nemotron2.json, max_tokens 30,000, effort: low): finish_reason: error via the Venice routing provider — the effort parameter is not portable to this provider on this model, extending method note (bps) to a new seat. Attempt 3 (max_tokens 32,000, default effort): ran past 11 minutes of wall-clock time with no content and no error, well past every prior latency this project has recorded for any seat on any task shape (S086's 96 calls topped out at 73.1s; this design's own §9 pricing table assumes tens of seconds), and was killed before it billed anything (confirmed against the key-usage snapshot: no change). This design proceeds with one critic pass, not two, and says so rather than treating a single successful review as though two had run. The charter's own bar is "at least one," so the design is not blocked, but the second-seat attempts and their cost ($0.047349 + $0 + $0, all recorded in config/budget.md) are disclosed rather than absorbed silently.