Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: wiki/decisions/resolved/D-20260905-01-tierD-primary-dose.md · rendered 2026-09-09

Page metadata (front matter)
typedecision
idD-20260905-01
statusresolved
created2026-09-05
updated2026-09-06
sensesaccuracy, naturalness, perceived-source-carriage, voice, style-correspondence, affect, cultural-mediation, consistency
internal-judgment-onlytrue
provisionaltrue
linksworkshop/experiments/E-20260905-tierD-design-v3/design.md, framework/tierD-repaired-rules.md, wiki/findings/results/RS-20260802-tierD-verdict.md, wiki/plan.md, PROJECT.md, config/models.md, wiki/goodness-senses.md, wiki/decisions/votes/2026-09-06/D-20260905-01-ratification-record.md

D-20260905-01 — Tier D's one authorized redesign: primary dose = light, specificity as a ratio

RESOLVED 2026-09-06 (S250) — Option C, no amendment

Independent adversarial review (qwen/qwen3.7-max, non-panel): RATIFY, no amendment. Routed non-Anthropic vote (P4 moonshotai/kimi-k3): RATIFY, no amendment, governing. Both voices agree the frozen design dispatches unchanged, subject to its own §12 pre-dispatch conditions (per-seat token probe, budget headroom, build reproducibility). Full record, both raw bodies: wiki/decisions/votes/2026-09-06/D-20260905-01-ratification-record.md. The session that opened this page (S247) took no part in review or vote, per charter §8.

Opened 2026-09-05 (S247), wiki/plan.md §W2 step 1, the single Tier D redesign Tom authorized 2026-09-04 (PROJECT.md §11) after RS-20260802-tierD-verdict found the S086 run's own reading NOT PASSED on one number (drop(naturalness) +1.12 against ≤ 0.75 at the heavy dose), passing at the light dose, with no route in that design to make the passing cell primary after the fact.

Context

Two runs now (S034, S086) have measured the same shape: detection of O4 accuracy-damage fires at ceiling at both a heavy (8-site) and a light (3-site) dose, and a specificity criterion — that the damage does not also depress an unrelated sense, naturalness — passes at light and fails at heavy. Both designs declared heavy primary before running, so both verdicts read NOT PASSED, and the S086 verdict page (§7) states plainly that making the light cell primary after seeing it pass would be exactly the move the project's verification discipline forbids. The fix has to be a fresh design, frozen before the new data exists, that decides the primary dose for a reason that does not cite either run's cell values — and Tom's authorization is for exactly one such redesign, with the approach declared exhausted if it fails.

The question

Which dose is Tier D's primary, and is the cross-sense specificity criterion stated as an absolute scale-point ceiling or as a ratio to the on-target effect (note (bke))?

Options considered

Provisional default

Option C. The frozen design (E-20260905-tierD-design-v3/design.md) is the provisional default this project proceeds on until ratified or overturned. It does not dispatch (wiki/plan.md §W2 step 2 is a separate, later session, on a fresh UTC day, after ratification).

Rationale

  1. The primary-dose reason is ecological, not statistical: a working evaluator's actual job is catching a handful of real errors in mostly-competent prose, which the light dose models and the heavy dose does not (design.md §2). This reason is available before any Tier D data exists and does not change if S086's numbers are re-read.
  2. The ratio form of specificity was already flagged as needed, independently of this redesign, at S129 (note (bke)) — this decision applies a fix the project had already reasoned its way to, rather than inventing one to produce a preferred outcome.
  3. The design was pre-run-critiqued by a non-Anthropic seat (qwen/qwen3.7-max, not a panel juror for this design) before this page was opened; two real defects it found (a held-out-arm call-count self-contradiction; imprecise wording on what the 75% share claim covers) are fixed in the frozen design, and the critique's own text is filed at workshop/experiments/E-20260905-tierD-design-v3/critique/qwen.json. A second critic seat (nvidia/nemotron-3-ultra-550b-a55b) was attempted three times and did not return usable content at up to 32,000 completion tokens on this task shape — recorded as an instrument finding (wiki/method-notes.md, new note this session's hand-off adds) rather than concealed as a missing review.

Contingent artifacts

What ratification would fix

Ratification (independent adversarial review + one routed non-Anthropic panel vote, charter §8) either confirms Option C as the design to dispatch, or returns an amendment the ratify-and-run session must apply before dispatching, or rejects it — in which case, per Tom's authorization, the approach is declared exhausted and config/models.md records Tier D as closed rather than merely NOT CALIBRATED. Never ratified by the session that opened it (this one); the ratify-and-run session (wiki/plan.md §W2 step 2) routes the vote before it dispatches anything.