Repository path: config/models.md · rendered 2026-09-09
Page metadata (front matter)
| type | note |
|---|---|
| id | models |
| status | active |
| created | 2026-07-23 |
| updated | 2026-09-07 |
| links | wiki/decisions/resolved/D-20260723-03-panel-v1-composition.md, config/budget.md, framework/tierD-repaired-rules.md, wiki/findings/results/RS-20260907-panel-judging-2.md |
The panel
The models the project reaches via OpenRouter (charter §5). Model slugs are configured here and nowhere else — specs and tools refer to roles; run records log resolved slugs as provenance. Revisit when the landscape changes (triggers below). This composition was ratified 2026-07-23 (decision D-20260723-03, resolved: independent adversarial review + non-Anthropic panel vote, both RATIFY) as a defensible, revisable working panel — not an optimum; changes go through the revisit triggers.
Calibration state, FINAL, set 2026-09-06 (S250): TIER D CALIBRATION IS EXHAUSTED. NOT CALIBRATED,
permanently, not pending recalibration. No jury verdict carries evidential weight; every workshop
self-assessment and every panel score anywhere in this repository is provisional and
internal-judgment-only, permanently. The handbook (framework/v0.3/) continues to rest on
workshop findings and traceability notes, never on a panel-jury verdict.
History of how this was reached (restructured 2026-07-25, charter A3; updated 2026-07-26, S034;
instrument defects recorded 2026-07-27, S040; stage-1 gate declared unrepairable 2026-07-28, S050;
repair arm closed resolved 2026-07-29, S055; the repaired instrument run to completion 2026-08-02,
S086; 2026-09-04 (S244), Tom authorized exactly one further redesign (PROJECT.md §11) after
S086's verdict, with the stated consequence that a failure declares the approach exhausted; that
one redesign was ratified and run 2026-09-06 (S250), and it FAILED, on two independent dispositive
gates — see the S250 entry below, which is now the final word on this row.**
- Tier D — THE ONE AUTHORIZED REDESIGN, RATIFIED AND RUN 2026-09-06 (S250). NOT PASSED, on two
independent dispositive gates, and this is EXHAUSTING: no further redesign is authorized.
RS-20260906-tierD-verdict-v3/E-20260905-tierD-design-v3/D-20260905-01. 96 calls in one session, $0.732950, 96 of 96 accepted, zero failures, zero retries, zero truncations — verifier 13/16 checks read as expected (the 3 "failures" are the reported non-firing gates themselves), 1 of 1 mutation tests caught. Ratified by independent adversarial review (qwen/qwen3.7-max, non-panel) and a routed non-Anthropic vote (moonshotai/kimi-k3,P4), both RATIFY with no amendment (wiki/decisions/votes/2026-09-06/D-20260905-01-ratification-record.md). What fired. The positive control (3 of 3, none −1). Detection at the heavy (secondary) dose, at ceiling (12 of 12 units, per-juror 3/3). The held-out arm correctly did not fire (2 of 6 consistent). What failed, and this is a new failure shape, not a repeat of S086's. (i) The sham broke into the UPPER branch for the first time in four Tier D runs (9 of 15 units preferred the unedited reference even against pure paraphrase substitutions carrying no content change) — on the design's own pre-registered table this alone reads NOT PASSED, cue not identified. (ii) Detection at the PRIMARY (light) dose did not fire — 9 of 12 units, one unit at −1 (the gate requires none), per-juror robustness 0 of 3. (iii) Cross-sense specificity fired at neither dose, and not because the ratio rule is defective:consistency— never scored in a Tier D run before this one — moved MORE thanaccuracyat both doses (+4.583 vs +3.500 heavy; +2.292 vs +2.125 light). The mechanism, checkable site by site, not merely asserted: this design (unlike S020/S034/S086) showed the jury NO source text except in the positive control. Three ofO4's four damage types (wrong referent, wrong sense, invented detail) are only legible as errors against a source the jury never saw here; the fourth (dropped negation) can contradict the passage's own surrounding sense and is exactly whatconsistencyis defined to catch. Every light-dose item whose independently-drawn 3-of-8 subset includes a dropped-negation site fired detection perfectly (9 of 9 units); the one item whose draw has none is the sole failure, and on it the jury actively preferred the damaged text. Consequence for any future instrument of this shape: withholding the source from the jury does not merely limit whatperceived-source-carriagecan claim (which is what it was designed to protect) — it appears to convertaccuracydetection intoconsistencydetection, catching only the subset of accuracy damage that happens to create an internal contradiction. This was not among the design's own six disclosed threats; it is a limit the run itself surfaced, not an execution defect. Full account, all five stages tabulated, the both-provenance check and the held-out per-sense pattern:RS-20260906-tierD-verdict-v3. Per Tom's 2026-09-04 authorization, a failure of this one redesign declares the approach exhausted. It is exhausted, as of this entry.ARM-first-judgmentand every other jury-gated arm stayprovisional-only, permanently. - Tier D — THE COMPLETE RUN UNDER THE REPAIRED RULES, EXECUTED 2026-08-02 (S086). NOT PASSED, on exactly one number, and it is the strongest run this project has produced.
RS-20260802-tierD-verdict/E-20260801f-tierD-run/ARM-tierD-run, closedresolved(a FAIL is completion). 96 calls across two sessions, $0.812454145, 48 of 48 this session accepted with zero failures and zero retries; verifier 1,664 checks, 0 failures, nine mutation tests, nine caught. Jurors P1/P2/P5, as at S020 and S034, so this extends the instrument rather than replacing it. - What fired. The prior positive control (2 of 3, R3's replacement for the unrepairable scale gate). The rebuilt sham IN BAND for the first time ever — 6/7/2 on 15 units, matched to the operator on edit kind and sized for power (0.8689 against a genuine 70% bias, against the 0.4202 of the rule it replaced). Detection at ceiling in both targeted cells: 12 of 12 units at +1, no −1 anywhere, all three jurors 4 of 4 taken separately (per-juror null P = 0.003906), and 8 of 8 with P5 excluded. The held-out arm did not fire (3 to Garnett, 0 to Hapgood, 3 split).
- What failed:
drop(naturalness)= +1.12 at the 8-site dose against §6.6 condition 3's ≤ 0.75. Cross-sense specificity therefore does not fire at the heavy dose, and §7's pre-registered reading gives NOT PASSED — "detects damage, not sense-calibrated onaccuracy." Conditions 1 and 2 pass comfortably in both cells (margins +2.16 heavy, +2.08 light). - At three sites everything fires — and it is NOT a pass, for a reason fixed before the run.
drop(naturalness)is +0.50 and all three specificity conditions hold. §7's first table is keyed on the heavy cell and its second table only names the dose a licensed statement would use ("the light dose is reported but does not gate"). Making the cell that passed into the primary after watching it pass is the move the verification discipline exists to stop, and it is not made here. The §5.5 form: on these 1,152 stored scores,drop(naturalness)at threshold 0.75 would have decided NOT PASSED at 8 sites and PASS at 3 sites had 3 sites been the primary. It was not. - The dose result reproduces S034 to two decimal places on the number that decides it — naturalness 1.11 → 1.12 at 8 sites, 0.56 → 0.50 at 3 — and this time the light cell is four independent draws fixed by
sha256(design id | item id) mod 56, not S034's single positional subset. Eight accuracy errors in a 300-word passage stop reading as accuracy errors and start reading as bad English; three do not. That is a finding about damage, not about the panel. - Charter §5.1's both-provenances requirement returns a null in the reassuring direction for the second run running. Lead-authored references behaved like the published ones at both doses: 6 of 6 units, accuracy-drop gaps 0.17 (heavy) and 0.25 (light) against a pre-registered ≤ 1.50. Per design §12 this is a fact about provenance, not evidence about accuracy damage as a category.
- The held-out per-sense pattern reproduced across runs: Hapgood higher on
accuracyon both items (−0.50, −1.00), Garnett higher onnaturalness(+0.67, +1.17) andliterary-quality(+1.17, +1.17) — per dimension, the ratified 1904/1906 reception record. Not a Tier P pass: two items, six units, canonicity uncontrolled by construction. - Predictions: five held, three failed — specificity at 8 sites (3), no detection at 3 sites (4, which fired at ceiling on four independent draws), and naturalness under 0.75 in both cells (7). The failures are on record because the design wrote them down first.
literary-qualitywas retired the day before this run and the verdict does not depend on it: recomputed on five senses, the heavy margin goes +2.16 → +2.17 and still fails c3, the light +2.08 → +2.17 and still fires.- One instrument defect, pinned in the verifier rather than described:
H-HA__o0__P5routed to SiliconFlow and billed 12,237 completion tokens against amax_tokensof 10,000 withfinish_reason: stop— $0.0414 against a declared per-call worst case of $0.038775. A worst case built frommax_tokensis not a guarantee (note (bgk), fired a second time on a different model and provider). No budget event: the stage reservation landed at 18% of $2.795. -
ARM-first-judgmentis unblocked asprovisional-only (charter §2.4);ARM-framework-v01stays blocked, becauseframework/closure.md§3 conditions a release on the gate. -
Tier D — AN OBLIGATION ADDED 2026-08-02 (S093), and it does not change the calibration state.
D-20260802-13was ratified (verdict C, review and vote agreeing) andnaturalness's definition changed underneath the instrument: the source-conditional escape clause is struck and the sense is now scored on the target text alone. Tier D's one failing number wasdrop(naturalness)= +1.12 against a ≤ 0.75 bar. Ratifying condition 4 therefore binds every future Tier D design: the specificity test must be re-run under the revised wording before any revised-sense jury verdict is called calibrated or compared with prior Tier D results. Nothing is currently being cited on Tier D's strength — it is NOT PASSED — so this bites the next design, not the next session. Also carried as the onlyowedrow inwiki/backlog.md. DISCHARGED 2026-08-05 (S113),RS-20260805e-wording-or-prose, and it changes no verdict. The test was re-run under the revised wording, on the S086 payloads byte-identical and on two new Portuguese items:drop(naturalness)comes back at 1.083 under the old string (against 1.11 at S034 and 1.12 at S086 — three runs, spread 0.037) and at exactly 0.750 under the revised one, and the 0.333 gap between them is not distinguishable from unit-level noise (exact paired permutation P = 0.138 on 12 units). Tier D stays NOT PASSED and this row stays NOT CALIBRATED. Two things a future Tier D design must carry: (i) say whichnaturalnessstring a condition-3 number was measured under — both strings are in the repository and they give 1.083 and 0.750 on the same twelve units; (ii) the revised string's smaller drop is the reference coming down, not the damaged text being spared (reference 6.125 → 5.417, damaged copy 5.042 → 4.667), and on new materials the same 0.333 shrinkage came entirely from the damaged copy rising. On those new materialsdrop(naturalness)is 1.917–2.250, so the two prior runs measured this statistic at its small end. A second consequence a future design must not miss: the new senseperceived-source-carriagehas never been through Tier D at all, and its own entry records a 0.40–0.60 false-positive rate on manufactured oddity; it is not a calibrated instrument and must not be presented as one. - Tier D — the stage-1 scale-usage gate, DECLARED UNREPAIRABLE 2026-07-28 (S050),
RS-20260728g-scale-usage/E-20260728g-scale-usage/ARM-tierD-repairstep 2. Nothing here changes the calibration state. $0.0197964, one independent pre-run critic call; the analysis is recomputation from 1,440 stored scores across both Tier D runs (S020 and S034, same three jurors), verifier 332 checks, 0 failures, including exact reproduction of S040's own §4 figures. S040 repaired §10's gate in form — per juror, not pooled — and left the number open. There is no number. All six (run × juror) cells clear the 0.75 criterion margin on their targeted arm (realised margins 1.500 to 4.083), so with the lowest observed sham dispersion at 0.143, any threshold above it fails a demonstrably capable juror and any threshold at or below it passes everybody; and the ordering by sham dispersion is not monotone in the realised margin in either run. One cell in six passes the inherited 0.75, and it is the cell whose statistic is most inflated by sense-level offset (0.866 → 0.617 with sense means removed; under the centred statistic none of the six passes). The gate would have blocked both runs, in each of which every juror separately detected the O4 damage at ceiling — an operating characteristic, not a proof that the statistic is wrong, which is the pre-run critic's correction to a claim this session's own design had made and withdrew before running. What a future Tier D design must therefore not do: import §10's gate with a tuned number. Its replacement is undecided and is step 5 of the arm — either a prior positive control on a known-difference pair (materials now exist:T-bargamot-R04-v1Unit B +variant-F1, eight sites, two of each of O4's four documented failure types, matched to the operator at +4.54% length) or a posterior threshold on realised differences, which is cheaper and is not prior.RS-20260727b§4's inference about P5 is superseded: "P5 cannot express a drop larger than one scale point at all" described the sham stage; on the targeted arm of the same run P5 produced a 3.500-point accuracy margin. - Tier D — the repair arm CLOSED
resolved2026-07-29 (S055), and the repaired rules now live in ONE place:framework/tierD-repaired-rules.md. Import that block; do not re-derive it from three result pages. Nothing here changes the calibration state — TIER D REMAINS NOT PASSED.ARM-tierD-repairdischarged all three conditions ofwiki/backlog.md'sowedentry: (i) repaired, (ii) unrepairable with the consequence stated, (iii) met, with one clause flagged. What the repair does: replaces the sham band's decision rule and sizes it for power (≈15 units, not 6); requires the sham to match the operator on edit kind; removes §10's scale-usage gate from the design vocabulary entirely, because no threshold on within-juror dispersion discriminates on this panel; and settles its replacement as a prior positive control on a known-difference pair (T-bargamot-R04-v1Unit B +variant-F1) rather than a posterior threshold, on the ground that being prior was the gate's whole function. What the repair does not do: it runs nothing, it calibrates nothing, and no jury verdict gains evidential weight. Two things a future design must carry rather than rediscover. (a) The stage-1 replacement is a specification choice, not a ratified decision, andframework/control-arm-spec.md's ratification trigger — "the next design that builds a paired comparison" — fired atE-20260729cand was not routed; that is declared rather than left to be noticed. (b) The evidence base licensing the only qualified pair's two admissible senses contains no printed exhibit from the work it licenses — see the condition (iii) entry below. - Tier D — condition (iii), DISCHARGED 2026-07-29 (S055):
D-20260725-07's outcome reproduces on evidence the lead did not write, and one of its clauses does not.RS-20260729c-neutral-summary/E-20260729c, $0.260345590, key-usage delta exact to 1e-9, verifier 50 checks, 0 failures, one pre-run critic call (moonshotai/kimi-k3, NEEDS-AMENDMENT, five findings, all five accepted, the design's own sentence "a one-cell question and one cell answers it" withdrawn before any call was dispatched). The primary was re-derived from the item id and every load-bearing S025 correction REPRODUCED. The ratification protocol was then re-run on two independently authored non-lead summaries of the whole page, no call told that a prior ratification existed. The routed vote returned C on both; the sense narrowing was identical three times out of three; the review/vote split reproduced. ⚠ The work restriction reproduced on one of two — on the other the vote licensed A House of Gentlefolk, the work clause 3 excludes — so clause 3 is flagged and a design that needs it should re-open it as a decision. The registered prediction that whole-page evidence would surface the citation-scope problem FAILED, and the argument came instead from the lead-written arm: the section both 2026-07-25 voices convicted of advocacy is the section that supplied the strongest attack on the outcome the advocate wanted. The-cfinding a future Tier D design must not inherit silently: all eighteen paired extracts the 1904 review prints are cited to Hapgood vol. iv / Garnett vol. ii — A Nobleman's Nest — and the review prints no style exhibit from the Memoirs of a Sportsman at all, so clause 2's exclusion of the four prose senses rests, in printed exhibits, entirely on the work clause 3 refuses to license. - Tier D — the instrument, REPAIRED IN SPECIFICATION 2026-07-27 (S040),
RS-20260727b-tierD-rules/ARM-tierD-repair. Nothing here changes the calibration state, and two of the three findings make a pass harder. No API call; everything recomputed from S034's stored outputs and frozen manifest. (i) §6.5's sham band reproduced by independent enumeration (0.004639 / 0.533936, ratio 115×) and repaired: state the lower branch on N₋ ≥ 5 of 6, matched at 0.004639 by construction, and stop collapsing the two branches into one conclusion — an upper firing means detection is confounded with edit-presence, a lower firing means the sham materials were not neutral and the false-alarm rate is unmeasured. (ii) The band's power had never been computed: 5-of-6 on 6 units catches a genuine 70% edit-presence bias 0.42 of the time, so a working sham needs ≈15 units (5 items), not 6. (iii) New, and not previously named anywhere: the sham and the targeted operator are not the same kind of edit — 0 of 16 sham sites change the word count against 10 of 24 targeted sites (+44 words, ≈4% per item), and 6 of 16 sham sites change no lexeme at all against 0 of 24 targeted. A floor measured on edits that cannot change length does not bound the false-alarm rate of edits that do. (iv) §10's pooled scale gate passed on variance that is not scale usage: 56.8% of the pooled sum of squares is between jurors, and pooled SD 0.783 falls to 0.515 with juror means removed, below the threshold it passed. All three jurors fail the per-juror form of the gate, so at S020's own numbers the S034 run would not have entered stage 2. The S034 verdict is not reopened and no repaired rule is applied to it retrospectively. - Tier D — the complete run, EXECUTED 2026-07-26 (S034),
RS-20260726d-tierD-heldout/E-20260726d-tierD-heldout. NOT PASSED, on two independent grounds. The first run in this project to carry all three mandatory controls on the same materials with the same jurors — sham, held-out and cross-sense specificity — which is what the S020 run could not do and whatARM-tierDwas constituted to build. 60 calls, 60/60 parsed, no truncation, $0.725719 (key-usage delta exact); every number recomputed byanalysis/verify.py, which imports nothing fromtools/and recomputes the null probabilities by exhaustive enumeration — 71 checks, 0 failures. Materials: Turgenev «Свидание» (Garnett 1897 / Hapgood 1904) split at a landmark named in the frozen design, plusT-son-makara-R04-v1(Korolenko, lead, contamination-measured at a longest run of 8 tokens). Jurors P1/P2/P5. - Ground 1 — the sham is out of band. 0 of 6 units preferred the unedited text in both orderings, so §6.5's pre-registered lower branch fires and no detection claim is licensed from the run. The verdict stands as pre-registered. But the rule is defective and the arithmetic is now on record: under the design's own null the lower branch has probability 0.534 against the upper branch's 0.0046 — a 116-fold asymmetry, and a branch that fires more often than not on a jury doing nothing. S020's sham measured 2 of 6, one unit above firing it. §6.5 is carried verbatim from S020 and must be replaced before any future run relies on it — a directional rule on the count of −1 units, or a two-sided rule with matched branch probabilities. This is a design task; it does not retroactively rescue this run.
- Ground 2 — cross-sense specificity fails at the heavy dose. Detection fired at ceiling — 9/9 units at both doses, in all 3 jurors taken separately, and still 6/6 with P5 excluded — but at 8 sites
drop(naturalness)= 1.11, above §6.6's 0.75 condition. At 3 sites it is 0.56 and specificity fires. The accuracy margin is +2.04 at both doses, identical to two decimals; what separates them is that eight accuracy edits accumulate into prose that reads wrong as English and stop being sense-localised. The smaller dose is the cleaner result. The light dose is one pre-registered nested subset {2,5,7}, not a sample, so no general dose claim follows. - The held-out arm did not fire (3 units to Garnett, 0 to Hapgood, 3 split; rule needs 5 of 6). This licenses exactly one sentence — "this 5-of-6 rule did not detect separation" — and nothing about chance, parity or equivalence: its power against an 80% preference is 0.56, and every non-split unit went the same way. Its per-sense pattern is the notable finding: Hapgood scores higher on
accuracyon both items (−0.17, −0.67) while Garnett scores higher onnaturalness(+1.17, +1.50) andliterary-quality(+1.17, +1.00) — per dimension, the ratified 1904/1906 reception record (D-20260725-07). This is not a Tier P pass and must not be read as one; it is two items and six units on one pair, whose forced overall preference still went 3–0 to the canonical translator. - Charter §5.1's both-provenances requirement returned a null, in the reassuring direction. The lead-authored reference behaved like the published ones:
T8-K3/3 units, accuracy drop 4.17 against a published mean of 4.00 (gap 0.17, pre-registered condition ≤1.50). Prediction 7 held. - Instrument notes. (i) The stage-1 scale-usage gate passed on the pool (SD 0.783, integers {4,5,6,7}) by 0.033, and P5 used two adjacent integers across 48 scores (SD 0.143; P1 0.672, P2 0.568) — a pooled gate carried by one juror, the same defect shape as §6.5's band, and it should be stated per juror next time. (ii) The build's six-part typographic audit found a second OCR artefact the design's repair list had missed —
lowgrowingin Hapgood HB — reported and deliberately not repaired, and now a named confound onH-HB. (iii) British-vs-American orthography, predicted by the critic as a live discriminator, is a measured null: 0 cross-variant tokens in all ten items. (iv) Punctuation density is not: Garnett 22 commas to Hapgood's 36 on the same passage. (v) P5 routed across seven providers inside 20 calls (Alibaba, Baidu, BaseTen, GMICloud, Ionstream, Novita, StreamLake); no call on any juror billed above its declared per-call worst case. -
When Tier D does pass on a sense, record here what it licenses — detects S-damage at dose D — and nothing broader. Nothing above licenses that sentence.
-
Tier D — detection calibration (the gate): FIRST RUN EXECUTED 2026-07-25 (S020),
RS-20260725-tierD-ladder. NOT PASSED, and it could not have been. The charter's held-out arm (reference vs an independent same-quality translation) is mandatory, and the independent critic established by measurement that the project's stored materials cannot supply one — the only candidate pair is Beowulf, where Gummere is alliterative verse and Kirtlan and the lead are prose, with 5 / 17 / 0 archaic tokens. The design therefore dropped the arm and pre-committed to claiming no pass. No sense is calibrated; no jury verdict carries evidential weight; all workshop self-assessments remainprovisional. What the run established: (i) the false-alarm floor is small but non-zero — eight quality-neutral substitutions moved no sense by more than 0.25 scale points, though the jury leaned 5/6 votes to the unedited text on a period passage and returned identical scores on a contemporary one; (ii)accuracycleared both legs — detection 6/6 units (exact P = 0.0002), specificity +3.10 against a 0.75 criterion, accuracy falling 6.83 → 1.75, replicating on both passages and with P5 excluded — and would be reportable as detection-calibrated if the held-out control existed; (iii)literary-qualityis detected but not localised (6/6 detection, per-sense drops flat at +1.08/+1.08/+1.08 against a +1.17 target); (iv)style-correspondence's apparently passing specificity was withheld by a pre-registered rule — the operator makes prose more conventionally English, so naturalness should not fall, and it fell +1.50; (v)voiceandcultural-mediationfailed detection, and the failure inverted by passage — on Shaw 1930 the jury unanimously preferred the de-marked text and scoredcultural-mediation4.00 → 6.67. Instrument notes: slot preference measured at P1 0.500, P2 0.600, P5 0.550, materially better than S010's 0.63–0.85 on this item format; on quality-neutral material two of three jurors used only the top three integers of a 1–7 scale, with the full range appearing only against real damage. When Tier D does pass on a sense, record here what it licenses — detects S-damage at dose D — and nothing broader. - Tier D — the held-out arm, 2026-07-25 (S021): still blocked, and the blocker has been re-diagnosed.
E-20260725-heldout-pair/RS-20260725-heldout-pair. The materialsNEXT.mdasked for were found — Chekhov's «Пари» in two PD English translations five years apart (Koteliansky & Murry 1915, Garnett 1920), same story, same form, both storable whole — and a six-measure qualification instrument was frozen and put through an independent non-Anthropic critic pass, which returned NEEDS-REDESIGN with ten blockers and was accepted. Three findings, none of which unblocks anything: (i) the 1915 translation is damaged — «за пять часов» → "five minutes", «Евангелие» → "the New Testament" while keeping "by no means thick", «в 12 часов дня» (noon) → "twelve o'clock midnight" in contradiction of its own contract clause; verified twice by independent code paths; (ii) the instrument would have certified a 111-year mismatch — run on the 1915 text against the lead's own 2026 translation of the same story, all six measures pass, because archaism and period-orthography counts detect deliberate archaising (the Beowulf renderings separate at 12.85–42.66 per 1,000) and are blind to a century of ordinary prose (all three Chekhov renderings sit at 0.00); (iii) canonicity is not a text property — Garnett is the default English Chekhov and K&M is not, so an LM jury may separate the pair on cadence familiarity, and no measurement can control for it. S021 concluded that "roughly same quality" is not a property this project can establish about any pair of texts it can reach, and escalated the operationalisation question toD-20260725-06(five options, no default acted on). - Tier D — the held-out arm, RATIFIED 2026-07-25 (S022):
D-20260725-06resolves to Q-A, the literal reading. The arm is a condition on the materials, and it is not satisfiable by construction, by surface measurement, or by relaxing the null. Independent adversarial review (openai/gpt-5.6-terra, P1)RATIFY-Q-A; non-Anthropic review vote (google/gemini-3.6-flash, P2)RATIFY. Both rejected the lead's stated preference. Full record:wiki/decisions/votes/2026-07-25/D-20260725-06-ratification-record.md. What is now binding: - "Same-quality" must be established from independent, pre-existing, external grounds before the arm runs. Reading it as a null hypothesis about the jury makes the control circular — the jury's behaviour would decide whether the materials were fit to test the jury's behaviour. Parity may not be created by choosing generators, by surface-measuring prose, or by setting a tolerance after seeing a damage effect.
- Two conditions bind every candidate pair: (i) comparative reception evidence — identifying both translations, addressing the same work or a materially relevant portion, supporting comparability rather than praising each translator separately; a lead-authored synthesis of diffuse remarks does not qualify; (ii) a mandatory pre-run factual-damage audit against the source — known material damage disqualifies a pair even if reception evidence exists.
- Rejected, with reasons on the decision page: Q-B (a constructed model/model arm is neither undamaged nor canonicity-free by construction, and fails population validity); the {lead translation, panel translation} variant (intensifies the provenance confound and makes the lead's own output a pole of the control licensing evaluation of that output — admissible only as a labelled exploratory stress test); Q-C (changes the null, and no defensible fraction exists — the damage gap ~5.08 and the sham maximum 0.25 estimate different phenomena, neither being ordinary between-translator variation); Q-D; Q-E.
- The correction that matters most: Q-A does not mean "blocked indefinitely." Both voices found the project's impossibility claim stronger than its evidence — one pair, one story, and one surface instrument that could not see a wrong number is not a survey, and no search protocol, search log, or list of reception sources considered-and-excluded was ever produced. The honest state is blocked until a properly documented search is run. The route named by the vote, now the top of the Tier D path: expand discovery to other PD short prose with multiple independent human translations; shift the candidate unit to passage level (100–500 words); run the factual audit as a filter first.
- Tier D — condition (ii), FIRST PAIR PARTIALLY DISCHARGED 2026-07-25 (S025):
D-20260725-07resolves to option C, narrowed. Independent adversarial review (openai/gpt-5.6-terra, P1) returned A; the routed non-Anthropic vote (google/gemini-3.6-flash, P2) returned C and governs. Full record:wiki/decisions/votes/2026-07-25/D-20260725-07-ratification-record.md. What is binding: - The Garnett 1894–99 / Hapgood 1903–05 pair satisfies condition (ii) for the Memoirs of a Sportsman / A Sportsman's Sketches cycle ONLY, and for the senses
accuracyandcultural-mediationONLY. With (i) and (iii) already discharged, that is the first candidate pair in this project to clear all three conditions on any sense. - Inadmissible for
naturalness,literary-quality,style-correspondence,voice— two independent 1904/1906 reviews agree on a directional Garnett advantage on English prose, and a held-out control on a pair with a documented gap on the measured sense imports the bias the control exists to exclude. - Inadmissible for A Nobleman's Nest / A House of Gentlefolk — the work the Nation's 40:12 style count is actually about.
- The evidence moved because the scans were read properly, not because more were found. The Nation page is three-column; archive.org's linearised OCR shuffles it. Reading order recovered from word geometry (
tools/deinterleave_djvu.py) and each load-bearing sentence then read off the page image. Two S024 quotation defects surfaced: a reconstruction wrong in three words, and the 40:12 count quoted without its scope or its hedge about the audited cycle. - Both voices found the submitting session's framing of its own corrections to be advocacy — sound corrections, overstated presentation — and named three overstatements, all since fixed in place. The un-taken check: a future session should re-derive the primary itself and re-run the vote on a summary the lead did not write.
- This calibrates nothing. It unblocks the design of a held-out arm on two senses. The arm has not been designed, let alone run.
- Nothing above changes the calibration state.
- Tier P — peer discrimination (a certification, not a gate): run three times, never certified — and 2026-08-04 (S108) the reason the second run gave for it being untestable was shown to be a property of the GRAIN, not of the materials. Case A below is run 1;
RS-20260804c-peer-record(S103) is run 2;RS-20260804i-name-or-prose(S108) is run 3,ARM-tierPstep 2. Nothing here changes the calibration state — Tier D remains NOT PASSED and no jury verdict carries evidential weight. - The blind that failed holds at a finer unit. Run 2 put whole loci of «Певцы» to P1/P2/P3 and all three named Garnett and all three named Hapgood, and concluded that this record cannot be tested blind on this instrument. Run 3 put the same story, same pair, same seats one to three sentences at a time under a false attribution and asked them to name the translators independently of it: P1 followed the false label at 24 of 24 cells and the true text at 0 of 24. So
D-20260725-07's ratified pair does not need replacing, and no search for a non-canonical pair is owed. Method note (biq). - What a Tier P statement must now carry: per-sense label-robustness. Like-for-like over the two seats present in both conditions, 24 cells each —
accuracy(record: Hapgood) 18–6 unattributed → 23–1 misattributed;naturalness(record: Garnett) 20–4 → 14–10, a coin. Neither direction reverses, so the run's registeredPR3holds; but a certification that reported one figure per sense without its label-robustness would be reporting a number a wrong name can halve. - The grain also changes which half of the record reproduces. Same seats, same loci: at locus level (run 2)
accuracywas 19–17 and not reproduced whilenaturalnesswas 30–6; at site level (run 3)accuracyis 26–10 andnaturalness28–8. The Nation's flat claim, which failed at locus level, comes back at site level. - What run 3 does NOT establish. Its two primaries were withheld by their own failure criterion: the two seats shown only the Russian agreed at 6 of 12 on whether a span forces a trade-off, below the registered bar of 4 FORK sites, and the bar was not moved. The recognition measurement that interprets everything is one seat (P3's answer was malformed), and the misattributed stage ran on two seats because P2 returned no body in four attempts. No certification is claimed and Tier P is not certified by this run.
- 22 bodies, $0.926131649, key cross-check exact to 2e-10; verifier 146 checks, 0 failures, 3 mutation tests, 3 caught. Pre-run critic
NEEDS-REDESIGN, six findings, three BLOCKING, all six accepted before any grading call. 65% of the spend bought nothing — see notes (bhf) sixth/seventh firing and (bgw) second firing.
Lead agent as translator (2026-07-25, charter §3/§5, A4). The lead translates as a labeled subject at no API cost. It never judges its own output; the non-Anthropic panel judges blind with authorship stripped. Panel membership stays non-Anthropic — the lead-correlation argument governs judges, and it binds harder now that the lead also produces subjects. This is the carve-out anticipated at D-03 ratification ("does not foreclose using an Anthropic model as a labeled translation subject whose output is judged by the non-Anthropic panel"), now exercised. Each lead translation declares contamination (whether published translations of that work plausibly sit in the lead's training data) and freezes its translator's log before evaluation is designed.
First calibration run executed 2026-07-25 (S014): Case A (Botchan, opening, upper-bounded) — wiki/findings/results/RS-20260725-calibration-caseA. This was a Tier P run. Outcome: the panel does not reproduce the documented Cohn-favoring reception record on the real test (Turney vs Cohn) — record-fit at/below chance on voice (0.55), affect (0.42), literary-quality (0.43); only naturalness reaches 0.80 and is register-cued. The sanity floor (Morri vs Cohn) passes near-ceiling. No sense is calibrated; no evidential weight on any sense. No US-taste cluster fired (Metric 2 firing rule not met). Instrument note: juror deepseek-v4-pro (P5) showed a 0.44 order-flip rate — down-weight until reps increase. Recalibration/next steps in the result page and NEXT.md.
Instrument note added 2026-09-01 (S237), from RS-20260901-inversion-habit. The first run to
probe reasoning: {"effort": "low"} on P3 against a key rather than against its own
full-effort answers. On a 523-item word-order coding task it cost $0.0073 a 24-item batch against
$0.094 at full effort — a factor of thirteen — and failed both calibration gates, calling an
INVERTED keyed line CANONICAL 7 times of 19 where P1 did so 3 times and P2 0. The errors
are one-sided: it under-detects the construct rather than scattering. So the S234 note's "worth
probing on P3" is narrowed: worth probing, and only against keyed items, reading the direction of
the errors — self-agreement at six of eight positions, which is what S234 had, does not detect a
seat that is systematically not looking. Note (bsq). Two further per-seat facts from the same
run, both on a 24-item batched coding prompt: P1 needs cap 12000 (it spends ~5,000 completion
tokens and returns clean; 4000 truncates before any answer), and P2 cannot do 24 items at any
cap probed — at 12000 it spent 11,996 tokens and answered 18 — so it was run at batch 12,
where it returns clean at ~5,700 tokens, and became the dearest seat of the three.
Instrument note added 2026-08-30 (S234), from RS-20260830b-rhyme-family. The first probe this project has run of all four reachable seats on the same two task shapes, at two caps and two reasoning settings, and it changes what a pre-flight may assume. On a nine-line structured answer at max_tokens 2500: P1 openai/gpt-5.6-terra clean at 9.0 s / $0.007056; P3 x-ai/grok-4.5 clean at 52.2 s / $0.016238, and with reasoning: {"effort": "low"} at 5.6 s / $0.002698 — six times cheaper, nine times faster, and its answer agreed with its own full-effort answer at six of eight positions; P2 google/gemini-3.6-flash truncated, and returned clean only at cap 6000 (16.3 s / $0.010747); the first reserve qwen/qwen3.7-max truncated, clean at 6000 but at 66.7 s / $0.026243; and the reserve z-ai/glm-5.2 returned an empty body after 6000 completion tokens, which extends note (brt) from short prompts to this shape. On a search-shaped prompt — find the one English rhyme best covering a list of senses — every seat collapsed: P1 empty at 2500 for $0.030528, unchanged by low effort, P3 no return in 100 s twice, P2 clean only at 6000 for $0.019511. Two conclusions for pre-flights. (i) A cap is a per-seat, per-task-shape measurement, not a number carried across a session — note (bsf), fourth firing. (ii) reasoning: {"effort": "low"} is worth probing on P3 and is worth nothing on P1; it is a per-seat property, not a lever. No panel composition changed and no reserve was promoted: stage F ran on P1 alone and stage S on P2 alone, disjoint by design, and the probe is why.
Instrument note added 2026-09-05 (S247), from E-20260905-tierD-design-v3's pre-run critic
dispatch. A third task shape — critique a ~50 KB frozen experimental design and return a
structured numbered-findings response — was probed on the two non-panel reserves plus
nvidia/nemotron-3-ultra-550b-a55b (a non-panel seat used before only for one-off adversarial
review, e.g. S108). qwen/qwen3.7-max returned a clean, on-topic critique at max_tokens 16,000
($0.0929, 205.5s). z-ai/glm-5.2 reproduced (brt) on this shape too, now at a much higher cap than
S234's: 66,674 characters of on-topic reasoning and zero returned content at max_tokens
20,000 ($0.108). nvidia/nemotron-3-ultra-550b-a55b failed three ways in three attempts: default
effort exhausted 16,000/16,000 tokens on reasoning alone with no answer; effort: low returned
finish_reason: error from the Venice routing provider (extending (bps) — the parameter is not
portable to this provider on this model); default effort at max_tokens 32,000 is recorded in
workshop/experiments/E-20260905-tierD-design-v3/critique/. This project now has three
independent seat/task-shape pairs where a large max_tokens cap alone does not fix a reasoning
seat that has not converged (rhyme-search at S234, this critique shape at S247, on two different
seats) — the fix that works is a different seat, not a bigger cap, once a cap in the tens of
thousands has already failed once.
Instrument note added 2026-07-25 (S015), from RS-20260725-anchor-verification. On a factual-adjudication task — not a quality judgment — all four non-Anthropic panel members used (P1, P2, P3, P5) performed strongly: each rejected 20/20 planted false claims, including 6/6 refutable only against the specific stored files, at false-alarm rates of 0.083–0.114, giving discrimination scores of 0.886–0.917; abstention was ~0 and held-out true controls were supported 3/3 by every model. All four also passed a six-item competence screen in Russian, French and Japanese (≥5/6 each; the panel had previously been probed on Japanese only). This does not bear on calibration — calibration is about matching a human reception record on matters of quality, and the panel remains NOT CALIBRATED — but it does mean the panel is usable as a checking instrument for textual and linguistic fact, in the failing direction (charter §4 still forbids treating their agreement as validation). Two cautions from the same run: P5 alone missed the jade/kingfisher polysemy of 翡翠 on the screen; and support for evaluative claims (0.958) ran higher than for descriptive ones (0.890), i.e. these models agree most readily where there is least to check.
Panel v1 (selected 2026-07-23; probe: config/probes/2026-07-23/)
| # | slug | lab (country) | list price in/out per M | probe notes (internal-judgment-only) |
|---|---|---|---|---|
| P1 | openai/gpt-5.6-terra |
OpenAI (US) | $2.00 / $12.00 (read from the API 2026-09-03, S242 — the first UPWARD move recorded here; this row read $2.50 / $15.00 from selection until 2026-07-30, $1.25 / $7.50 until 2026-08-04, and $1.00 / $6.00 until 2026-09-03) | 3/3 accurate; fast (3.1s), terse token use → cheapest frontier call in practice ($0.0028) |
| P2 | google/gemini-3.6-flash |
Google (US) | $0.75 / $3.75 (read from the API 2026-08-14, S182; this row read $1.50 / $7.50 from selection until then) | 3/3 accurate; heavy hidden reasoning (~1.7k tok, $0.013); clean idiomatic register |
| P3 | x-ai/grok-4.5 |
xAI (US) | $2.00 / $6.00 | 3/3 accurate; clean on the Kajii litotes ("This was rather bad" — though P5 produced the identical rendering, so not uniquely best); $0.0037 |
| P4 | moonshotai/kimi-k3 |
Moonshot (CN) | $3.00 / $15.00 | 3/3 accurate with literary flair ("a twenty-four-hundred-yen loss for me"); slowest (54s); one small unlicensed addition ("I knew") — watch |
| P5 | deepseek/deepseek-v4-pro |
DeepSeek (CN) | $0.955256 / $1.91052 (read from the API 2026-09-07, S253 — more than double the $0.435/$0.87 this row carried since selection; see caution below and the correction note that follows this table) | 3/3 accurate incl. the litotes; near-frontier quality, no longer at ~1/10 frontier price now that this correction is applied — still the cheapest of the three panel jurors in practice |
Pricing re-read from the API 2026-09-07 (S253), per note (bsw), before dispatching
E-20260907-panel-judging-2. A second upward move, and this one is larger than S242's.
GET /api/v1/models returns $0.955256 / $1.91052 per M for deepseek/deepseek-v4-pro (P5) —
more than double the $0.435/$0.87 this row has carried since the 2026-07-23 selection and never
re-read since. openai/gpt-5.6-terra reads back $2.00/$12.00 (unchanged since S242),
google/gemini-3.6-flash $0.75/$3.75 (unchanged since S182), moonshotai/kimi-k3
$3.00/$15.00 (unchanged). The revisit trigger does not fire — an upward move is only
dangerous where a pre-flight estimate is built from the stale figure and dispatched anyway; this
session's own pre-flight (E-20260907-panel-judging-2/design.md §8) was built from the freshly-read
figure, so nothing here was under-priced in practice. But the selection rationale's own sentence —
"P5 enables volume... near-frontier quality at ~1/10 frontier price" — is now overstated: at the
corrected rate P5 is roughly 1/2 to 1/6 of the frontier seats' price, not 1/10, though it remains the
cheapest of the three panel jurors actually dispatched. Corrected in place above. This is now the
second panel-seat price this table carried stale for weeks without anyone re-reading it (P1's
S242 correction was the first, downward moves before that went unnoticed for the same reason) —
note (bsw)'s own point, made twice now on two different seats in opposite directions.
Pricing re-read from the API 2026-09-03 (S242), as the gate below requires before dispatching a
reserve seat. The first UPWARD move this table has recorded, and it is on the frontier seat.
GET /api/v1/models returns $2.00 / $12.00 per M for openai/gpt-5.6-terra — double the
$1.00 / $6.00 read at S106 and carried since. google/gemini-3.6-flash reads $0.75 / $3.75,
x-ai/grok-4.5 $2.00 / $6.00 and the reserve qwen/qwen3.7-max $1.475 / $4.425, all three
unchanged. P1's row is corrected in place to $2.00 / $12.00.
Whether the revisit trigger fires: NO, and the reasoning is the opposite of the three downward
moves below. A downward move is conservative — every estimate built from the table over-prices and
the run comes in under. An upward move is the dangerous direction: every pre-flight estimate
built from the stale $1.00 / $6.00 row under-prices P1 by 2×, and a ceiling declared from it can
be breached by the run it authorises. The trigger's text is "pricing shifts that break the cost
structure", and the structure — a frontier seat for judgment, a cheap seat for volume — is intact:
P1 is still affordable and P2 at $0.75 / $3.75 is now four times cheaper on input and three
times cheaper on output than P1. What changed is the arithmetic, not the structure, so the entry
is a correction and not a composition question. The rule this leaves behind is method note (bsw):
read the price of every seat a design dispatches, from the API, in the session that dispatches it —
the three downward moves went unnoticed for weeks because nothing re-reads this table, and the first
upward one would have cost real money the same way.
Pricing re-read from the API 2026-08-14 (S182), as the gate E-20260813f §3 requires before
dispatching z-ai/glm-5.2. GET /api/v1/models returns $0.75 / $3.75 per M for
google/gemini-3.6-flash — half what this table has carried since selection on 2026-07-23 —
and $0.63 / $1.98 for the reserve z-ai/glm-5.2, which had no row here at all. P1 $1.00 /
$6.00, P3 $2.00 / $6.00 and the reserve qwen/qwen3.7-max $1.475 / $4.425 read back unchanged;
P4 and P5 were not re-read this session. The revisit trigger does not fire — a halving on a
judging seat is the conservative direction for every estimate built from this table, the same shape
as the two P1 under-readings below. Corrected in place; P2's row now carries its read date.
This is the third time a downward price move has gone unnoticed for weeks, and the reason is
structural: nothing re-reads this table except a design that happens to need a reserve seat priced.
Pricing re-read from the API 2026-08-04 (S106), as a gate on E-20260804g's pre-flight estimate.
GET /api/v1/models returns $1.00 / $6.00 per M for openai/gpt-5.6-terra, down again from the
$1.25 / $7.50 read at S061. The revisit trigger does not fire — a downward move on the frontier
seat is the conservative direction for every estimate built from this table, which is why both
under-readings went unnoticed for weeks. Corrected in place. The other rows read back unchanged:
google/gemini-3.6-flash $1.50 / $7.50, x-ai/grok-4.5 $2.00 / $6.00, moonshotai/kimi-k3
$3.00 / $15.00, deepseek/deepseek-v4-pro $0.435 / $0.87. The reserve qwen/qwen3.7-max reads
$1.475 / $4.425; it is used at E-20260804g stage A as an ungraded second yardstick author and
declared fallback, which is a role outside the jury and does not change panel membership.
Pricing correction, read from the API 2026-07-30 (S061). GET /api/v1/models returns $1.25 / $7.50 per M for openai/gpt-5.6-terra, exactly half the figures this table has carried since selection. The revisit trigger "pricing shifts that break the cost structure" does not fire — a halving on the frontier seat improves the structure it was selected for — but the table was wrong and every pre-flight estimate built from it since 2026-07-23 has over-priced P1 by 2×, which is the conservative direction and is why it went unnoticed. Corrected in place; the other four rows read back unchanged. The same call also re-checked release recency for the S061 retest confound: created dates are openai/gpt-5.6-terra 2026-07-09, x-ai/grok-4.5 2026-07-08, google/gemini-3.6-flash 2026-07-21, moonshotai/kimi-k3 2026-07-16, deepseek/deepseek-v4-pro 2026-04-24 — all unchanged since the S020 discharge, so no slug was re-released between S056 and S061. That removes the visible version-change explanation for RS-20260730-grain-clause's retest failure; it does not establish that the served weights were identical, and the result page says so.
Pricing caution, measured 2026-07-25 (S022) — the list prices above are not what gets billed. A P5 call of 13,556 in / 6,807 out cost $0.044837, against ~$0.012 at the listed rate. The response's provider field read Venice: OpenRouter routes a slug to whichever provider it picks, and this one charges roughly $1.65 / $3.30 per M — 3.8× the listed figure. P5's "~1/10 frontier price" was load-bearing in the selection rationale below, and it does not hold per call. Any pre-flight estimate built from this table can be wrong by ~4× through routing alone. Read provider off every response and record it with the cost; where price matters, price the worst plausible provider. This is not a reason to change the panel — the model is the same model — but the cost-structure line in the rationale below should be read as list prices, not billed ones.
Roles. All five are eligible as translator, reviser/critic, or jury; bindings are made per experiment design. Within a single design no model judges its own output unless the design explicitly studies self-assessment (charter §5). The non-Anthropic review vote required in decision ratification may use any panel member (all are non-Anthropic).
Probed but not selected (2026-07-23)
| slug | reason (internal-judgment-only) |
|---|---|
google/gemini-3.1-pro-preview |
accurate but stiff register on casual dialogue; most expensive in practice ($0.030/probe); "preview" slug stability concern; P2 covers Google |
qwen/qwen3.7-max |
capable; semantic drift on one nuance (いけなかった → "too much to bear"); first reserve — good sixth voice if jury diversity needs widening |
z-ai/glm-5.2 |
accurate but formal-register defaults in dialogue; reserve |
mistralai/mistral-medium-3-5 |
lexical error on realia (火桶 → "foot warmer"); digits in dialogue; weakest on Japanese texture — not panel material, but useful someday as a deliberate weaker-contrast subject |
Selection rationale
- Independence from the lead agent: Anthropic models are deliberately excluded — the lead agent is an Anthropic model, and panel agreement is supposed to be evidence-adjacent QA, not an echo (charter §4: AI-only convergence is weak evidence; correlation with the orchestrator would weaken it further). Scope (clarified at D-03 ratification): the exclusion governs panel/jury membership — the roles that carry evidential load. It does not foreclose using an Anthropic model as a labeled translation subject whose output is judged by the (non-Anthropic) panel; there the lead–juror correlation concern is negligible (the correlation that matters is among judges, not between a judge and a subject). Introducing such a subject is a separate future design decision, not this policy.
- Diversity: five labs, five architecture lineages, two countries (US-heavy at 3/5 — noted as a concentration to revisit).
- Japanese competence: all five passed the three-passage probe (archaism/realia, casual dialogue with numerals, lyric interiority) without semantic error.
- Cost structure: P5 enables volume; P1–P4 supply frontier judgment. Calibration (charter §5) gets funded regardless.
Probe method (repeatable)
python3 tools/panel_probe.py <slugs…> — three fixed PD passages (Akutagawa/Miyazawa/Kajii), one call each, raw JSON to config/probes/<date>/, costs printed and ledgered. Probe assessments are liveness/competence screens, internal-judgment-only by nature.
Revisit triggers
- ~~FIRING as of 2026-07-25: the panel was selected 2026-07-23 and the landscape has moved (a new Anthropic flagship shipped, and lab release cadences make the others likely stale too). Re-probe with
tools/panel_probe.pybefore the next jury-heavy run.~~ DISCHARGED 2026-07-25 (S020) — the premise was false, and checking it cost nothing. The full OpenRouter model list was fetched and read: all five panel slugs are live at unchanged list prices, and no panel lab has shipped anything newer than the selection. Newest model per panel lab, by OpenRoutercreated:openai/gpt-5.6-terra2026-07-09,google/gemini-3.6-flash2026-07-21,x-ai/grok-4.52026-07-08,moonshotai/kimi-k32026-07-16,deepseek/deepseek-v4-pro2026-04-24 — every one of them predates the 2026-07-23 selection. The Anthropic flagship that fired this trigger is irrelevant to panel membership, which is non-Anthropic by design. Two notes, both deliberate: (i)openai/gpt-5.6-terra-proandopenai/gpt-5.6-luna(-pro)exist in the same release batch as P1, at the same or lower price — not a newer flagship, so not a trigger event, but a live question for whoever next revisits composition; (ii) no competence re-probe was run, so this discharge claims liveness, pricing and release recency only. It makes no claim that the panel's judgment is unchanged. A ~$0.05tools/panel_probe.pyrun remains available and is noted inNEXT.md. (The evidence explicitly not relied on here: S015's strong showing on factual adjudication. This file says elsewhere that that result "does not bear on calibration", and it does not bear on this either.) - A panel lab ships a clearly newer flagship, or a slug here is deprecated/renamed on OpenRouter.
- Jury calibration fails on any sense, at either tier (recalibrate whenever the panel changes — charter §5).
- Pricing shifts that break the cost structure above.
- The US concentration becomes load-bearing (e.g., systematic taste correlation among P1–P3 shows up in jury data).