Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: config/models.md · rendered 2026-09-09

Page metadata (front matter)
typenote
idmodels
statusactive
created2026-07-23
updated2026-09-07
linkswiki/decisions/resolved/D-20260723-03-panel-v1-composition.md, config/budget.md, framework/tierD-repaired-rules.md, wiki/findings/results/RS-20260907-panel-judging-2.md

The panel

The models the project reaches via OpenRouter (charter §5). Model slugs are configured here and nowhere else — specs and tools refer to roles; run records log resolved slugs as provenance. Revisit when the landscape changes (triggers below). This composition was ratified 2026-07-23 (decision D-20260723-03, resolved: independent adversarial review + non-Anthropic panel vote, both RATIFY) as a defensible, revisable working panel — not an optimum; changes go through the revisit triggers.

Calibration state, FINAL, set 2026-09-06 (S250): TIER D CALIBRATION IS EXHAUSTED. NOT CALIBRATED, permanently, not pending recalibration. No jury verdict carries evidential weight; every workshop self-assessment and every panel score anywhere in this repository is provisional and internal-judgment-only, permanently. The handbook (framework/v0.3/) continues to rest on workshop findings and traceability notes, never on a panel-jury verdict.

History of how this was reached (restructured 2026-07-25, charter A3; updated 2026-07-26, S034; instrument defects recorded 2026-07-27, S040; stage-1 gate declared unrepairable 2026-07-28, S050; repair arm closed resolved 2026-07-29, S055; the repaired instrument run to completion 2026-08-02, S086; 2026-09-04 (S244), Tom authorized exactly one further redesign (PROJECT.md §11) after S086's verdict, with the stated consequence that a failure declares the approach exhausted; that one redesign was ratified and run 2026-09-06 (S250), and it FAILED, on two independent dispositive gates — see the S250 entry below, which is now the final word on this row.**

Lead agent as translator (2026-07-25, charter §3/§5, A4). The lead translates as a labeled subject at no API cost. It never judges its own output; the non-Anthropic panel judges blind with authorship stripped. Panel membership stays non-Anthropic — the lead-correlation argument governs judges, and it binds harder now that the lead also produces subjects. This is the carve-out anticipated at D-03 ratification ("does not foreclose using an Anthropic model as a labeled translation subject whose output is judged by the non-Anthropic panel"), now exercised. Each lead translation declares contamination (whether published translations of that work plausibly sit in the lead's training data) and freezes its translator's log before evaluation is designed.

First calibration run executed 2026-07-25 (S014): Case A (Botchan, opening, upper-bounded) — wiki/findings/results/RS-20260725-calibration-caseA. This was a Tier P run. Outcome: the panel does not reproduce the documented Cohn-favoring reception record on the real test (Turney vs Cohn) — record-fit at/below chance on voice (0.55), affect (0.42), literary-quality (0.43); only naturalness reaches 0.80 and is register-cued. The sanity floor (Morri vs Cohn) passes near-ceiling. No sense is calibrated; no evidential weight on any sense. No US-taste cluster fired (Metric 2 firing rule not met). Instrument note: juror deepseek-v4-pro (P5) showed a 0.44 order-flip rate — down-weight until reps increase. Recalibration/next steps in the result page and NEXT.md.

Instrument note added 2026-09-01 (S237), from RS-20260901-inversion-habit. The first run to probe reasoning: {"effort": "low"} on P3 against a key rather than against its own full-effort answers. On a 523-item word-order coding task it cost $0.0073 a 24-item batch against $0.094 at full effort — a factor of thirteen — and failed both calibration gates, calling an INVERTED keyed line CANONICAL 7 times of 19 where P1 did so 3 times and P2 0. The errors are one-sided: it under-detects the construct rather than scattering. So the S234 note's "worth probing on P3" is narrowed: worth probing, and only against keyed items, reading the direction of the errors — self-agreement at six of eight positions, which is what S234 had, does not detect a seat that is systematically not looking. Note (bsq). Two further per-seat facts from the same run, both on a 24-item batched coding prompt: P1 needs cap 12000 (it spends ~5,000 completion tokens and returns clean; 4000 truncates before any answer), and P2 cannot do 24 items at any cap probed — at 12000 it spent 11,996 tokens and answered 18 — so it was run at batch 12, where it returns clean at ~5,700 tokens, and became the dearest seat of the three.

Instrument note added 2026-08-30 (S234), from RS-20260830b-rhyme-family. The first probe this project has run of all four reachable seats on the same two task shapes, at two caps and two reasoning settings, and it changes what a pre-flight may assume. On a nine-line structured answer at max_tokens 2500: P1 openai/gpt-5.6-terra clean at 9.0 s / $0.007056; P3 x-ai/grok-4.5 clean at 52.2 s / $0.016238, and with reasoning: {"effort": "low"} at 5.6 s / $0.002698 — six times cheaper, nine times faster, and its answer agreed with its own full-effort answer at six of eight positions; P2 google/gemini-3.6-flash truncated, and returned clean only at cap 6000 (16.3 s / $0.010747); the first reserve qwen/qwen3.7-max truncated, clean at 6000 but at 66.7 s / $0.026243; and the reserve z-ai/glm-5.2 returned an empty body after 6000 completion tokens, which extends note (brt) from short prompts to this shape. On a search-shaped prompt — find the one English rhyme best covering a list of senses — every seat collapsed: P1 empty at 2500 for $0.030528, unchanged by low effort, P3 no return in 100 s twice, P2 clean only at 6000 for $0.019511. Two conclusions for pre-flights. (i) A cap is a per-seat, per-task-shape measurement, not a number carried across a session — note (bsf), fourth firing. (ii) reasoning: {"effort": "low"} is worth probing on P3 and is worth nothing on P1; it is a per-seat property, not a lever. No panel composition changed and no reserve was promoted: stage F ran on P1 alone and stage S on P2 alone, disjoint by design, and the probe is why.

Instrument note added 2026-09-05 (S247), from E-20260905-tierD-design-v3's pre-run critic dispatch. A third task shape — critique a ~50 KB frozen experimental design and return a structured numbered-findings response — was probed on the two non-panel reserves plus nvidia/nemotron-3-ultra-550b-a55b (a non-panel seat used before only for one-off adversarial review, e.g. S108). qwen/qwen3.7-max returned a clean, on-topic critique at max_tokens 16,000 ($0.0929, 205.5s). z-ai/glm-5.2 reproduced (brt) on this shape too, now at a much higher cap than S234's: 66,674 characters of on-topic reasoning and zero returned content at max_tokens 20,000 ($0.108). nvidia/nemotron-3-ultra-550b-a55b failed three ways in three attempts: default effort exhausted 16,000/16,000 tokens on reasoning alone with no answer; effort: low returned finish_reason: error from the Venice routing provider (extending (bps) — the parameter is not portable to this provider on this model); default effort at max_tokens 32,000 is recorded in workshop/experiments/E-20260905-tierD-design-v3/critique/. This project now has three independent seat/task-shape pairs where a large max_tokens cap alone does not fix a reasoning seat that has not converged (rhyme-search at S234, this critique shape at S247, on two different seats) — the fix that works is a different seat, not a bigger cap, once a cap in the tens of thousands has already failed once.

Instrument note added 2026-07-25 (S015), from RS-20260725-anchor-verification. On a factual-adjudication task — not a quality judgment — all four non-Anthropic panel members used (P1, P2, P3, P5) performed strongly: each rejected 20/20 planted false claims, including 6/6 refutable only against the specific stored files, at false-alarm rates of 0.083–0.114, giving discrimination scores of 0.886–0.917; abstention was ~0 and held-out true controls were supported 3/3 by every model. All four also passed a six-item competence screen in Russian, French and Japanese (≥5/6 each; the panel had previously been probed on Japanese only). This does not bear on calibration — calibration is about matching a human reception record on matters of quality, and the panel remains NOT CALIBRATED — but it does mean the panel is usable as a checking instrument for textual and linguistic fact, in the failing direction (charter §4 still forbids treating their agreement as validation). Two cautions from the same run: P5 alone missed the jade/kingfisher polysemy of 翡翠 on the screen; and support for evaluative claims (0.958) ran higher than for descriptive ones (0.890), i.e. these models agree most readily where there is least to check.

Panel v1 (selected 2026-07-23; probe: config/probes/2026-07-23/)

# slug lab (country) list price in/out per M probe notes (internal-judgment-only)
P1 openai/gpt-5.6-terra OpenAI (US) $2.00 / $12.00 (read from the API 2026-09-03, S242 — the first UPWARD move recorded here; this row read $2.50 / $15.00 from selection until 2026-07-30, $1.25 / $7.50 until 2026-08-04, and $1.00 / $6.00 until 2026-09-03) 3/3 accurate; fast (3.1s), terse token use → cheapest frontier call in practice ($0.0028)
P2 google/gemini-3.6-flash Google (US) $0.75 / $3.75 (read from the API 2026-08-14, S182; this row read $1.50 / $7.50 from selection until then) 3/3 accurate; heavy hidden reasoning (~1.7k tok, $0.013); clean idiomatic register
P3 x-ai/grok-4.5 xAI (US) $2.00 / $6.00 3/3 accurate; clean on the Kajii litotes ("This was rather bad" — though P5 produced the identical rendering, so not uniquely best); $0.0037
P4 moonshotai/kimi-k3 Moonshot (CN) $3.00 / $15.00 3/3 accurate with literary flair ("a twenty-four-hundred-yen loss for me"); slowest (54s); one small unlicensed addition ("I knew") — watch
P5 deepseek/deepseek-v4-pro DeepSeek (CN) $0.955256 / $1.91052 (read from the API 2026-09-07, S253 — more than double the $0.435/$0.87 this row carried since selection; see caution below and the correction note that follows this table) 3/3 accurate incl. the litotes; near-frontier quality, no longer at ~1/10 frontier price now that this correction is applied — still the cheapest of the three panel jurors in practice

Pricing re-read from the API 2026-09-07 (S253), per note (bsw), before dispatching E-20260907-panel-judging-2. A second upward move, and this one is larger than S242's. GET /api/v1/models returns $0.955256 / $1.91052 per M for deepseek/deepseek-v4-pro (P5) — more than double the $0.435/$0.87 this row has carried since the 2026-07-23 selection and never re-read since. openai/gpt-5.6-terra reads back $2.00/$12.00 (unchanged since S242), google/gemini-3.6-flash $0.75/$3.75 (unchanged since S182), moonshotai/kimi-k3 $3.00/$15.00 (unchanged). The revisit trigger does not fire — an upward move is only dangerous where a pre-flight estimate is built from the stale figure and dispatched anyway; this session's own pre-flight (E-20260907-panel-judging-2/design.md §8) was built from the freshly-read figure, so nothing here was under-priced in practice. But the selection rationale's own sentence — "P5 enables volume... near-frontier quality at ~1/10 frontier price" — is now overstated: at the corrected rate P5 is roughly 1/2 to 1/6 of the frontier seats' price, not 1/10, though it remains the cheapest of the three panel jurors actually dispatched. Corrected in place above. This is now the second panel-seat price this table carried stale for weeks without anyone re-reading it (P1's S242 correction was the first, downward moves before that went unnoticed for the same reason) — note (bsw)'s own point, made twice now on two different seats in opposite directions.

Pricing re-read from the API 2026-09-03 (S242), as the gate below requires before dispatching a reserve seat. The first UPWARD move this table has recorded, and it is on the frontier seat. GET /api/v1/models returns $2.00 / $12.00 per M for openai/gpt-5.6-terra — double the $1.00 / $6.00 read at S106 and carried since. google/gemini-3.6-flash reads $0.75 / $3.75, x-ai/grok-4.5 $2.00 / $6.00 and the reserve qwen/qwen3.7-max $1.475 / $4.425, all three unchanged. P1's row is corrected in place to $2.00 / $12.00.

Whether the revisit trigger fires: NO, and the reasoning is the opposite of the three downward moves below. A downward move is conservative — every estimate built from the table over-prices and the run comes in under. An upward move is the dangerous direction: every pre-flight estimate built from the stale $1.00 / $6.00 row under-prices P1 by 2×, and a ceiling declared from it can be breached by the run it authorises. The trigger's text is "pricing shifts that break the cost structure", and the structure — a frontier seat for judgment, a cheap seat for volume — is intact: P1 is still affordable and P2 at $0.75 / $3.75 is now four times cheaper on input and three times cheaper on output than P1. What changed is the arithmetic, not the structure, so the entry is a correction and not a composition question. The rule this leaves behind is method note (bsw): read the price of every seat a design dispatches, from the API, in the session that dispatches it — the three downward moves went unnoticed for weeks because nothing re-reads this table, and the first upward one would have cost real money the same way.

Pricing re-read from the API 2026-08-14 (S182), as the gate E-20260813f §3 requires before dispatching z-ai/glm-5.2. GET /api/v1/models returns $0.75 / $3.75 per M for google/gemini-3.6-flash — half what this table has carried since selection on 2026-07-23 — and $0.63 / $1.98 for the reserve z-ai/glm-5.2, which had no row here at all. P1 $1.00 / $6.00, P3 $2.00 / $6.00 and the reserve qwen/qwen3.7-max $1.475 / $4.425 read back unchanged; P4 and P5 were not re-read this session. The revisit trigger does not fire — a halving on a judging seat is the conservative direction for every estimate built from this table, the same shape as the two P1 under-readings below. Corrected in place; P2's row now carries its read date. This is the third time a downward price move has gone unnoticed for weeks, and the reason is structural: nothing re-reads this table except a design that happens to need a reserve seat priced.

Pricing re-read from the API 2026-08-04 (S106), as a gate on E-20260804g's pre-flight estimate. GET /api/v1/models returns $1.00 / $6.00 per M for openai/gpt-5.6-terra, down again from the $1.25 / $7.50 read at S061. The revisit trigger does not fire — a downward move on the frontier seat is the conservative direction for every estimate built from this table, which is why both under-readings went unnoticed for weeks. Corrected in place. The other rows read back unchanged: google/gemini-3.6-flash $1.50 / $7.50, x-ai/grok-4.5 $2.00 / $6.00, moonshotai/kimi-k3 $3.00 / $15.00, deepseek/deepseek-v4-pro $0.435 / $0.87. The reserve qwen/qwen3.7-max reads $1.475 / $4.425; it is used at E-20260804g stage A as an ungraded second yardstick author and declared fallback, which is a role outside the jury and does not change panel membership.

Pricing correction, read from the API 2026-07-30 (S061). GET /api/v1/models returns $1.25 / $7.50 per M for openai/gpt-5.6-terra, exactly half the figures this table has carried since selection. The revisit trigger "pricing shifts that break the cost structure" does not fire — a halving on the frontier seat improves the structure it was selected for — but the table was wrong and every pre-flight estimate built from it since 2026-07-23 has over-priced P1 by 2×, which is the conservative direction and is why it went unnoticed. Corrected in place; the other four rows read back unchanged. The same call also re-checked release recency for the S061 retest confound: created dates are openai/gpt-5.6-terra 2026-07-09, x-ai/grok-4.5 2026-07-08, google/gemini-3.6-flash 2026-07-21, moonshotai/kimi-k3 2026-07-16, deepseek/deepseek-v4-pro 2026-04-24 — all unchanged since the S020 discharge, so no slug was re-released between S056 and S061. That removes the visible version-change explanation for RS-20260730-grain-clause's retest failure; it does not establish that the served weights were identical, and the result page says so.

Pricing caution, measured 2026-07-25 (S022) — the list prices above are not what gets billed. A P5 call of 13,556 in / 6,807 out cost $0.044837, against ~$0.012 at the listed rate. The response's provider field read Venice: OpenRouter routes a slug to whichever provider it picks, and this one charges roughly $1.65 / $3.30 per M — 3.8× the listed figure. P5's "~1/10 frontier price" was load-bearing in the selection rationale below, and it does not hold per call. Any pre-flight estimate built from this table can be wrong by ~4× through routing alone. Read provider off every response and record it with the cost; where price matters, price the worst plausible provider. This is not a reason to change the panel — the model is the same model — but the cost-structure line in the rationale below should be read as list prices, not billed ones.

Roles. All five are eligible as translator, reviser/critic, or jury; bindings are made per experiment design. Within a single design no model judges its own output unless the design explicitly studies self-assessment (charter §5). The non-Anthropic review vote required in decision ratification may use any panel member (all are non-Anthropic).

Probed but not selected (2026-07-23)

slug reason (internal-judgment-only)
google/gemini-3.1-pro-preview accurate but stiff register on casual dialogue; most expensive in practice ($0.030/probe); "preview" slug stability concern; P2 covers Google
qwen/qwen3.7-max capable; semantic drift on one nuance (いけなかった → "too much to bear"); first reserve — good sixth voice if jury diversity needs widening
z-ai/glm-5.2 accurate but formal-register defaults in dialogue; reserve
mistralai/mistral-medium-3-5 lexical error on realia (火桶 → "foot warmer"); digits in dialogue; weakest on Japanese texture — not panel material, but useful someday as a deliberate weaker-contrast subject

Selection rationale

Probe method (repeatable)

python3 tools/panel_probe.py <slugs…> — three fixed PD passages (Akutagawa/Miyazawa/Kajii), one call each, raw JSON to config/probes/<date>/, costs printed and ledgered. Probe assessments are liveness/competence screens, internal-judgment-only by nature.

Revisit triggers