Repository path: workshop/experiments/E-20260907-panel-judging-2/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260907-panel-judging-2 |
| status | frozen |
| created | 2026-09-07 |
| senses | accuracy, naturalness, voice, style-correspondence, cultural-mediation, affect |
| purpose | Readers of literary fiction in English who cannot read the source, meeting these texts as reading editions rather than as cribs (carried verbatim from E-20260803-a4-set, D-20260801-10). |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/plan.md, wiki/arms/ARM-first-judgment.md, workshop/experiments/E-20260803-a4-set/design.md, workshop/experiments/E-20260803-a4-set/materials/items.json, workshop/experiments/E-20260803-a4-set/analysis/scores.json, wiki/goodness-senses.md, config/models.md, workshop/translations/hirurgiya/R04-v1/translation.md, workshop/translations/monelle-paroles/R04-v1/translation.md |
E-20260907-panel-judging-2 — W2 step 4: re-judge the A4 set, and judge two more filed W3 translations, permanently provisional
Frozen 2026-09-07 (S253) before any call is dispatched. wiki/plan.md §W2 step 4: "Judge
translations. Proceeds on provisional labels only, permanently ... re-judge the A4 set, judge
every W3 translation filed since S099 that has a frozen log, and file scores on the translation
pages. Three blind non-Anthropic jurors, authorship stripped, order-swapped, meaning-preserving
micro-paraphrase null (S094's working null), under $2 per session."
Standing, stated first because it governs every number below. Tier D calibration is EXHAUSTED
(config/models.md, S250): NOT PASSED, permanently, not pending recalibration. No score this design
produces carries evidential weight; every artifact it touches stays provisional and
internal-judgment-only.
Scope of this installment. "Every W3 translation filed since S099" is, by count, 208 candidate
files (workshop/translations/**/translation.md with created > 2026-08-03) — far more than one
$2 session can judge. This design does two things and states plainly what it defers: (1) re-judges
the five A4 primary items plus the null pair, to measure whether scores are stable across a long
gap (S094 → S253, 159 sessions); (2) adds two new W3 translations filed since S099, one per two
of the three screened languages, chosen for a frozen translator's log and a single, non-ladder
regime. The positive control (BAR-D) and the log-prediction stage are dropped from this
installment to hold the declared cost near the plan's $2 figure — see §8. The remaining ~206
candidates are not this session's to judge; wiki/backlog.md carries the queue (§Hand-off).
1. The question, and the subject-rule sentence
Does this jury's scoring of the project's own translations survive a long gap unchanged, and what
do two more W3 translations — filed since the A4 promise was first kept — score under the same
instrument? Subject-rule sentence: this unit teaches whether the panel's quality judgments of
this project's translations are stable enough to trust as an ongoing record, and what two more
translated passages score on the same six-sense instrument used since S094 — a claim about
evaluating translations, not about the project's own apparatus for its own sake (wiki/tracks.md).
2. Materials
materials/items.json, SHA-256 recorded by the runner in every raw response file.
| item | work | pair | regime | role | words | status here |
|---|---|---|---|---|---|---|
| TAK | Ōgai, 高瀬舟 (1916) | JA→EN | R04 | primary | 517 | re-judged, verbatim from A4 |
| KUS | Sōseki, 草枕 ch. VII (1906) | JA→EN | R06 | primary | 587 | re-judged, verbatim from A4 |
| BAR | Andreyev, «Баргамот и Гараська» (1898) | RU→EN | R04 | primary | 441 | re-judged, verbatim from A4 |
| MAR | Sand, La Mare au Diable ch. II (1846) | FR→EN | R04 | primary | 569 | re-judged, verbatim from A4 |
| PAN | Arène, «La Mort de Pan» (1876) | FR→EN | R04 | primary | 477 | re-judged, verbatim from A4 |
| CLO | Arène, «Le Clos des Ames» (1876) | FR→EN | R04 | secondary | 503 | re-judged, verbatim from A4 |
| CLO-P | CLO + 10 frozen micro-paraphrase edits | — | — | null | 499 | re-judged, verbatim from A4 |
| HIR | Chekhov, «Хирургия» (1884), first two-thirds | RU→EN | R04 | primary | 1,074 | new, filed S246 (T-hirurgiya-R04-v1) |
| MON | Schwob, «Paroles de Monelle» (1894), 4 litanies | FR→EN | R04 | primary | 531 | new, filed S099/S130 (T-monelle-paroles-R04-v1) |
TAK/KUS/BAR/MAR/PAN/CLO/CLO-P are reused byte-for-byte from E-20260803-a4-set/materials/items.json
— same source slice, same translation, same SHA-256-verifiable text — so that any score difference
from S094 is attributable to the jury and not to a changed item. BAR-D (the positive control) is
excluded from this run, a cost-driven choice named in §8 and F1 below.
HIR and MON are drawn from the 208-file "since S099" set on these criteria: a screened language
(Russian and French are two of the three panel-competence-screened languages, config/models.md
S015), a single non-ladder R04 regime (not one cell of a multi-arm rule-set study), status:
frozen with a ## Translator's log section already committed, and no overlap in author with the
A4 set (Chekhov and Schwob are new authors; Arène appears in A4 but not among the new items). HIR
is out of the 441-587-word band (1,074 English words, the full two-thirds excerpt its own file
declares) — named here as a covariate, not concealed; §6 tests for a length effect. A third
candidate in the screened third language (Japanese) exists in principle but its Aozora Bunko source
text is not yet fetched and decoded in this session — deferred to the next installment rather
than adding fetch-and-decode work this session did not budget for.
Contamination, carried from each artifact's own front matter: TAK/KUS/BAR/MAR none-to-suspected
as at A4 (unchanged, not re-measured here); MON suspected (basis: monelle-paroles/contamination.md);
HIR suspected, declared on a search rather than a measurement (basis: the artifact's own note —
eight Gutenberg Garnett-Chekhov volumes searched by name and character, no comparator found).
Contamination bears on whether the lead's own rendering could be inflated by memorized published
English; it does not bear on whether this jury — which never produced any of these texts — can
score them, so it is not a selection gate here (unlike a design that needs an independent
comparator).
3. Procedure
Carried verbatim from E-20260803-a4-set §3 except where noted.
- Jurors P1, P2, P5 (
openai/gpt-5.6-terra,google/gemini-3.6-flash,deepseek/deepseek-v4-pro) — the same three as S020, S034, S086, S089, S094. Non-Anthropic (charter §5). The lead never judges its own translation, and every item here is the lead's. - Absolute rating, one text per call, stateless. No comparison, no shared context between items.
- Blind and authorship-stripped.
- Two passes, independently ordered (
hashlib.sha256of design-id/pass/item, as at A4), no shared context — the retest floor. - Strictly sequential (charter §6).
- Six senses, the post-
D-20260802-13wording, carried byte-for-byte from A4's runner (§0 above;consistencyomitted, length-dependent;perceived-source-carriageomitted, its own departure-level record is not collected here). - Dropped from this installment, and named rather than silently absent: the log-reading seat (P3) and the lead's own parallel prediction (A4's log-prediction test, already established and not the question this run asks); the edits neutrality check (already run at A4 on this exact edit list — re-running it on unchanged edits would be repetition, not verification); the positive control BAR-D (§8 names the cost trade-off).
4. Analysis, fixed before the run
Let x(i,s,j,p) be the score for item i, sense s, juror j, pass p; x̄(i,s,j) the mean over
p; X(i,s) the mean over j. PRIMARY = {TAK, KUS, BAR, MAR, PAN, HIR, MON} (seven quality
items; CLO/CLO-P are the null pair, reported separately per A4's amendment A10).
- Primary table:
X(i,s)for all nine items × six senses, printed whole. - Retest floor
R(s) = mean over (i,j) of |x(i,s,j,1) − x(i,s,j,2)|overPRIMARY; per-juror floors reported separately. - Paraphrase floor, per A4 amendment A1:
P(s,j) = |x̄(CLO,s,j) − x̄(CLO-P,s,j)|, compared against that juror's ownR(s,j). Pooled figure secondary only. - The cross-session comparison, the reason this design re-judges rather than only extends. For
the seven items shared with
E-20260803-a4-set(TAK, KUS, BAR, MAR, PAN, CLO, CLO-P), computeΔ_session(i,s) = X(i,s) − X_S094(i,s), readingX_S094fromE-20260803-a4-set/analysis/scores.json. Report the mean absoluteΔ_sessionper sense and overall, and compare it against this run's own within-session retest floorR(s)— if the across-session gap is no larger than the within-session gap, the instrument is stable across a 159-session gap; if materially larger, something drifted (a re-routed provider, a changed sampling default, or genuine rater drift) and the result says so rather than averaging over it. - Nuisance checks: (a) Spearman between item word count and
X(i,s)overPRIMARY, flagging that HIR's 1,074 words sit well outside the other six; (b) ceiling/floor fraction per juror; (c) pass-1-minus-pass-2 drift per juror; (d) per-juror mean. - No difference is described as a difference unless it exceeds the operative floor
(
max(R(s), mean_j P(s,j))), printed beside every reported gap.
5. Registered predictions
| # | prediction |
|---|---|
| 1 | The retest floor R (over PRIMARY, seven items) is greater than 0 and less than 1.00 scale points, comparable to A4's 0.233. |
| 2 | The paraphrase floor P(s,j) ≤ 2×R(s,j) on at least 5 of 6 senses for every juror (A4's A7 form). |
| 3 | The mean absolute cross-session gap Δ_session does not exceed twice this run's own R, on at least 4 of 6 senses — the instrument reads the same seven items about the same way 159 sessions later. |
| 4 | HIR and MON score no lower than 4.0 on any sense, per juror — this jury does not find either new item bad. |
| 5 | The between-item spread on PRIMARY exceeds the operative floor on at least 3 of 6 senses (A4's A6 form). |
6. Failure criteria
- F1 — no positive control this round. Because
BAR-Dis not dispatched, no gate exists to confirm this jury is still capable of detecting gross damage in this exact run; the result relies on A4's and S089's establishment of that capability and says so rather than re-certifying it. A future installment that wants a fresh capability check should re-includeBAR-Dor a successor. - F2 — the null. If
P(s,j) > 2×R(s,j)on more than one sense for any juror, prediction 2 is recorded FAILED andPbecomes the operative floor. - F3 — drift. Any juror's pass-1-minus-pass-2 mean exceeding 0.50 is reported and that juror's passes are shown separately as well as pooled.
- F4 — ceiling. Any juror at 7 on more than 50% of cells has its comparisons reported separately.
- F5 — cost. Any call billed above its declared per-call worst case is named in the result.
A hard stop, additional to F5 and specific to this session's $2 target: if the running sum of
usage.costacross dispatched calls reaches $1.90 before all cells are filled, no further calls are dispatched and the result reports exactly which cells are missing and why, rather than running over the plan's declared figure. - F6 — parse. A
finish_reason: lengthbody is a seat failure (note (b)); rejected, ledgered, retried once at the runner's built-in escalation.
7. What this design cannot do, written before it runs
- The jury is not calibrated; nothing here carries evidential weight (Tier D EXHAUSTED, permanent).
- No positive control this round (F1) — capability rests on A4/S089, not re-demonstrated here.
- n = 7 primary items, two new. Three languages of sixteen, as at A4.
- HIR is more than twice the word-length of the shortest item; any HIR-specific reading must account for that before comparing it to the others.
- The lead wrote every item. One translator, one agent.
- This is one installment of a 208-file backlog. It establishes stability and adds two items; it does not clear the plan's "every W3 translation since S099" clause, which the hand-off names as ongoing at cadence (matching W2 step 5's own precedent).
8. Cost, and the reserve declared before dispatch
Prices re-read from GET /api/v1/models 2026-09-07 (S253): openai/gpt-5.6-terra $2.00/$12.00
per M (unchanged from S242's reading), google/gemini-3.6-flash $0.75/$3.75 (unchanged since
S182), deepseek/deepseek-v4-pro $0.955256/$1.91052 — more than double the $0.435/$0.87 this
project's table has carried since 2026-07-23 selection, corrected in config/models.md by this
session. moonshotai/kimi-k3 (critic) $3.00/$15.00, unchanged.
| stage | calls | cap (max_tokens) |
worst case |
|---|---|---|---|
pre-run critic (P4, reasoning: low) |
1 | 16,000 | $0.24 |
| rating, pass 1 (9 items × 3 jurors) | 27 | 6,000 | $0.951 |
| rating, pass 2 | 27 | 6,000 | $0.951 |
| declared worst case | 55 | $2.14 |
Worst case built from max_tokens × each seat's highest list price (note (abc)); split across
jurors: P1 18 calls × $0.072 = $1.296, P2 18 × $0.0225 = $0.405, P5 18 × $0.01147 = $0.206, critic
$0.24 → $2.147. This exceeds the plan's stated $2 figure on worst case alone, which is why
BAR-D and the log stage are already dropped (§2, §3) and why F5's hard stop at $1.90 actual exists.
Realistic expectation, from A4's own ratio of actual to declared worst case (0.65 / 2.21 ≈ 0.29)
applied here: ≈ $0.62. 2026-09-07 is a fresh UTC day with $0 spent; the full $5.00 daily cap is
available regardless, so the binding constraint is the plan's own $2 step-budget line, not the
daily cap.
9. Independent pre-run critic
Required by CLAUDE.md §6 (any run over $0.50). Dispatched to moonshotai/kimi-k3 (P4) before any
rating call, reasoning: {"effort": "low"} per note (b)'s prescribed fix. runs/out/critic.json,
$0.0293462, VERDICT NEEDS-AMENDMENT, 10 findings, 2 BLOCKING. All ten accepted, applied below before
any rating call.
10. Amendments from the independent pre-run critic pass
F1 (BLOCKING) — cross-session comparison confounded by unverifiable juror-side drift. Accepted.
call.py already records provider and model (the dispatched slug) per response; OpenRouter does
not expose a finer version string than that. §4.4's prediction 3 is reclassified: it tests
"this project's panel, addressed by the same slugs, on the same provider-routing service" — not
"the identical served weights" — and the result states this qualification verbatim (matching S061's
practice: a slug/provider match does not establish identical served weights). Any provider mismatch
against A4's per-item provider is reported per item.
F2 (BLOCKING) — no mechanism catches a dead or flatlined juror once BAR-D is dropped. Accepted,
fixed analytically rather than by reinstating BAR-D (which would breach the cost line F7 also
flags). A degeneracy check is added to analysis/score.py: for each juror, the standard
deviation of that juror's scores across all (item, sense, pass) cells must exceed 0.30 on at least 4
of 6 senses. A juror failing this is flagged in the result as not usable for the stability
question, and its cells are reported but excluded from prediction 3's pooled figure.
F3 (non-blocking) — payload identity is claimed, not hashed. Accepted as a documentation fix.
call.py is copied byte-for-byte from E-20260803-a4-set/call.py (verified: diff on the two
files, no output) and rate_prompt() is copied byte-for-byte from A4's run.py (same PURPOSE,
SENSE_DEFS, JSON schema string). The one structural difference is item order, which is
independently re-randomised per pass by design (§3) and is therefore not a confound. This is stated
here rather than hashing full request bodies, which would not change the conclusion.
F4 (non-blocking) — HIR's length may distort the pooled retest floor. Accepted. analysis/score.py
reports R two ways: R_6 over the six items in the 441-587-word band (TAK, KUS, BAR, MAR, PAN,
MON) and R_7 over all seven PRIMARY items including HIR. Prediction 1 is tested against
R_6, matching A4's band; R_7 is reported beside it and the gap between them is itself reported as
a finding about whether length changes measured retest variance.
F5 (non-blocking) — the length check is underpowered and decorative at n=7 with one outlier. Accepted. The Spearman figure in §4.5(a) is reported as a descriptive flag only; no claim of a length effect (or its absence) is drawn from it, stated explicitly in the result.
F6 (non-blocking) — prediction 4 is trivially satisfiable under ceiling tendency. Accepted. Prediction 4's verdict is reported beside the F4 ceiling-fraction figures for the same items, and the result states plainly that a pass on prediction 4 alongside a high ceiling fraction is uninformative about quality and informative only about juror generosity.
F7 (non-blocking) — no pre-registered drop order if the $1.90 hard stop fires mid-run. Accepted.
Pre-registered now, before any rating call: the runner dispatches stage rate1 (all nine items,
all three jurors) to completion before stage rate2 begins. If the hard stop fires during rate2,
the drop is truncate to pass 1 only, for every item — the design falls back to a single-pass
absolute-rating result with no retest floor, reported as such, rather than a partial second pass
that would break the pairing §4.2-§4.4 depend on.
F8 (non-blocking) — the paraphrase floor's numerator (CLO/CLO-P, reused from A4) and the retest floor's denominator are on different footings. Accepted. §4.3's comparison is stated explicitly in the result as "a cross-session-stable null re-measured this session, compared against this session's own within-session retest floor" — not a same-session apples-to-apples ratio, which is what A4's original A1/A7 amendments assumed when both quantities came from one run.
F9 (non-blocking) — the installment's title and the plan's citation promise more than two items deliver, and item selection was not a blind pre-specification. Accepted, disclosed rather than hidden. Item selection was made by inspecting the front matter and content (language, regime, frozen-log status, word count, author overlap) of several of the 208 candidates before freezing this design — not by inspecting any score, which did not exist yet. That is ordinary material selection (as at A4 and every translation-selection step in this project), not selection on the dependent variable, and it is named here rather than implied to be a blind draw. §0's title already states "two more W3 translations" rather than "every"; no retitling is needed beyond that.
F10 (non-blocking) — contamination as a possible inflator of HIR/MON's own scores, not only a
selection property. Accepted. The result's write-up states, beside HIR's and MON's scores: both
are contamination: suspected (declared on a search, not a dependence-check measurement, since no
comparator was found to measure against), so a naturalness/style-correspondence score for either
that is unusually high should be read with that caveat, exactly as wiki/goodness-senses.md's
accuracy entry already asks for a comparable caveat on multi-sense halo.