Repository path: workshop/experiments/E-20260803-a4-set/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260803-a4-set |
| status | frozen |
| created | 2026-08-03 |
| updated | 2026-08-03 |
| senses | accuracy, naturalness, voice, style-correspondence, cultural-mediation, affect |
| purpose | Readers of literary fiction in English who cannot read the source, meeting these texts as reading editions rather than as cribs. Declared because every evaluation must state its purpose parameter (D-20260801-10). |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-first-judgment.md, wiki/findings/results/RS-20260802c-regime-scoring.md, wiki/goodness-senses.md, config/models.md, framework/control-arm-spec.md, workshop/translations/clos-des-ames/R04-v1/translation.md, wiki/findings/results/RS-20260802-tierD-verdict.md |
E-20260803-a4-set — the A4 promise: five filed lead translations judged on their own terms, on a floor that means something
Frozen 2026-08-03 (S094) after the fresh translation and its log were committed at 9ab4c89 and
before any call was dispatched. ARM-first-judgment step 2. Translation limb:
T-clos-des-ames-R04-v1 (Arène, FR→EN), translated in session and frozen before this file existed.
0. Standing, stated first because it governs every number below
Tier D is NOT PASSED (RS-20260802-tierD-verdict, S086; config/models.md). No score this
design produces carries evidential weight, none may support a framework recommendation, and every
artifact it touches stays provisional and internal-judgment-only. ARM-first-judgment was
unblocked by that failure and pre-committed to running in exactly this mode.
Two further obligations, both new since step 1 and both binding here. (i) D-20260802-13 struck
naturalness's source-conditional escape clause on 2026-08-02; the wording put to the jury in §3
is therefore NOT the wording step 1 used, and the two runs' naturalness numbers are not
comparable. (ii) The new sense perceived-source-carriage is deliberately not scored here — its
own entry requires a departure-level record which this design does not collect, and it has never
been through Tier D. Scoring it as a bare 1–7 item would violate the sense's own rule.
1. The question, and the subject-rule sentence
What do this project's own translations actually score, per sense, when read blind against their sources — and does the translator's own frozen record of what was hard predict where blind readers find the translation weakest?
Fifty-eight filed translations, sixteen source languages, and until S089 not one had ever been judged
for quality (wiki/reassessment-2026-08-01.md §1.3). S089 scored six pairs — a draft against its
own revision. No translation of this project has ever been judged on its own terms, which is what
charter §3 A4 promises and what wiki/tracks.md §Deliverables names as T3's current deliverable.
Subject-rule sentence (wiki/tracks.md; continue-prompt.md §4.5): this unit teaches whether a
translator's own written account of the difficulties of a passage locates where readers of the
finished translation find it weakest — asked on six translations, five of them made years of sessions
before anyone thought of scoring them. That is a claim about the practice of translating and about
how translations are evaluated. The jury is the means, not the subject.
The wire between the limbs, in one sentence. The study limb cannot say that any difference
between two of these translations is a difference at all without a resolution floor; a floor requires
edits that are certainly quality-neutral; and only prose whose every decision the translator has just
recorded can be edited with that certainty — so the session translated «Le Clos des Ames» in order
to have a text it could paraphrase without changing its quality, and that paraphrase is the null
this instrument has never had (RS-20260802c §3, §10.1; note (bhk)).
2. Materials
materials/build.py asserts every slice and refuses to emit a manifest if any assertion fails.
materials/items.json carries SHA-256 0d32bd3ca022ec8b931bafc2f06e46e2b96b195d56ce78e2ffee99583b5e0053,
recorded by the runner in every raw response file. Sources and translations are the filed
artifacts, sliced source-and-target together.
| item | work | pair | regime | role | words |
|---|---|---|---|---|---|
| TAK | Ōgai, 高瀬舟 (1916), opening | JA→EN | R04 | primary | 517 |
| KUS | Sōseki, 草枕 ch. VII (1906), first 3 of 12 ¶ | JA→EN | R06 | primary | 587 |
| BAR | Andreyev, «Баргамот и Гараська» (1898), unit B | RU→EN | R04 | primary | 441 |
| MAR | Sand, La Mare au Diable ch. II (1846) | FR→EN | R04 | primary | 569 |
| PAN | Arène, «La Mort de Pan» (1876) | FR→EN | R04 | primary | 477 |
| CLO | Arène, «Le Clos des Ames» (1876) — secondary | FR→EN | R04 | secondary | 503 |
| CLO-P | CLO + 10 frozen micro-paraphrase edits | — | — | null | 499 |
| BAR-D | BAR + variant-F1, 8 documented damage sites |
RU→EN | R04 | ctrl-pos | 461 |
Three languages, not sixteen, and the reason is not laziness. Only Russian, French and Japanese
have ever been screened for panel competence (config/models.md, S015: six items, ≥5/6 per juror).
An accuracy judgment from a juror who cannot read the source is not a weak measurement but a
different one, so the banked German, Chinese, classical-Chinese, Polish, Bengali, Italian, Hungarian,
Swedish, Danish, Korean, Portuguese, Spanish, Finnish and Dutch artifacts are all excluded by that
rule. This is a real narrowing of the "cross-language set" the arm asked for and is declared as
one. Within the screened three, the regimes that exist are R04, R06 and R07; the one R07 artifact
in a screened language with an aligned source span (T-may-night-golova-R07-v1) is 1,457 words,
outside the length band, so the set is R04 ×4 and R06 ×1.
Length band 441–587 words, by construction. KUS is sliced to its first 3 of 12 paragraphs, source and target together, by the pre-registered rule take the fewest leading paragraphs whose English reaches 400 words; every other item is the whole filed span. Absolute rating of a longer text gives more opportunities to find a fault, so length is a confound and it is bounded here rather than argued away.
Two items share an author. PAN and CLO are both Arène, both R04, both the lead. That is a declared limitation on work-spread and a small compensation: it is the one place where a translation made before this design existed and one made inside it can be compared on the same author.
2.1 The null, and the edit list that makes it a null
RS-20260802c §3 established that step 1's null did not work: two byte-identical texts, all eighteen
cells recognised the identity, the floor came out at exactly 0.0000, and "above the noise floor"
became a sentence the instrument could not say. The remedy its own critic proposed and this design
executes: a meaning-preserving micro-paraphrase with a frozen edit list.
CLO-P is CLO with exactly ten substitutions, each asserted by build.py to match once and only
once, all listed in materials/edits.json:
| # | from | to |
|---|---|---|
| 1 | the whole of his close | his whole close |
| 2 | very old and badly kept | very old and poorly kept |
| 3 | right at the bottom | at the very bottom |
| 4 | must at one time have formed | must once have formed |
| 5 | meeting it for the first time | coming on it for the first time |
| 6 | having long since worn away | having long ago worn away |
| 7 | in cherry time | at cherry time |
| 8 | of every bourgeois felicity | of all bourgeois felicity |
| 9 | has gone on trembling ever since | has been trembling ever since |
| 10 | in a group around him | in a group about him |
Every edit is a function-word or near-synonym substitution inside one clause. None changes propositional content, sentence structure, paragraph shape, punctuation, the handling of any culture-bound item, or any decision entered in the translator's log — the thirteen logged decisions were checked one by one against the list before it was frozen. Net −4 words on 503.
Why this null works where step 1's did not: the jurors never see the two texts together. Every call in this design rates one text. CLO and CLO-P arrive in different calls, in different passes, with no shared context, so no juror can recognise that a relationship exists — which is exactly what made step 1's identical-text null measure identity detection instead of resolution.
2.2 The contamination gate, which fired twice before admitting the fresh work
Run before any locus was chosen and before anything was translated (CLAUDE.md; note (bcd)), on
a discard rule frozen in advance: longest common run ≥ 12 tokens, or ≥ 1 shared 12-gram → discard.
| candidate | comparator (whole, never displayed) | run | 12-grams | nulls | verdict |
|---|---|---|---|---|---|
| Daudet, «Le Phare des Sanguinaires» | anon., Letters from My Windmill (PG 30442) | 14 | 3 | 4/4/3, 0 | DISCARDED |
| Sand, La Mare au Diable ch. XI | Ives (PG 12816) | 16 | 13 | 4/4, 0 | DISCARDED |
| Sand, La Mare au Diable ch. XI | Sedgwick (PG 21993) | 13 | 2 | 3/2, 0 | DISCARDED |
| Arène, «Le Clos des Ames» | none exists (two search routes) | — | — | — | admitted |
The second firing lands on a work this design also scores, and it forced a re-check of a published
figure. T-mare-au-diable-R04-v1 is filed contamination: none, measured at S030 on ch. II against
those same two comparators. Ch. XI, same translator, same comparators, returns 16 tokens and 13
shared 12-grams. That does not falsify the ch. II figure, so the ch. II figure was recomputed on
the filed text before MAR was admitted as an item: materials/gate3.py returns 9 tokens against
Sedgwick and 10 against Ives, 0 twelve-grams in both, nulls 2–4 — the published figure reproduces
and MAR stays. What the pair of measurements establishes is that contamination: is a property of
the span, not of the work, which the project's own front-matter field does not distinguish. That is
method work and is recorded in the result's limits, not made this session's subject.
3. Procedure
- Jurors P1, P2, P5 (
openai/gpt-5.6-terra,google/gemini-3.6-flash,deepseek/deepseek-v4-pro) — the same three as S020, S034, S086 and S089, so this extends the instrument rather than replacing it. Non-Anthropic (charter §5). The lead never judges its own translation, and every translation here is the lead's. - Absolute rating, one text per call. The juror sees the source passage and one English
translation, and scores it 1–7 on six senses with a one-sentence
why. No comparison, no preference, no second text. This item format has never been used in this project, and §9 says what follows from that. - Blind and authorship-stripped. The juror is not told who made the translation, what regime produced it, that other translations are being scored, or that any two items are related.
- Two passes. Every (item, juror) is dispatched twice, in two passes with different item orders and no shared context. The two passes are byte-identical prompts, so the difference between them is the instrument's own noise — the retest floor.
- Strictly sequential. Judgment is never parallelized (charter §6); the runner enforces it.
- Six senses, from
wiki/goodness-senses.mdas it stands afterD-20260802-13.accuracy,voice,style-correspondence,cultural-mediationandaffectare carried verbatim fromE-20260802c's runner, which carried them fromtools/run_calibration.py, so the same words go to the same jury as in the Tier D runs.naturalnessis rewritten to the post-D-20260802-13definition: scored on the target text alone, against a named register anchor — unmarked / literary-contemporary (A-mchugh-presence), which is the register the declared purpose implies.consistencyis omitted as in step 1 (its own entry makes its weight a function of length, and these are 441–587-word spans);perceived-source-carriageis omitted for the reason in §0. - A separate, earlier stage: the log-reading seat. P3 (
x-ai/grok-4.5), which takes no part in the rating stage, is given the frozen translator's log of one item and nothing else — no source, no translation, no title beyond the work's name — together with the six sense definitions, and is asked to rank the six senses from most to least likely to receive the lowest score. Dispatched in full before the first rating call, so no rating can influence it. - The lead's own parallel prediction is written into
analysis/lead-predictions.jsonand committed before the critic pass and before any dispatch. It isinternal-judgment-onlyand is reported beside P3's, not instead of it.
4. Analysis, fixed before the run
Let x(i, s, j, p) be the score for item i, sense s, juror j, pass p. Let
x̄(i, s, j) = mean over p, and X(i, s) = mean over j.
- Primary table:
X(i, s)for all 8 items × 6 senses, printed whole, no aggregation hidden. - The retest floor
R(s) = mean over (i, j) of |x(i,s,j,1) − x(i,s,j,2)|, andR = mean over s. Per-juror floors reported separately. This is the instrument's noise on identical text. - The paraphrase floor
P(s) = |X(CLO, s) − X(CLO-P, s)|, andP = mean over s. The comparison ofPwithRis the null's whole point: ifP ≈ Rthe ten edits are quality-neutral to this jury andRis the operative floor; ifP > Rthe edits are not quality-neutral andPis the operative floor. Both are reported whatever they show. - The positive control
D(s) = X(BAR, s) − X(BAR-D, s). Registered gate — see F1. - No difference between two items on a sense is described as a difference unless it exceeds the operative floor, and every reported difference is printed with the floor beside it.
- The log-prediction test. For each of the six logged items, rank the six senses by
X(i, s)ascending (1 = lowest), ties given the mean rank. Letr(i)be the rank of the sense the log-reading seat named first. The statistic isr̄ = mean over the six items. Under the null of no relation, eachr(i)is uniform on {1..6}, soE[r̄] = 3.5; the exact null distribution is computed by full enumeration of 6^6 = 46,656 outcomes and the achieved one-sided p is reported. Criterion, fixed here:r̄ ≤ 2.20fires. The lead's own prediction is scored the same way and reported beside it, labelledinternal-judgment-only. - Nuisance checks, computed and reported whatever they show: (a) Spearman between item word
count and
X(i, s)over the five primary items, per sense; (b) the fraction of cells at the ceiling (7) and at the floor (1), per juror; (c) pass-1 minus pass-2 mean per juror, as a drift diagnostic; (d) per-juror mean score, since an absolute scale has no order-swap to cancel juror-level offset. control-arm-specR3/R4: five primary items against R4's 22–30 per stratum. No estimate, no interval and no stratified summary is claimed anywhere. Every number is descriptive.
5. Registered predictions
| # | prediction |
|---|---|
| 1 | The retest floor R is greater than 0 and less than 1.00 scale points. |
| 2 | The paraphrase floor P is within 0.50 of R — the ten edits are quality-neutral to this jury. |
| 3 | F1 fires: D(accuracy) ≥ 2.00 and D(accuracy) is the largest of the six D(s). |
| 4 | No primary item scores below 4.0 on any sense — this jury does not find any of these translations bad. |
| 5 | accuracy has the highest mean over the five primary items of the six senses; affect or style-correspondence the lowest. |
| 6 | The log-prediction test fires: r̄ ≤ 2.20 for the P3 seat. |
| 7 | The between-item spread on any one sense over the five primary items exceeds the operative floor — i.e. this instrument can tell these five translations apart at all. |
6. Failure criteria
- F1 — the format gate. If
D(accuracy) < 3 × R(accuracy), the absolute-rating format has not been shown to separate grossly damaged prose from clean prose, and no between-item comparison is reported at all; only the floors, the control and the failure. - F2 — the null. If
P(s) > 2 × R(s)on more than one of the six senses, the paraphrase is not quality-neutral,Pbecomes the operative floor, and prediction 2 is recorded FAILED. - F3 — drift. If any juror's pass-1 minus pass-2 mean exceeds 0.50 scale points, it is reported as a drift and that juror's two passes are additionally reported separately.
- F4 — ceiling. Any juror at 7 on more than 50% of its cells has its between-item comparisons reported separately as well as pooled.
- F5 — cost. Any call billed above its declared per-call worst case is named in the result.
- F6 — parse. A
finish_reason: lengthbody is a seat failure, never a partial answer (note (b)); it is rejected, its cost ledgered, and it is retried once at a raised cap.
7. Cost, and the reserve declared before dispatch
| stage | calls | cap | worst case from max_tokens |
|---|---|---|---|
| pre-run critic | 1 | 24,000 | $0.19 |
| log-reading seat (P3) | 6 | 6,000 | $0.22 |
| rating, pass 1 | 24 | 6,000 | $0.90 |
| rating, pass 2 | 24 | 6,000 | $0.90 |
| declared worst case | 55 | $2.21 |
Worst cases are built from max_tokens and the highest list price on the seat, per note (abc), and
sized from measured reasoning appetite per note (bhq) — the S092/S093 bodies spent 2,056–7,119
completion tokens on reasoning for short answers, so 6,000 is the answer plus that appetite.
2026-08-03 is a fresh UTC day with no rows; the full $5.00 is available. The realistic
expectation from step 1's 60 calls at $0.516 is ≈ $0.50.
8. What this design cannot do, written before it runs
- The jury is not calibrated. Tier D NOT PASSED. Nothing here carries evidential weight.
- The item format is new. Tier D and step 1 both used paired presentation; nothing established about this jury transfers automatically to absolute rating, which is why F1 exists as a gate rather than a diagnostic.
- n = 5 primary items, against
control-arm-specR4's 22–30. No estimate is claimed. - Three languages, four regimes reduced to two, two items sharing an author (§2).
- The log-prediction stage is partly a readback. Logs written after the project adopted its own
sense vocabulary use that vocabulary —
T-mort-de-pan-R04-v1's log has section headings named Accuracy against the source, Style correspondence, Naturalness and consistency. A seat reading that log is being handed the answer in the project's own words. Per item, whether the log names senses explicitly is recorded in the result table, and the test is reported both overall and restricted to the logs that do not. - One translator, one agent. Nothing here is about translation in general.
9. Amendments from the independent pre-run critic pass
runs/critic_kimi-k3_try1.txt, seat P4 (moonshotai/kimi-k3), verdict NEEDS-AMENDMENT, eleven
findings, three BLOCKING. All eleven accepted; one accepted with its mechanism corrected and half of
its remedy refused in writing. Applied before any rating or log call was dispatched. Note (rr)'s
streak, broken at S093, resumes. Note (b)'s prescribed fix — cap the reasoning effort on this
seat, do not rely on a generous max_tokens — was applied on the first dispatch this time, and
the body came back stop at $0.07026 against S093's $0.2215 for zero characters on the same slug.
A1 (finding 1, BLOCKING) — the null's two quantities are put on the same footing. R was a mean
absolute per-call retest difference and P a difference of juror-averaged means, so juror
disagreement in direction cancels in P and noise shrinks it. P is now computed per juror,
P(s, j) = |x̄(CLO, s, j) − x̄(CLO-P, s, j)|, and compared against that same juror's retest
differences. The pooled figure is retained as a secondary line only.
A2 (finding 2, BLOCKING) — accepted in its risk, corrected in its mechanism, half its remedy
refused with a reason. The finding is that a juror seeing the same source twice with near-identical
targets can infer the relationship, converting absolute rating into implicit comparison. The
mechanism requires memory across calls and there is none: call.py issues one stateless HTTP
request per rating with a single user message and no conversation history, so no juror can know that
any other item exists. This is the same correction E-20260802c made to the same argument (its A1).
Refused: showing CLO and CLO-P once each, because it would destroy the per-juror pairing A1
requires. Adopted: the dispatch order within a pass is now forced to place CLO and CLO-P, and BAR
and BAR-D, at maximal separation, as belt-and-braces; and the result page says recognition is
precluded by statelessness, not merely mitigated.
A3 (finding 3, BLOCKING) — the log-prediction null is now conditional on the observed tie
structure. Enumerating 6^6 uniform outcomes is invalid when X(i, s) has ties, which ceiling
clustering makes likely. The null is now: for each item, the six senses' ranks are the observed
rank multiset (ties at mean rank), and under the null the seat's named sense is one of the six
uniformly at random. The exact null is the enumeration of the product of the six items' observed
multisets — still 6^6 = 46,656 outcomes, but drawn from the real ranks. The uniform-{1..6} figure is
not reported.
A4 (finding 4) — the claim is bounded to this translator. Every statement of the log-prediction result is restricted to this lead's logs for these six items; §1's subject-rule sentence is read under that restriction, and no sentence about translators in general may be written from it.
A5 (finding 5) — prediction 3 is aligned with gate F1. It now reads: D(accuracy) ≥ max(2.00,
3 × R(accuracy)) and D(accuracy) is the largest of the six D(s). The registered prediction
and the format gate can no longer diverge.
A6 (finding 6) — two predictions tightened and one reframed. Prediction 7 now requires the between-item spread to exceed the operative floor on at least three of six senses, not one. Prediction 4 is restated per juror: no juror's mean for a primary item falls below 4.0 on any sense, with the juror offsets reported beside it. Prediction 1's band is left as registered and is labelled in the result as the weak prediction the critic says it is.
A7 (finding 7) — F2 governs the null's status, and prediction 2 is restated in F2's terms.
Prediction 2 now reads: P(s, j) ≤ 2 × R(s, j) on at least 5 of 6 senses for every juror. The
aggregate |P − R| ≤ 0.50 form is withdrawn.
A8 (finding 8) — the edit list gets an independent neutrality check before dispatch. A new stage
edits, dispatched before any rating call: seat P3 is given the ten edits in context and the
six sense definitions, and asked, for each edit, which senses it would move and in which direction.
P3 has not seen the design, the predictions or the critic. Its verdict is recorded whatever it says
and is reported beside the measured P. It is a check, not a gate: the edit list is frozen and
is not revised on it, because revising the null after an opinion about it is the move this project's
discipline exists to stop.
A9 (finding 9) — absence of a comparator is not absence of contamination. CLO is admitted because
no English rendering exists to collide with; that says nothing about whether the lead carries the
French source from training. Recorded in §8 and on T-clos-des-ames-R04-v1. The word
"uncontaminated" is not used of CLO anywhere.
A10 (finding 10) — R is computed over the five primary items only. The retest differences of
CLO, CLO-P and BAR-D are reported separately, since damage and paraphrase may interact with pass.
A11 (finding 11) — KUS's fragment status is a per-item covariate. KUS is the first 3 of 12
paragraphs; every other item is a whole filed span. No KUS-versus-whole-item comparison is made on
affect or style-correspondence, both of which are properties of a completed unit.