Repository path: workshop/experiments/E-20260802c-regime-scoring/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260802c-regime-scoring |
| status | frozen |
| created | 2026-08-02 |
| updated | 2026-08-02 |
| senses | accuracy, naturalness, voice, style-correspondence, cultural-mediation, affect |
| purpose | Readers of literary fiction in English who cannot read the source, meeting these texts as reading editions rather than as cribs. Declared because every evaluation must state its purpose parameter (D-20260801-10). |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-first-judgment.md, framework/control-arm-spec.md, framework/closure.md, config/models.md, wiki/goodness-senses.md, workshop/regimes/R04-lead-close.md, workshop/regimes/R06-lead-single-pass.md, workshop/translations/mort-de-pan/R06-v1/translation.md, workshop/translations/mort-de-pan/R04-v1/translation.md, wiki/findings/results/RS-20260802-tierD-verdict.md |
E-20260802c — what does one self-revision pass actually buy, per sense, on this project's own translations?
Frozen 2026-08-02 (S089) after both arms of every pair were frozen and before any call was
dispatched. AMENDED before the first scoring call by the independent pre-run critic pass —
critic.md, verdict NEEDS-AMENDMENT, ten findings, six BLOCKING, all ten accepted (amendments
A1–A9). Every amendment is marked in place below and materials/items.json was rebuilt; its
SHA-256 is now b4c25b5c847895655bc422bdd352c6ac653fdf2f4f88d8ec4b65b0cea494d5b7. ARM-first-judgment step 1, study limb. Translation limb: T-mort-de-pan-R06-v1 and
T-mort-de-pan-R04-v1, both frozen at commits 67a700b and its successor before this file
existed (charter §3, A4).
0. Standing, stated first because it governs every number below
Tier D is NOT PASSED (RS-20260802-tierD-verdict, S086; config/models.md). No score this
design produces carries evidential weight, no result of it may support a framework
recommendation, and every artifact it touches stays provisional. That is the condition
ARM-first-judgment pre-committed to running under, and it is why the arm was unblocked by a
failure rather than by a pass.
What the Tier D run does license and this design uses: the panel separates a reference from an
8-site and a 3-site accuracy-damaged variant at ceiling, in every juror taken separately. It is
not at chance on this item format. That bounds nothing about ranking two competent translations,
which is exactly what this design asks — and the distance between those two things is the reason
everything here is provisional.
1. The question, and why it is not a question about the apparatus
What does a translator's own second pass over their own draft actually change, and on which of the ways a translation can be good?
framework/traceability-inventory.md's candidate C12 — one self-revision pass buys
naturalness without moving accuracy — is the only prescriptive statement about translating
this project has ever produced, and it is inadmissible (framework/closure.md §2). It came from
S010, an API-regime comparison (R01/R02) on an uncalibrated jury with a control whose inference was
withdrawn. It has never been tested on a human-shaped translator, on more than one language pair,
or on a translation anybody read twice.
Meanwhile fifty-seven filed translations have never been judged for quality by anything
(wiki/reassessment-2026-08-01.md §1.3). The project has an opinion about what revision buys and no
evidence about what any of its own prose is worth.
The subject-rule sentence (wiki/tracks.md; continue-prompt.md §4.5), narrowed by
amendment A8: this unit teaches what this translator's second pass changed, on these five
pairs, as reported by this uncalibrated jury — descriptively and provisionally — and whether
the one prescriptive claim the project has ever made about revision survives contact with any
evidence at all. C12 is treated here as an inadmissible prior being stress-checked, not as a
hypothesis this design could accept. That is still a claim about the practice of translating, not
about this project's instruments; the instruments are the means.
The wire between the limbs, in one sentence: the five pairs that carry every primary number are pairs whose revision was made by a translator with no reason to think anyone would ever score it, and the sixth — «La Mort de Pan», translated and revised in this session — is the first R04 revision this project has made knowing the pair would be scored, which is why the critic's finding 5 put it outside every primary aggregate and why it is still reported: it is the one cell that shows what scoring-awareness does to a revision.
2. Materials
Every arm is a frozen artifact sliced by exact section marker by materials/build.py, which
asserts each slice against the word count its own artifact states and refuses to emit a manifest if
any slice is short, empty, or restructured. materials/items.json carries SHA-256
b4c25b5c847895655bc422bdd352c6ac653fdf2f4f88d8ec4b65b0cea494d5b7 (post-amendment; the pre-critic
manifest was a3535f7a…), recorded by the runner in every raw response file.
| item | work | pair | draft w | revision w | Δ | stratum |
|---|---|---|---|---|---|---|
| TAK-A | Ōgai, 高瀬舟 (1916), span A | JA→EN | 513 | 517 | +4 | pos |
| TAK-B | Ōgai, 高瀬舟, span B | JA→EN | 382 | 388 | +6 | pos |
| KUS | Sōseki, 草枕 ch. VII, the bath | JA→EN | 1210 | 1206 | −4 | neg |
| BAR-A | Andreyev, «Баргамот и Гараська» (1898), unit A | RU→EN | 236 | 247 | +11 | pos |
| BAR-B | Andreyev, «Баргамот и Гараська», unit B | RU→EN | 437 | 441 | +4 | pos |
| PAN | Arène, «La Mort de Pan» (1876) — secondary, A4 | FR→EN | 478 | 477 | −1 | neg |
| CTRL-NULL-TAK | 高瀬舟 span B revision against itself | — | 388 | 388 | 0 | null |
| CTRL-NULL-KIS | Akutagawa 「煙管」 §一 R04 against itself (held out, A1) | — | 473 | 473 | 0 | null |
| CTRL-NULL-SVI | Turgenev «Свидание» opening R04 against itself (held out, A1) | — | 644 | 644 | 0 | null |
| CTRL-POS | Andreyev unit B revision vs variant-F1 |
RU→EN | 441 | 461 | +20 | control |
Amendment A4 — the primary set is FIVE items, not six. PAN is the one pair whose revision was
made in knowledge of this design, and the critic's finding 5 is accepted in full: it is excluded from
every primary aggregate and from predictions 1, 2, 5 and 7, and reported as a labelled secondary
cell. The negative length stratum therefore has n = 1 (KUS), which under control-arm-spec R3
yields its raw value and no summary at all.
Amendment A5 — CTRL-POS's stimulus is specified, not named.
workshop/translations/bargamot/R04-v1/variant-F1.txt, spec in variant-F1.md: eight sites, two
each of wrong referent, wrong word sense, invented detail, dropped negation, net +20 words
(+4.5%). Its SHA-256 is carried in the manifest. That +4.5% is a length confound and it is
declared: a fire on this control is not attributable to accuracy alone (A2).
Two of the eighteen banked R06/R04 pairs per work, and not more, and here is the rule. Only
language pairs on which the panel's competence has actually been screened are admitted — Russian,
French and Japanese (config/models.md, S015: a six-item screen, ≥5/6 for every juror). That rule
excludes the banked German (bettelweib-locarno), Chinese (wang-liulang, yingyi-jiejixing) and
classical-Chinese (cuzhi) pairs, which would otherwise have doubled the item count. It is a real
cost and it is paid deliberately: an accuracy judgment from a juror who cannot read the source is
not a weak measurement, it is a different one. svidanie is excluded from the scored items separately, for
contamination: high (ARM-first-judgment §Constraints) — and is used as a null control, where
that exclusion does not bite: a null presents one text against itself, so nothing about its
independence from a published rendering can affect the floor it measures.
The length property this design has and S010 did not. On the five primary items, mean |Δ| is
5.8 words, 1.65%, and Metric A = 0.800 (4 pos, 1 neg; chance 0.500; the six-item figures,
0.667 and 5.0 words, are stored beside them; materials/freeze-stats.json, computed at freeze time
per control-arm-spec R5). S010's TEMP arm spanned 1 to 151 words and its jury's preference tracked
|Δ| at ρ = +0.470. These arms are nearly length-matched by accident of the regime — a
self-revision is not a rewrite — so the confound that made S010 unreadable is small here by
construction rather than by matching, which control-arm-spec F1 shows produces a third text.
2.1 The fresh pair, and the two works that did not survive their own gate
The contamination gate of note (bcd) was run before the locus was chosen and before anything was translated, on a discard rule frozen in advance: longest common run ≥ 12 tokens, or ≥ 1 shared 12-gram, and the work is discarded as a limb.
| candidate | comparator (whole story, never displayed) | longest run | 12-grams | nulls | verdict |
|---|---|---|---|---|---|
| Maupassant, «Menuet» | anon. English, Original Short Stories vol. 10 (PG #3086) | 14 | 3 | 3 / 3 / 2 tokens, 0 twelve-grams | DISCARDED |
| Villiers, «La Torture par l'espérance» | anon. English, Riddle Stories (PG #29704) | 12 | 1 | 3 / 2 / 2, 0 | DISCARDED |
| Arène, «La Mort de Pan» | none exists | — | — | — | admitted |
Both gate translations are filed and both were frozen before their measurement ran. The rule
was not relaxed after watching it fire twice, which is the move this project's verification
discipline exists to stop. The third work was admitted on a different basis: a search of Project
Gutenberg (six Arène volumes, all French) and the Internet Archive (creator:"Arène, Paul", 16
records, 15 French and 1 German) found no English rendering of any work by Paul Arène, so there
is no published English for the translator to have remembered. That is an argument from absence and
is weaker than a number, and the artifacts say so.
One exposure, declared. While debugging the second gate's heading extraction, roughly twenty words of the «La Torture par l'espérance» comparator's opening were printed to the lead's console. That work had already been gated and discarded; the exposed text belongs to the gate span, whose translation was frozen before the exposure, so the 12-token figure above is a genuine pre-exposure measurement. Nothing of «La Mort de Pan» was exposed at any point. Recorded because a contamination event that only the operator can see is exactly the kind that does not get recorded.
3. Procedure
- Jurors P1, P2, P5 (
openai/gpt-5.6-terra,google/gemini-3.6-flash,deepseek/deepseek-v4-pro), the same three as S020, S034 and S086, so this extends the instrument rather than replacing it. Non-Anthropic (charter §5). The lead never judges its own translation, and every translation here is the lead's. - Blind and authorship-stripped. The juror sees the source passage and two English translations labelled A and B. It is not told that one is a draft and one a revision, that they are related at all beyond sharing a source, who made either, or in what order they were produced.
- Order-swapped. Every (item, juror) is dispatched twice, slots swapped. The reported delta for a cell is the mean over the two orders, which cancels slot preference by construction.
- Strictly sequential. Judgment is never parallelized (charter §6); the runner enforces it.
- Six senses, from
wiki/goodness-senses.mdas it stands after S085 and S088:accuracy,naturalness,voice,style-correspondence,cultural-mediation,affect.literary-qualityis retired andpurpose-fitis a parameter, so neither is a legal id.consistencyis deliberately omitted: its own entry makes its weight a function of length, and these are 236- to 1210-word spans. Sense wordings are carried verbatim fromtools/run_tierD.pywhere they exist there, so the same words go to the same jury as in the Tier D runs. - Runner:
run_scoring.py, adapted fromtools/run_tierD.pywith the sense list changed and nothing else.tools/run_tierD.pyis left untouched so the Tier D instrument stays reproducible. - Scale 1–7 per sense per text, plus a forced overall preference (no ties).
4. Analysis, fixed before the run
Let d(i, s, j) = mean over the two orders of score(revision) − score(draft) for item i, sense
s, juror j. Positive means the revision scored higher.
- Per item and per sense,
d(i, s)= mean over the three jurors. Reported as a table, all 36 real cells, with no aggregation hidden. - The noise floor is
N= max over senses and over all three nulls of|d(CTRL-NULL-*, s)|(amendment A1). Per-null floors are reported separately, because a floor that differs by work, language or length is itself the finding. Every delta is reported againstNand no delta belowNis described as a movement. A1's mechanism correction is on the record: the critic argued that a null sharing prose with a scored item invites cross-payload memory; every call here is a stateless request with no shared context, so that mechanism does not exist, and the amendment was made on the finding's other and stronger ground — generalisation. control-arm-specR1 (ratified as amended 2026-08-01): every headline aggregate is split by length-sign stratum — pos = {TAK-A, TAK-B, BAR-A, BAR-B}, neg = {KUS} — with per-stratum item count and per-stratum raw aggregate. R3: both strata have fewer than 5 items, so no intervals are computed and none is reported; raw per-item values only, and the neg stratum is a single value. R4: this design has 5 primary items against R4's 22–30 per stratum, claims no stratified estimate, and uses the split as a diagnostic. F3: the within-stratum magnitude check is computed and reported as description only, at n = 4 and n = 1. 3a. Amendment A9, the mandatory nuisance check. The rank correlation between per-itemd(naturalness)and per-itemΔwordsover the five primary items is computed and reported whatever it shows, as description only at n = 5. Four of five revisions are longer, so anaturalnessgain that tracks length is a live alternative to one that tracks revision, and at this n the design cannot separate them.control-arm-specR5: Metric A and the sign test are recorded at freeze time (§2) and are not recomputed after the scores are seen.- Slot preference is measured per juror over its 20 calls and reported alongside S014's and S020's figures (0.500–0.600).
- Amendment A6 — forbidden wordings, binding on the result page. Prediction 2's outcome may be written only as "no accuracy movement detectable above the null floor", never as "accuracy unchanged". The phrases "confirms C12", "supports C12" and "C12 holds" may not appear: C12 is inadmissible and this design cannot make it admissible. Prediction 7's ceiling is a descriptive bound, not a licence to set a large number aside as jury pathology.
- Amendment A3 — identity detection in a null. A null cell whose
whystates or implies that the two texts are the same is retained, and the count of such cells is reported per null and per juror. A juror that notices identity and returns equal scores is returning the floor correctly; excluding those cells would select for the jurors that failed to notice.
No pooled figure is the headline of anything. Where a pooled number appears it appears beside both strata, per R1.
5. Predictions, registered so that failures are on the record
Predictions 1, 2, 5 and 7 are over the five primary items, PAN excluded (A4).
naturalnessmoves up. Meand(naturalness)over the five primary items is > 0 and > N.- No
accuracymovement is detectable above the floor.|mean d(accuracy)| ≤ Nwhile (1) holds. Together (1) and (2) would reproduce C12's pattern — descriptively, on a jury with no authority, and without making C12 admissible (A6). - CTRL-POS fires:
d(CTRL-POS, accuracy) ≤ −1.00and the forced preference goes to the undamaged text in ≥ 5 of 6 unit-orders. A fire establishes that the panel separates grossly damaged prose from clean prose on this material today — nothing about near-peer discrimination, and not attributable toaccuracyalone, since the damaged arm is +4.5% longer (A2). - The nulls are quiet:
N ≤ 0.50scale points across all three. - Descriptive only (A7). The revision is preferred overall in more than half of the 30 primary calls. Reported per juror; not a headline.
- Secondary observation, not a prediction (A4).
d(PAN, naturalness)is smaller than the meand(naturalness)over the primary items — a revision has less to repair where the languages are structurally close. Reported with PAN's confound named beside it. - Nothing moves far. No sense's mean
|d|over the five primary items exceeds 1.00 scale points. These are two drafts of one text by one translator, and a jury that reports a large gap is telling us about itself.
6. Failure criteria — what voids what
| id | condition | consequence |
|---|---|---|
| F1 | N > 0.75 — the panel reports a larger difference between a text and itself than three quarters of a scale point |
No directional claim is made from any item. The instrument's resolution is coarser than any effect this design could see; the run reports the floor and the controls and nothing else |
| F2 | CTRL-POS does not fire (prediction 3 fails on either clause) | The panel is not separating gross accuracy damage on this material today, notwithstanding S086 (wording narrowed by A2 — a fire would not have licensed "the instrument works", so a failure does not license "the instrument is broken"). Every real-item number becomes descriptive only and is labelled so on the result page |
| F3 | a juror's slot-A preference rate falls outside [0.30, 0.70] over its 20 calls (band tightened by A7) | That juror's forced-preference data is excluded; its per-sense scores are retained, because the order-swap mean already cancels slot preference in them. If all three fall outside, no preference result is reported at all |
| F4 | more than 5 of 60 calls fail to parse after one retry | the cell structure is holed; report the achieved cells and claim nothing about any sense with a missing juror |
| F5 | any call bills above its declared per-call worst case | not a budget event but a routing finding (note (bgk), fired twice already); the runner aborts the stage and it is reported |
F1 is the criterion this design is most likely to fail, and failing it is the informative outcome. A jury that cannot tell a text from itself to better than a point is a jury that cannot be asked what a self-revision buys, and that fact would be worth more than any of the six deltas.
7. Budget
Worst case built from max_tokens, not from an expected output length (note (abc)), and
inflated by 1.3× on the output leg because note (bgk) has twice recorded a call billing above its
max_tokens.
| juror | max_tokens | worst in | worst out (×1.3) | worst $/call |
|---|---|---|---|---|
P1 openai/gpt-5.6-terra |
3000 | 6000 tok @ $1.25/M | 3900 @ $7.50/M | $0.03675 |
P2 google/gemini-3.6-flash |
4000 | 6000 @ $1.50/M | 5200 @ $7.50/M | $0.04800 |
P5 deepseek/deepseek-v4-pro |
6000 | 6000 @ $1.65/M † | 7800 @ $3.30/M † | $0.03564 |
† P5 priced at the worst plausible provider, not its list price: OpenRouter routing has billed
this slug at 3.8× list (config/models.md, S022 measurement).
20 payloads × 3 jurors = 60 calls → $2.408, plus a 10% retry allowance at the dearest seat
(6 × $0.048 = $0.288) = full-stage reservation $2.696, against today's headroom after the critic
call of $3.5402 (UTC 2026-08-02; $1.3975 spent by S086–S088, plus this design's critic call at
$0.0623044). Reserved in full before the first call by --reserve, or the stage is not entered.
The call count rose from 48 to 60 with amendment A1's two extra nulls; the reservation was recomputed before dispatch and still fits. The critic call itself came in at $0.0623 against the $0.08 this section originally budgeted — an overrun of the stated line, recorded rather than quietly re-based; the figure above is the actual.
8. What is NOT being measured
- Not whether these translations are good. Tier D is not passed; §0 governs. What is measured is a difference between two arms, under a jury with no established authority to rank near-peers.
- Not whether R04 is a better regime than R06. Five pairs by one translator is not a regime
comparison in the sense
framework/closure.md§3 means; it is the first evidence of any kind. - Not the second limb of the release gate.
ARM-framework-v01staysblockedon the gate's first limb, which the Tier D run answered against. A scored comparison on an uncalibrated jury does not open a release, and this design makes no claim that it does. - Not the A4 promise. That is
ARM-first-judgmentstep 2 — a cross-language set of filed translations judged blind on their own terms, not as pair members.