Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260730c-revision-close/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260730c-revision-close
statusfrozen
created2026-07-30
updated2026-07-30
sensesaccuracy, naturalness, style-correspondence
provisionaltrue
internal-judgment-onlytrue
linkswiki/arms/ARM-revision.md, workshop/experiments/E-20260729e-revision-pass/design.md, wiki/findings/results/RS-20260729e-revision-pass.md, wiki/findings/results/RS-20260730-grain-clause.md, workshop/experiments/E-20260730b-c16-redraw/design.md, workshop/translations/niewola-tatarska/R04-v1/translation.md, config/models.md, config/budget.md, wiki/method-notes.md

E-20260730c — close ARM-revision: rebuild the control that failed, and find out whether the instrument holds still long enough for the rebuild to mean anything

Frozen before the independent pre-run critic pass and before any reader call. The translation limb was frozen earlier and separately (T-niewola-tatarska-R06-v1 at 14c8989, T-niewola-tatarska-R04-v1 and its revision log at 73ba25c, contamination figures at 1bf2831). Amendments made after the critic pass are recorded in §11 with the finding that caused each; nothing else in this file changes after the freeze commit.

1. What this is

ARM-revision step 2, the arm's last, three items in the order the arm names them:

And one thing the arm did not ask for, which the previous session made unavoidable. RS-20260730-grain-clause (S061) found that a reader instrument of the same family returned κ 0.452 → 0.737/0.808 on a byte-identical request one day apart. S057's F3 diagnosis and this session's rebuild are separated by exactly that gap. So the rebuild carries its own cross-day control, and the control is a gate on attribution rather than a curiosity: see F5.

2. The wire between the limbs, in one sentence

The Polish pair was translated as translation, in session, and it is the first prospective instance of item (b)'s statistic — draft frozen, revision frozen, comparator never opened, run measured afterwards — where all five other pairs in that measurement are retrospective.

3. Materials

The reader batch is S057's batch with FIVE items substituted in place. build_items.py copies S057's frozen materials/items.json (221 items: 205 real edits over nine R06/R04 pairs, 8 meaning controls, 8 surface controls), preserves every key and every slot, and replaces the four text fields of five of the eight surface controls. No reshuffle, no reseed. 216 of 221 items are byte-identical at byte-identical positions — asserted by build_items.py rather than described — and that is what makes F5 measurable. (As frozen this said eight substituted and 213 identical; the redesign of §5 reduced the change. The lower number is what ran.)

Item (b)'s corpus, and n is stated here before any threshold is set — which is the failure A4 withdrew P5 for. Mechanical inclusion rule: a pair enters if (i) its R06 and R04 prose are extractable and (ii) the artifact already records a comparator reachable at a Project Gutenberg ebook id. That rule gives n = 6:

pair source language comparator Gutenberg
bargamot Russian (as recorded on the artifact) #49598
enfermeiro Portuguese Goldberg 1921 #21040
kiseru Japanese Shaw 1930 #78105
kusamakura-vii-bath Japanese Takahashi 1927 #73131
wang-liulang Chinese (classical) Giles #43628
niewola-tatarska Polish Curtin 1898 #36583

Excluded, each with the reason, because a silently truncated corpus reads as a complete one: bettelweib-locarno — its artifact states no free comparator exists; alfred-preface and saigo-no-ikku — contamination UNMEASURED, no comparator recorded; takasebune, yingyi-jiejixing, senilia-clean — comparators exist and are reachable but not at a Gutenberg id, so they fail clause (ii) and are excluded for uniformity of retrieval rather than for unavailability; malory-worship — intralingual, and #1252 is its source, not a comparator translation. Six of thirteen candidate pairs enter. That is the number A4 predicted and it is arrived at by rule, not by choosing.

4. What is not re-run, stated so no borrowed number sneaks in

The log-versus-diff primary of RS-20260729e §3 (42 of 47) is untouched and is not recomputed here. The edit list is not rebuilt. The Polish pair is not added to the reader batch — adding it would break the byte-identity F5 depends on — so it contributes to item (b) and to nothing else.

5. The surface controls — WITHDRAWN AND REBUILT BLIND after the critic pass

§5 as frozen described eight surface controls the lead had built by hand, under a criterion ("match the corpus on mechanical kind") derived from S057's failure scores. The pre-run critic returned BLOCKING finding 2 against exactly that, and it is right: controls built by an agent that has seen the failed threshold are selection-confounded by construction, and no amendment repairs that after the fact. The eight were WITHDRAWN UNSCORED. They remain in git history at 913f5f7. Full record: critic/dispositions.md.

What runs instead. The surface band is split into two declared bands at the same eight slots, with no reshuffle and no change in batch size:

F3's surface mean is computed on band L only. Band N supplies the within-run contrast gate F7.

The diagnosis, now measured rather than asserted (critic finding 3). §5 as frozen said the rebuild would "match the corpus on mechanical KIND"; the critic showed that claim was false of the lead's own materials. The replacement criterion was computed before the blind build:

Does the revision span introduce a word type absent from the draft span? 184 of the 205 real edits do — 89.8%. Four of S057's eight surface controls do, and all four are two contractions and two misspellings — the same word in another form. Not one of S057's eight surface controls substitutes a different English word for a word. That is the operation nine tenths of the corpus consists of, and its total absence from the control set is a sharper statement of note (bdu) than "too weak".

Span, realised blind (build_items.py), against the target the builder was given:

draft-side span, median / mean revision-side, median / mean
the 205 real edits 2.0 / 2.717 2.0 / 3.195
S057's surface controls 2.5 / 4.500 2.0 / 4.125
the lead's withdrawn controls 3.0 / 3.750 3.0 / 3.750
band L, built blind 2.0 / 2.800 2.0 / 3.200

Matched to within 0.083 and 0.005 of a word, by a builder that could not see a threshold.

The meaning controls are byte-identical to S057's and are not touched. With band N kept, 216 of the 221 items are byte-identical at byte-identical slots.

The residual confound, declared: the builder's prompt was written by a lead that had seen S057's scores. The builder's ignorance is of the scores and thresholds, not of the framing.

6. Procedure

  1. Freeze this file, build_items.py and materials/items.json. (This commit.)
  2. Independent pre-run critic pass — P4, the only seat that is neither a reader nor the declared reserve. Findings accepted or declined in writing before any reader call.
  3. Reader pass: 3 seats × 2 axes = 6 calls, prompts and parameters unchanged from E-20260729e/run_readers.py except the eight substituted items.
  4. Item (b): re-fetch the six comparators, locate each by heading offset, measure tools/dependence_check.py on (R06 vs comparator) and (R04 vs comparator). No comparator prose is displayed at any point (method note abm).
  5. Analysis, then independent verification recomputing every reported number from the stored raw bodies.

7. The axes and the prompt

Unchanged from E-20260729e §7, including the limitation stated there: readers do not see the source, so M measures change in what the English asserts, not correctness against the source.

8. Registered gates

9. Predictions, registered

10. Panel bindings and pre-flight

role seat why
pre-run critic P4 moonshotai/kimi-k3 neither a reader nor the declared reserve; the S053 role-collision fix, tenth session running
readers ×3 P1, P3, P5 S057's seats exactly — F5 requires the same seats or it measures nothing
surface-control builder qwen/qwen3.7-max added by amendment A2. Probed-but-not-selected in config/models.md; not a reader, not the readers' reserve, not the critic. A materials-construction role, not a judging role
declared reserve P2 as S057, declared here and not chosen later (notes (b), (bdl))

Pre-flight, built from max_tokens and not from an assumed output length (note abc), at the corrected P1 price:

Worst case ≈ $1.05. Every output-dominated run in this ledger has landed at 15–34% of worst case. Headroom at this freeze: $4.3973500561 (today's $5.00 less S061's $0.5664084814 and this session's gate at $0.0362414625).

Free, and never ledgered: the Polish translation and both its logs, all six contamination measurements, the item rebuild, every analysis and the verifier.

11. Amendments after the pre-run critic pass

moonshotai/kimi-k3 (P4), one call, stop, in 5,398 / out 4,771, provider Together, $0.0876726. Verdict NEEDS-REDESIGN — two BLOCKING, four MANDATORY, one ADVISORY. All seven accepted, none declined. Note (rr), twentieth consecutive session, and the first NEEDS-REDESIGN this project has taken since S021.

Amendments A1–A7, each with the finding that caused it and what it cost, are in critic/dispositions.md. In summary: A1 tightened F5 from ≤10/≥0.80 to ≤3/≥0.90 and made attribution a comparison against the measured drift rather than a threshold; A2 withdrew the lead's eight hand-built surface controls unscored and replaced them with five built blind by a seat that saw no score; A3 replaced the "match on mechanical kind" criterion with the measured new-word-type criterion (89.8% of real edits, 0 of 8 S057 surface controls); A4 matched span blind to within 0.083 of a word; A5 made F5's failure two-way confounded in writing; A6 registered that a failed F5 makes P4 within-session-only and item (a) unanswerable; A7 added band N and gate F7.

A design that loses its own control set to its own critic is the mechanism working. Nothing else in this file changed, and no reader call had been dispatched when these were applied.

12. What this design cannot establish