Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260730h-strict-coverage/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260730h-strict-coverage
statusfrozen
created2026-07-30
updated2026-07-30
trackT2
sensesnaturalness, style-correspondence, cultural-mediation, accuracy
provisionaltrue
linkswiki/arms/ARM-rule-coverage.md, wiki/findings/results/RS-20260730e-rule-coverage.md, workshop/regimes/R07-fluency.md, workshop/translations/osso-di-morto/R07-v1/translation.md, workshop/translations/postmaster/R07-v1/translation.md, workshop/translations/petits-poemes/R07-v1/translation.md, config/models.md, config/budget.md

E-20260730h — the strict test, applied backwards to 39 published D codes and forwards at the moment of decision

ARM-rule-coverage step 2. Frozen before the Italian and Bengali logs were opened and before the new source was read for translation. Session S067.


0. The wire, in one sentence

The study limb re-derives all 39 published R07 D codes against §5's exactly one survivor test; the translation limb is a fourth R07 run in which the live options are enumerated and both tests are applied at the moment of each decision — so the two limbs together say whether the defect S064 found is in the coding or in the option-recording the coding was done from.

1. The question

R07 §5 defines D as: a rule names the feature at issue and only one of the live options satisfies it. RS-20260730e §2 established, on three sites, that a D has in practice been recorded whenever a rule excluded an option — a weaker condition — and that at two of the three sites the lead was wrong and three independent readers were right.

That is a demonstration on three items chosen by the lead. It is not an estimate. Three published coverage rates — 13.6% (Italian), 36.1% (French), 66.7% (Bengali) — stand as upper bounds of unknown tightness, under a binding condition written into R07 §Measured defects.

Q1 (retrospective). Of the 39 published D codes, how many leave exactly one live option standing?

Q2 (prospective). On fresh material, coded under both tests at the moment of decision by a translator who knows the distinction, how large is the gap between the exclusion-test rate and the strict-test rate?

Q3 (reachability). Is the strict test answerable by a reader who did not make the choice — when it is decomposed into per-option judgments rather than asked as a three-way classification? RS-20260730e §5 measured the three-way version at 0.757 self-agreement on a byte-identical same-day prompt, which is why the arm's condition 2 is still open.


2. Priming, declared before anything else

The French log was read in full during this session's orientation, before this design was written. All 36 sites of T-petits-poemes-R07-v1, with their option sets, chosen renderings, codes and cited rules, were displayed. So for the French run the recount procedure below is not blind to the material it will be applied to, and the 13 French D codes are declared primed.

The Italian and Bengali logs have not been opened. Their 26 D codes are the blind stratum. T-osso-di-morto-R07-v1 and T-postmaster-R07-v1 are opened only after this file is committed.

Consequence, registered: every headline figure is reported twice — over all 39, and over the blind 26 alone. If the two disagree in direction, the blind 26 govern.

Also declared: the lead has read RS-20260730e §2 and therefore knows the correct strict-test answer at three sites (I23 «খোল-করতাল» → P, I27 «les gargoulettes» → D, I34 «শ্মশান» → P). Those three sites are excluded from the primary statistic and reported separately.


3. The procedure — frozen, and mechanical by construction

Applied to each of the 39 published D codes, in the order the logs print them.

  1. Read the site's live options cell verbatim. The option set is whatever the log records. It is not extended, corrected or supplemented. Where the log records an option as a description rather than a string (e.g. "keep the figure"), it counts as one option.
  2. Read the site's rules cell and take the set of rule numbers it cites. Rule text is R07 §The rule set, verbatim, v1.0 frozen.
  3. For each option, for each cited rule, ask: does this option satisfy the rule as written? Three values: SATISFIES, VIOLATES, NOT-BORNE (the rule does not name the feature this option turns on). NOT-BORNE counts as satisfying — a rule that does not bear cannot exclude.
  4. Survivors = options that are not VIOLATES on any cited rule.
  5. Strict code: D iff survivors == 1. P iff survivors ≥ 2. ANOMALY iff survivors == 0 (the chosen rendering is always among the options, so a zero means the log's own reasoning is internally inconsistent, and it is reported as such rather than silently coded).
  6. A rule cited as bearing-but-not-excluding is still cited. The logs distinguish these in prose; the procedure does not read the prose. It reads the option strings and the rule numbers.
  7. The chosen rendering's identity is not used. Which option the translator picked is not an input to steps 3–5.

What this procedure cannot do, stated now. It cannot recover options the translator never wrote down. A site whose log records two options where three were live will be recounted on two. That is exactly the limitation Q2 exists to price, and it is why the translation limb enumerates prospectively.


4. Conditions

C1 — the lead's recount (free, no call)

The procedure of §3, applied by the lead to all 39 sites. Output: analysis/recount.json, one row per site with option-level judgments, survivor count, strict code, and a one-clause reason per VIOLATES.

C1 runs AFTER the panel prompts are built and dispatched (§4.2). The item set is frozen and sent before the lead's own answers exist, so the item construction cannot be fitted to the recount.

C2 — the per-option panel instrument (the rebuild RS-20260730e §9 item 2 asks for)

The task is not the S064 task. Raters are never shown the labels D, P or S, are never told what a coverage rate is, and are never asked to classify a site. They are asked, per option, whether it satisfies a verbatim rule. The code is derived arithmetically afterwards by §3 steps 4–5.

C3 — the synthetic controls, built in BOTH directions

RS-20260730e §2's control failed because it was mis-built: it forced items where the lead believed the answer was D and the correct answer was P. The fix is controls whose answers are forced by the rule text and which point in both directions.

id rule options forced answer why it is forced
X1 F5, American declared "harbour" · "harbor" 1 survivor F5 forbids mixing; one form is the declared dialect
X2 F4 «samovar» untranslated · "tea urn" 1 survivor F4 names untranslated source-language words in so many words
X3 F1 "big" · "large" 2 survivors neither is dated; F1 excludes neither
X4 F3 "very tired" · "extremely tired" 2 survivors neither is slang or dialect; F3 excludes neither

X1, X2 force exactly one survivor. X3, X4 force two. A rater that cannot separate these two directions cannot do the task at all.

C4 — the translation limb (free, no call)

A fourth R07 run, on material selected after this file is committed, under the regime unchanged and frozen. Two things distinguish it from the three existing runs and both are the point:

  1. Live options are enumerated exhaustively at each contested site, before any code is assigned — every rendering the translator was actually willing to write, not only the ones the deciding rule bore on.
  2. Two codes are recorded at every site: code_excl (a cited rule excludes ≥ 1 live option) and code_strict (exactly one live option survives all cited rules). The translator applies both tests in the same breath.

Contamination gate per the standing rule: run on the first unit before the rest is drafted, and the passage is not selected until the gate has been read.

Material selection criteria, frozen here so the passage is not chosen after seeing what it would give. In this order, and the first candidate meeting all four is taken:

  1. Realia-dense, so F4 fires repeatedly — because §2's demonstration is that the exclusion/survivor gap opens exactly where a culture-bound item has several acceptable English renderings. Choosing culturally-close material would test the strict test where it does not bite.
  2. Source and a published English translation both freely reachable, so contamination is a measurement and not a placeholder (CLAUDE.md, standing rule 2026-07-26).
  3. A language already used in this project, so the run is comparable; breadth is not this arm's question.
  4. Not a passage any earlier lead translation has touched.

Registered before selection: the target length is 700–1,000 source words, and the passage is translated whole — no site is dropped for being awkward to code.


5. Registered predictions

Frozen before the Italian and Bengali logs are opened, before the new source is read, and before any call is dispatched.

# prediction fails if
P1 Fewer than 25 of the 39 published D codes survive the strict test (≥ 36% fail). ≥ 25 survive
P2 The Bengali run loses the largest share of its D codes and the Italian run the smallest. any other ordering
P3 The per-option instrument's byte-identical repeat agreement (option-level) is ≥ 0.85, against the three-way task's 0.757. < 0.85
P4 Panel-derived strict codes agree with the lead's strict recount on the 36 scored items at ≥ 0.75. < 0.75
P5 On the new run, code_excl rate − code_strict rate ≥ 15 percentage points. < 15 pp
P6 Mean live options per contested site on the new run is ≥ the mean of the three published runs. < that mean

P6 is the discriminating prediction and is stated as such. If P6 holds and P5 holds, the defect is in both the coding and the logging, and the retrospective recount of §4.1 is a lower bound on how many D codes are wrong. If P6 fails — the new run enumerates no more options than the old ones did — then the logs were not under-recording and the recount is an estimate rather than a bound.

Registered no-effort nulls, on the pre-run critic's standing lesson from S066 finding 7.


6. Failure and void criteria

F1 — control failure. A rater scoring < 3 of 4 on X1–X4 is excluded; its data is published as description carrying no weight. If fewer than two raters survive F1, C2 is VOID and no agreement figure is reported as a measurement.

F2 — repeat floor. If C2R option-level agreement is < 0.85, no C2 agreement figure is a measurement of anything, exactly as RS-20260730e applied its own Q8′. The C2 codes are then published as description only and the arm's condition 2 closes retired on the ground that the task is not reachable.

F3 — degenerate marginals. If any surviving rater's VIOLATES share is < 15% or its SATISFIES share is < 15% across its 40 items, that rater is reported as degenerate and excluded from P3 and P4.

F4 — anomaly rate. If ANOMALY (zero survivors) exceeds 5 of 39 in C1, the procedure of §3 is mis-specified rather than the codes being wrong, and §3 — not the logs — is what the result page reports on.

F5 — parse. Any rater body that does not yield 40 complete rows is re-issued once to the declared reserve; a second failure drops the seat, and the seat count is reported.


7. Pre-run critic

One call, P4 moonshotai/kimi-k3, the standing critic seat (S053 role-collision fix). P4 is not a rater and is not the raters' reserve. Declared reserve for the critic seat: P2. All findings are recorded in amendments.md with accept/reject and reason, before any rater call is dispatched. Note (rr).

8. Pre-flight cost estimate

Built from max_tokens, not from an assumed output length — note (abc) — and built from the call count that can actually occur, which is where S066 overran at 459%.

stage calls priced max_tokens worst case
critic 1 + 1 reserve 12,000 $0.36 (kimi-k3 out $15/M; 1.5× provider premium; ~11k prompt)
C2 raters 3 + 3 reserves 8,000 $0.42 (worst seat P1 at $7.50/M out; 6 × 8k out + 6 × ~9k in)
C2R repeat 1 + 1 reserve 8,000 $0.14
total 10 dispatches priced $0.92

Today's headroom (UTC 2026-07-30) is $1.9977568927. The worst case fits. If it did not, C2R would be dropped first and the deferral recorded.

The translation limb, the recount, the item build and every analysis cost $0 and are never ledgered (charter §3, A4).

9. Verification

analysis/verify.py imports nothing from analyse.py, re-parses the stored .raw bytes, re-derives every code from the option-level judgments by an independently written implementation of §3 steps 4–5, and asserts every figure in the result page. At least two mutation tests.

10. What this experiment may not conclude