Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260728i-coverage-replication/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260728i-coverage-replication
statusfrozen
created2026-07-28
updated2026-07-28
sensesaccuracy, naturalness, style-correspondence, voice, cultural-mediation, consistency, purpose-fit
provisionaltrue
linksframework/traceability-inventory.md, wiki/findings/results/RS-20260726e-framework-coverage.md, wiki/arms/ARM-framework.md, workshop/translations/mare-au-diable/R04-v1/translation.md, workshop/translations/takasebune/R04-v1/translation.md, config/models.md, config/budget.md

Design — does the framework's coverage figure survive an independent reader, and does it replicate off the pair it was measured on?

Frozen 2026-07-28 (S051) before any reader call and before the lead's mapping of Log B was written. ARM-framework step 4. Charter §8 discipline: frozen design → independent pre-run critic → run with raw outputs preserved → post-run verification recomputing every reported number.

1. Question

RS-20260726e-framework-coverage reports that the project's candidate recommendations cover 10 of 21 decisions in a real translation and decide 0 of 21. ARM-framework's closure statement rests on that figure. The figure has three known weaknesses, two of them named by the page itself:

  1. One reader. The mapping from decision to candidate is the lead's, unchecked. The page calls it "the weakest link in this page and the cheapest thing to check" and makes it revision trigger 2: a second reader re-maps the twenty-one decisions and disagrees on more than two. ARM-framework step 4 absorbed this check into the arm's completion criterion at S046.
  2. One pair, and the wrong one. The measured translation is FR→EN, and the inventory's §3 item 4 records that zero of fourteen candidates are evidenced on French→English. Every hit was therefore an application outside its evidenced pairs, carrying untested under D-20260724-04. The figure has never been measured where the inventory is strongest.
  3. One text, one translator, 542 words.

This design attacks (1) and (2). It cannot attack (3).

2. Materials

id what frozen at
Log A the 21-decision translator's log of T-mare-au-diable-R04-v1 (George Sand, La Mare au Diable ch. II, FR→EN) commit 9079930, 2026-07-26, before the inventory existed
Log B the translator's log of T-takasebune-R04-v1 (Mori Ōgai 高瀬舟, JA→EN, 1,506 source characters) this session, committed before the lead's mapping of it was written and before any reader call
the candidates all fourteen rows of framework/traceability-inventory.md §2, each with the operative sentence of its claim page where one exists 2026-07-28

Why JA→EN, and it is the choice that matters. JA→EN is the pair the inventory is best evidenced on: 4 of the 14 candidates name it in their evidenced-pairs column (#2 fluency cost, #3 handling set, #5 grammar transcoding, #7 forked-class non-uniformity), against 0 of 14 for FR→EN. If coverage does not rise on the inventory's best pair, the emptiness is not a pair-coverage artefact. The design is built to give the inventory its best case.

Candidate #12 is included in the list given to the readers, with its text as written and no note that it is inadmissible. Withholding it would rig the prescriptive count to zero. It is scored separately, because what it would contribute if Tier D ever passed is a number the closure statement needs.

Contamination. Measured before the study limb was designed, per the standing rule (CLAUDE.md, note (bcd)) — reported on the translation artifact. The relevant threat here is not baseline inflation but authenticity of the log: a translator reproducing a remembered English is not deciding.

3. The scheme

Each reader labels every decision in both logs on a three-way scheme, and names the candidate id(s) invoked:

Coverage = DECIDES + INFORMS. Prescriptive coverage = DECIDES. These reproduce RS-20260726e §1's two figures: its "covered by a surviving claim" is coverage, its "covered by something that would have changed what was written: 0 of 21" is prescriptive coverage.

4. Procedure

  1. Freeze this design. Done before step 2.
  2. Translate 高瀬舟 spans A and B under R06 (draft, frozen as its own artifact) then R04 (self-revision); write Log B at translation time; commit. The commit is the freeze.
  3. The lead writes its own mapping of Log B on the scheme in §3 and commits it before any reader call. The lead's Log A mapping already exists, published, in RS-20260726e §1.
  4. Independent pre-run critic — P2 google/gemini-3.6-flash, one call. P2 is neither reader. Findings are dispositioned in writing before dispatch.
  5. Two readers, one call each, blind to each other and to both lead mappings — P1 openai/gpt-5.6-terra and P3 x-ai/grok-4.5. Identical prompt, identical materials, temperature: 0, reasoning: {"effort":"low"}, brevity instructed (note (b), note (abc)). Raw request/response JSON preserved under runs/.
  6. Analysis in analysis/analyse.py; independent verification in analysis/verify.py, which imports nothing from analysis/analyse.py and recomputes every reported number from the stored raw bodies.

Blinding, stated exactly. The readers see the fourteen candidates and the two decision lists. They do not see RS-20260726e, the lead's mapping, this design, or each other's output. They are not told which log is which pair beyond what the decision texts themselves reveal — which is a good deal, and is not controllable.

5. Registered predictions

Registered before any reader call. The lead's Log B mapping is frozen first so that it cannot be adjusted toward these.

# prediction falsified if
P1 On Log A, RS-20260726e's revision trigger 2 fires: at least one reader disagrees with the published mapping on more than 2 of 21 decisions, on the binary covered/not both readers disagree on ≤ 2
P2 Reader-vs-reader agreement on the binary is no higher than the mean of the two lead-vs-reader agreements, +0.10 tolerance R1~R2 exceeds mean(lead~R) by more than 0.10
P3 DECIDES ≤ 1 per reader per log any reader returns ≥ 2 DECIDES on either log
P4 Log B coverage (mean over readers) does not exceed Log A coverage (mean over readers) by more than 0.15, despite JA→EN being the best-evidenced pair the gap exceeds +0.15
P5 ≥ 50% of Log B's decisions that both readers label NONE are about a single word or a single figure (Log A: 9 of 11 = 0.818, RS-20260726e §3) below 50%

P3 is close to entailed and is registered anyway, because the only candidate shaped do X rather than Y addressed to a translator is #12, so a DECIDES can essentially only come through it. What is not entailed is which decisions #12 reaches, and that is the number ARM-framework's closure statement needs.

P4 is the one this session exists to test. The prediction is a null: that giving the inventory its best pair buys nothing.

6. Failure criteria, pre-committed

  1. A reader returns fewer labels than decisions, labels outside the scheme, or a malformed body → the call is void; rerun once; a second failure drops that reader and the run reports on one, saying so.
  2. Reader-vs-reader agreement on the binary below 0.50 → no coverage figure from this run is reportable as an estimate. The run's sole output is then the reliability finding, and the closure statement must say that the project cannot measure its own coverage rather than quoting a number. This is pre-committed and is the outcome the lead considers most likely after P1.
  3. A reader's call costs more than $0.40 → the second reader is dropped and the run reports on one.
  4. finish_reason: length with empty content → drop that slug rather than retry it (note (b), seven sessions of evidence). Declared reserve for either reader: qwen/qwen3.7-max (config/models.md, first reserve). A reserve substitution is reported as such.

7. What this design cannot establish

8. Cost

Pre-flight, built from max_tokens at list out-price plus the prompt at list in-price (note (abc)):

call model max_tokens worst case
pre-run critic google/gemini-3.6-flash (P2) 8,000 $0.072
reader 1 openai/gpt-5.6-terra (P1) 12,000 $0.203
reader 2 x-ai/grok-4.5 (P3) 12,000 $0.090
total $0.365

Day headroom at design time: $3.493220 of $5.00 (2026-07-28, seven prior sessions). Fits. The translation limb costs $0 and is never ledgered.


9. Amendments after the pre-run critic pass (2026-07-28, before dispatch)

Critic P2 google/gemini-3.6-flash, verdict NEEDS-REDESIGN, nine findings, seven accepted, one accepted-as-limitation, one declined. Full dispositions with the reasoning: critic/dispositions.md. Sections 1–8 above are preserved verbatim, including the sentences withdrawn below, so what was withdrawn is visible.

  1. §5 P4 is evaluated on three Log B subsets — all 42, draft-only B1–B27, revision-only B28–B42 — because the two logs differ in size and composition and no subset is a clean counterpart to Log A's 21.
  2. §7's "the bias runs toward coverage … P4's null is the hard direction and a rise in coverage is uninterpretable" is WITHDRAWN. Priming changes not only which decisions are noticed as inventory-shaped but which decisions get recorded at all; the two effects push the proportion in opposite directions and neither is measured, so the null is as uninterpretable as the rise. What survives: the reliability measurement on Log A (published mapping, log written before the inventory existed) and the DECIDES count. Log B's coverage proportions are description of one log, not a replication of RS-20260726e's figure, and the word replication is withdrawn from this design's claim about itself.
  3. §6 failure criterion 2 is restated on Cohen's κ: coverage figures are unreportable as estimates at κ ≤ 0.20 on the binary. Raw agreement and the majority-class baseline (0.524 on the lead's Log B base rates) are reported alongside. The original 0.50 raw-agreement criterion sat at chance and is withdrawn.
  4. New P6, registered before dispatch and capable of defeating the closure this arm intends: if either reader returns DECIDES ≥ 2 on Log B through a candidate other than C12, the inventory holds prescriptive-about-translating content the lead's mapping missed, and ARM-framework must be extended, not closed.
  5. The reader instruction gains one clause (rule 2): a candidate that merely lists options is INFORMS even where one listed option fits; DECIDES requires the candidate to do the ranking.
  6. §4's blinding claim is corrected. There is no pair blinding: the source strings identify both pairs outright.
  7. P3 is restated as a check on reader discipline, not on the inventory.
  8. P5 is WITHDRAWN. "About a single word or a single figure" has no operational definition, no assigned classifier and no independent check. RS-20260726e §3's grain finding is not re-tested here.
  9. Analysis and verifier join on an explicit D<n> ↔ B<n> map with assertions on both key sets, and report three-way agreement alongside the binary.
  10. §6 criterion 3 rises to $0.50 per reader call; a finish_reason: length body containing all 63 parseable lines is accepted rather than voided.