Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260730e-rule-coverage/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260730e-rule-coverage
statusfrozen
created2026-07-30
updated2026-07-30
trackT2
sensesnaturalness, style-correspondence, cultural-mediation, accuracy
linkswiki/arms/ARM-rule-coverage.md, workshop/regimes/R07-fluency.md, workshop/translations/petits-poemes/R07-v1/translation.md, workshop/translations/postmaster/R07-v1/translation.md, workshop/translations/osso-di-morto/R07-v1/translation.md, config/models.md, config/budget.md

E-20260730e — do readers who are not the translator assign R07's coverage codes as the translator did?

Frozen before any call is dispatched. ARM-rule-coverage step 1, study limb.

Question

R07 measures its own success by a coverage rate: the fraction of contested sites at which a numbered rule decides (D), as against merely permitting (P) or being silent (S). That rate is this project's operational answer to Tymoczko's objection that Venuti supplies no criteria. Every coverage figure the project holds was assigned by the agent that wrote the rules and made the choices, which R07 §Known limitations declares about itself and nothing has ever measured.

Two questions, one design:

  1. Is a coverage code a property a reader can recover from the rule set, the site and the live options?
  2. Does F10 forbid a translation from reproducing a figure the source contains? — the IR1 question, which the Bengali run had to rule on and which the rule set does not answer.

Materials

Condition C1 — the item pool. 37 sites drawn from the three frozen R07 logs, built by build_items.py (seed 20260730) and stored in items.json. Stratified, not proportional:

Italian (osso) Bengali (postmaster) French (petits) total
D 3 5 4 12
P 4 4 4 12
S 8 3 2 13 (every S site in the project)
total 15 12 10 37

Natural marginals across all 110 logged sites: D 39, P 58, S 13. The pool is balanced deliberately, so raw agreement on this pool is not an estimate of population agreement and is never reported as one; the re-weighted estimate below is.

Each item shows the rater: the site, and the live options in shuffled order (seed-fixed; permutation stored). It does not show the chosen option, the code, or the rules the translator cited. R07 §5 defines the code as a textual test over the live options — does a rule name this feature, and does only one live option satisfy it — so this is exactly the information the code is defined on.

Condition C2 — the IR1 sites. Four figurative sites, source string plus two renderings (figure carried over / figure flattened), and F10's text. Sites: postmaster #14 («প্রকৃতির দরবারে» — nature's court), postmaster #18 (the canvas of her heart), petits #4 (the point of the Infinite), petits #20 (the eternal heat lolling).

Roles (config/models.md)

Procedure

  1. C1: one call per rater, all 37 items, temperature 0, max_tokens 8000.
  2. C1R: P1's C1 prompt re-sent byte-identical, same day. Prompt token counts asserted identical to the digit.
  3. C2: one call per rater, all 4 sites, independent context, max_tokens 4000.
  4. Recognition question appended to C1 and C2 (note (bes)): do you recognise any of this material — name the work and author if so. Two output lines; the rate is published whatever it is.

Registered predictions and failure criteria

F1 — the positive control, and the run is void without it. Three items (postmaster #11 «খোল-করতাল», postmaster #25 «শ্মশান», petits #22 «gargoulettes») put one thing in contest: keep the source-language word, or translate it. F4 says "No untranslated source-language words" in so many words. Condition: at ≥ 2 of the 3 items, ≥ 2 of 3 raters return D. If it fails, no agreement figure from this run is reported as a measurement of the rule set; the run is reported as an instrument failure and the arm's condition 2 stays open.

# prediction fails if
Q2 — PRIMARY Krippendorff's α (nominal, 3 raters × 37 items) ≥ 0.40. Reference: this project's categorical instruments run 0.51 (four-class drift, three raters, RS-20260729b) down to 0.066 once an unreachable label is removed. α < 0.40 — and that is a first-class result, not a null: it says the coverage code is not reader-shareable, which answers the arm's question in the negative
Q3 Re-weighted agreement with the lead's code ≥ 0.50, where re-weighted = Σ_c (natural_c / 110) × (agreement within stratum c). < 0.50
Q4 Raters return S less often than the lead did — mean rater S-rate on the pool < 35.1% (the lead's 13/37). Mechanism: the ten rules are broadly worded, and a reader asked does any rule bear will reach for one where the translator, who knew which rule it meant, recorded silence. mean rater S-rate ≥ 35.1%
Q5 — the gloss control, which can embarrass this design Lead-agreement on glossed items (those whose site cell carries a lead-written English gloss, 15 of 37 — all 12 Bengali plus 3 others) is not more than 0.15 above unglossed. difference > 0.15 — in which case the gloss is doing the work and the Bengali agreement figure is a measurement of the lead's prose
Q6 — reachability (notes (bdq), (beb)) Each of D, P, S is used at least once by every rater. any rater omits a label — α is then reported as corrected against a category absent from that rater's data, and flagged
Q7 — C2 The three raters converge (all three return the same one of KEEP / FLATTEN / F10-DOES-NOT-DECIDE) at ≤ 2 of 4 sites. Registered direction: no convergence, i.e. the rule does not decide. convergence at ≥ 3 of 4

Q8 — the repeat control and Q9 — the recognition rate are reported, not predicted: neither has a prior this project can defend.

Declared limitations, before the critic sees this

Pre-flight cost (note (abc) — worst case from max_tokens, not from expected output)

call slug max_tokens worst case
C1 ×3 P1, P3, P5 8,000 $0.065 + $0.056 + $0.033 = $0.154
C1 reserve allowance P2 8,000 $0.066
C1R ×1 P1 8,000 $0.065
C2 ×3 P1, P3, P5 4,000 $0.074
critic ×1 P4 16,000 $0.264

Cap-literal worst case $0.623. Reasoning tokens can exceed the cap — S062 measured 3× the output cap on one slug — so the reasoning-inflated figure on the output component is ≈ $1.16. The larger is declared: $1.16. P5 is priced at the worst plausible provider per the S022 routing caution, not at list. Today's headroom is $3.4762; three sessions have already run on 2026-07-30.