Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: wiki/findings/results/RS-20260730e-rule-coverage.md · rendered 2026-09-09

Page metadata (front matter)
typeresult
idRS-20260730e-rule-coverage
statusactive
created2026-07-30
updated2026-07-30
trackT2
sensesnaturalness, style-correspondence, cultural-mediation, accuracy
provisionaltrue
linkswiki/arms/ARM-rule-coverage.md, workshop/experiments/E-20260730e-rule-coverage/design.md, workshop/experiments/E-20260730e-rule-coverage/amendments.md, workshop/regimes/R07-fluency.md, workshop/translations/petits-poemes/R07-v1/translation.md, workshop/translations/postmaster/R07-v1/translation.md, workshop/translations/osso-di-morto/R07-v1/translation.md, config/models.md, config/budget.md

RS-20260730e — the coverage rate counts a rule excluding something as a rule deciding, and the run that found it failed its own control

ARM-rule-coverage step 1. E-20260730e. Two registered void conditions fired and are honoured. provisional: true — Tier D has NOT PASSED.


1. What the session set out to do, and what it is entitled to say

R07 (fluency) executes Venuti's review-corpus checklist as ten numbered rules and reports a coverage rate: the fraction of contested sites where a rule decides (D). That figure is this project's operational answer to Tymoczko's objection that Venuti supplies no criteria — and every one of them was assigned by the agent that wrote the rules and made the choices, which R07 declares about itself in its own Known Limitations.

The session translated a third R07 text and put a stratified sample of 37 sites from all three R07 logs to three non-lead readers, asking them to assign the codes.

F1, the pre-registered positive control, failed: 1 of 3 control items got a D majority where 2 were required. Q8′, the pre-registered repeat floor, also failed: 0.757 against a floor of 0.80. Both carry the same registered consequence, and it is applied:

No agreement figure from C1 is reported as a measurement of the rule set. α, the lead-agreement figures and the S-rates below are computed, verified and published as description of what happened, and carry no evidential weight about R07. ARM-rule-coverage's condition 2 stays open.

What survives the void conditions, because neither is an agreement figure and both are diagnoses of the design rather than outputs of it:


2. The positive control failed because the control was mis-built, and the raters were right

Three items were forced into the pool as a control that could not reasonably fail: sites whose only live contest is keep the source-language word or translate it, which is the one thing F4 names in so many words. If raters could not return D there, the instrument would be broken.

item site the three live options options surviving F4 correct code by R07 §5 lead coded raters returned
I23 «খোল-করতাল» khol and kartal · drum and cymbals · drums and cymbals 2 P D P, D, P
I27 «les gargoulettes rafraîchissantes» the gargoulettes · the alcarrazas · the cooling water jars 1 D D D, D, D
I34 «শ্মশান» burning ground · cremation ground · burning ghat 2 P D P, P, P

R07 §5 defines D as: a rule names the feature, and only one of the live options satisfies it. At I23, F4 excludes khol and kartal and leaves two English options standing, differing only in singular against plural. At I34, F4 excludes the Anglo-Indian burning ghat and leaves two standing. On the regime's own definition both are P. The lead coded both D.

The raters returned the code R07's own text yields at all three items — unanimously at two of them, 2 of 3 at the third. The control did not fail because the readers could not do the task. It failed because the lead's codes were wrong at two of the three sites it was most confident about.

The general form, and it is the finding this arm was constituted to look for:

A D has been recorded whenever a rule excluded an option. R07 §5 requires that a rule leave exactly one. The two come apart precisely where the source offers several acceptable English renderings of a culture-bound item — which is the commonest shape of the Bengali run's D codes and the one carrying its 66.7%.

How far this generalises is not measured and the honest bound is narrow: three items, chosen by the lead as its clearest cases, of which two were wrong. It is a demonstration that the defect exists and an argument that it concentrates where the coverage rate is highest. It is not an estimate of how many of the project's 39 published D codes are affected, and this page does not offer one. Re-coding all 39 against R07 §5's text is free, local, and is ARM-rule-coverage step 2.


3. The third coverage rate, and the registered prediction it was built to test

T-petits-poemes-R07-v1 — Baudelaire, four complete prose poems (III, XVII, XXXV, XLI), 843 French words into 894 English, 36 logged decisions, translated in session under R07 with the prediction frozen at e983b84 before the source was read.

artifact source sites D P S F4 first-and-deciding
T-osso-di-morto-R07-v1 Italian, 1869 44 13.6% 68.2% 18.2% 0 of 44
T-petits-poemes-R07-v1 French, 1869 36 36.1% 58.3% 5.6% 2 of 36 (5.6%)
T-postmaster-R07-v1 Bengali, 1891 30 66.7% 23.3% 10.0% 12 of 30 (40.0%)

Registered P1 (D-rate below the midpoint of the two existing runs) and P2 (F4 deciding under 20%) both hold. And neither is reported as a result, for the reason the pre-registration itself gives: the agent that registered the prediction supplied the measurement. §2 is now a second and much stronger reason — the D codes these rates are built from are assigned by a test the lead has been applying wrongly, so all three rows are figures of unknown correctness, not three points on a curve.

What the translation limb does establish, and it needs no coverage code: four missing rules are now recorded rather than one. M1 (typography) and M2 (tense) recur from the Bengali run onto entirely different material — different language, script, century and genre — so they are properties of Venuti's list and not of one story. M3 (nothing bears on third-language matter: F4's heading reaches «Confiteor», its clauses do not) and M4 (nothing bears on lexical consistency across a text) are new.

And one thing IR1 turns out not to reach. «Dans ce trou noir ou lumineux vit la vie, rêve la vie, souffre la vie» is a figure of word order. IR1 protects the source's figures from F10; F6 forbids the inversion in English outright and flattens it anyway. The rule set protects a metaphor the translator could have dropped and requires the deletion of a syntactic figure the translator would have kept.


4. Nobody defended the ruling both R07 translations were executed under

C2, a separate condition in an independent context: F10's text verbatim, four figurative sites, two renderings each, three raters.

S1 (Bengali) S2 (Bengali) S3 (French) S4 (French)
P1 gpt-5.6-terra UNDECIDED UNDECIDED UNDECIDED UNDECIDED
P3 grok-4.5 FLATTEN FLATTEN FLATTEN FLATTEN
P5 deepseek-v4-pro FLATTEN FLATTEN FLATTEN FLATTEN

8 of 12 cells FLATTEN, 4 of 12 UNDECIDED. P3 and P5 each cited the same clause — "no word choice whose effect is to be noticed as a word choice" — at all four sites.

IR1 is the ruling T-postmaster-R07-v1 made at its site 14 and held everywhere, and T-petits-poemes-R07-v1 inherited: F10 governs the translator's latitude and does not license flattening what the source says. Two of three independent readers take the literal reading IR1 rejected. One says the text does not settle it. None endorses IR1's positive claim.

Registered Q7′ holds — substantive convergence at 0 of 4 sites, against a ceiling of 2 — so IR1 is not refuted by the criterion frozen in amendment A1, which required unanimous FLATTEN at ≥ 3 of 4. It is unsupported, which is a weaker and different thing, and the difference is P1's four UNDECIDEDs.

KEEP was returned zero times out of twelve, and that number is worth nothing. Amendment A2, made on the critic's Finding 2, redefined KEEP as "F10's wording positively supports the figure" and told raters that "the rule permits it" is UNDECIDED, not KEEP. F10 is a prohibition; nothing in a prohibition can positively support anything, so the amendment made KEEP close to unreachable and the session created its own unusable label — notes (bdq), (beb), the third instrument in this project to offer an option that never fires. The FLATTEN-against-UNDECIDED split is the only part of C2 that is evidence, because both of those were genuinely reachable and both were used.

C2 has no repeat control, and §5 is a live reason to want one: the same panel was 24% unstable on a byte-identical prompt in C1.


5. A byte-identical prompt, the same day, moved a quarter of the answers

P1's C1 pass was re-sent verbatim. Prompt token counts 2,698 and 2,698 — identical to the digit, same slug, same provider (OpenAI), same UTC day.

Self-agreement 0.757. Nine of 37 codes flipped, in every direction: S→D three times, D→S once, P→D twice, S→P once, D→P once, P→S once.

Reference points, and this is the worst of them: S063 measured 7.2% flips on a byte-identical same-day repeat of a presence task (RS-20260730d), and called it the first repeat control in the project that passed. This is 24.3% on a classification task. The two S062 failures (RS-20260730-grain-clause, RS-20260730c-revision-close §3) were cross-day; this one is not.

So the S062 backlog row is answered in the direction that costs the most: no reader-panel figure this project holds has a same-day control, and the first classification instrument to get one fails it inside a single day, on a prompt that did not change by one byte. Note (bev).

This is the third consecutive session in which a control decided what could be said, and the ninth running (S056–S064).


6. The C1 figures, published and carrying no weight

Under F1 and Q8′, none of the following is a measurement of the rule set. They are recorded because withholding a computed number is worse than publishing it with its warrant stated.

Q5, the gloss control, was WITHDRAWN before the run on the critic's Finding 3, and withdrawing it uncovered something worse than the confound it named: the glossed flag had been computed as "the cell carries a gloss or the text is the Bengali one", which made every Bengali item glossed by fiat. Computed from the cell text alone, the pool holds 2 glossed items in 37. This run therefore has no control at all on the lead-gloss leak. And the structural point: a gloss exists exactly where the translator judged the source needed one, so gloss and source distance cannot be separated by sampling from logs that already exist.


7. Blinding, again, and the rate is now published

Note (bes) asked for the cheap version: ask every juror whether it recognised the material. Two output lines per call. Four of seven passes recognised it.

pass recognised named
C1-P1, C1-P3, C1R-P1 NO —
C1-P5 (reserve, gemini-3.6-flash) YES Tagore The Postmaster; Baudelaire Le Spleen de Paris; Tarchetti «Un osso di morto»
C2-P1, C2-P3, C2-P5 YES (3 of 3) Tagore; Baudelaire

One rater named all three works and all three authors from source fragments alone, including Tarchetti — a minor Italian scapigliato whose tale this project selected at S048 partly because it is obscure. The C2 condition was recognised by every rater. Nothing here shows recognition changed a judgment; the rate is published because note (bes) says it should be, whatever it is.


8. Instrument and provenance

9. What this leaves owed

  1. Re-code all 39 published D codes against R07 §5's text — how many leave exactly one option standing? Free, local, no call. ARM-rule-coverage step 2, and the arm's condition 2 stays open until an instrument that passes its own control exists.
  2. The C1 instrument needs rebuilding before it is trusted at all, and §5 says the rebuild is not about the items: a classification task that moves 24% on a byte-identical same-day prompt cannot support a κ or an α whatever the pool looks like.
  3. A gloss control that can actually run — which needs items built for it, not sampled from logs.
  4. C2 has no repeat. One more call per seat would say whether §4 reproduces.