Repository path: wiki/findings/results/RS-20260730e-rule-coverage.md · rendered 2026-09-09
Page metadata (front matter)
| type | result |
|---|---|
| id | RS-20260730e-rule-coverage |
| status | active |
| created | 2026-07-30 |
| updated | 2026-07-30 |
| track | T2 |
| senses | naturalness, style-correspondence, cultural-mediation, accuracy |
| provisional | true |
| links | wiki/arms/ARM-rule-coverage.md, workshop/experiments/E-20260730e-rule-coverage/design.md, workshop/experiments/E-20260730e-rule-coverage/amendments.md, workshop/regimes/R07-fluency.md, workshop/translations/petits-poemes/R07-v1/translation.md, workshop/translations/postmaster/R07-v1/translation.md, workshop/translations/osso-di-morto/R07-v1/translation.md, config/models.md, config/budget.md |
RS-20260730e — the coverage rate counts a rule excluding something as a rule deciding, and the run that found it failed its own control
ARM-rule-coverage step 1. E-20260730e. Two registered void conditions fired and are honoured. provisional: true — Tier D has NOT PASSED.
1. What the session set out to do, and what it is entitled to say
R07 (fluency) executes Venuti's review-corpus checklist as ten numbered rules and reports a coverage rate: the fraction of contested sites where a rule decides (D). That figure is this project's operational answer to Tymoczko's objection that Venuti supplies no criteria — and every one of them was assigned by the agent that wrote the rules and made the choices, which R07 declares about itself in its own Known Limitations.
The session translated a third R07 text and put a stratified sample of 37 sites from all three R07 logs to three non-lead readers, asking them to assign the codes.
F1, the pre-registered positive control, failed: 1 of 3 control items got a D majority where 2 were required. Q8′, the pre-registered repeat floor, also failed: 0.757 against a floor of 0.80. Both carry the same registered consequence, and it is applied:
No agreement figure from C1 is reported as a measurement of the rule set. α, the lead-agreement figures and the S-rates below are computed, verified and published as description of what happened, and carry no evidential weight about R07.
ARM-rule-coverage's condition 2 stays open.
What survives the void conditions, because neither is an agreement figure and both are diagnoses of the design rather than outputs of it:
- §2 — why the positive control failed, which is a mechanical fact about three item option-sets and is the session's substantive result.
- §4 — the C2 marginal, a tally of which label was used, on a separate condition with its own registered criterion.
2. The positive control failed because the control was mis-built, and the raters were right
Three items were forced into the pool as a control that could not reasonably fail: sites whose only live contest is keep the source-language word or translate it, which is the one thing F4 names in so many words. If raters could not return D there, the instrument would be broken.
| item | site | the three live options | options surviving F4 | correct code by R07 §5 | lead coded | raters returned |
|---|---|---|---|---|---|---|
| I23 | «খোল-করতাল» | khol and kartal · drum and cymbals · drums and cymbals | 2 | P | D | P, D, P |
| I27 | «les gargoulettes rafraîchissantes» | the gargoulettes · the alcarrazas · the cooling water jars | 1 | D | D | D, D, D |
| I34 | «শ্মশান» | burning ground · cremation ground · burning ghat | 2 | P | D | P, P, P |
R07 §5 defines D as: a rule names the feature, and only one of the live options satisfies it. At I23, F4 excludes khol and kartal and leaves two English options standing, differing only in singular against plural. At I34, F4 excludes the Anglo-Indian burning ghat and leaves two standing. On the regime's own definition both are P. The lead coded both D.
The raters returned the code R07's own text yields at all three items — unanimously at two of them, 2 of 3 at the third. The control did not fail because the readers could not do the task. It failed because the lead's codes were wrong at two of the three sites it was most confident about.
The general form, and it is the finding this arm was constituted to look for:
A
Dhas been recorded whenever a rule excluded an option.R07§5 requires that a rule leave exactly one. The two come apart precisely where the source offers several acceptable English renderings of a culture-bound item — which is the commonest shape of the Bengali run's D codes and the one carrying its 66.7%.
How far this generalises is not measured and the honest bound is narrow: three items, chosen by the lead as its clearest cases, of which two were wrong. It is a demonstration that the defect exists and an argument that it concentrates where the coverage rate is highest. It is not an estimate of how many of the project's 39 published D codes are affected, and this page does not offer one. Re-coding all 39 against R07 §5's text is free, local, and is ARM-rule-coverage step 2.
3. The third coverage rate, and the registered prediction it was built to test
T-petits-poemes-R07-v1 — Baudelaire, four complete prose poems (III, XVII, XXXV, XLI), 843 French words into 894 English, 36 logged decisions, translated in session under R07 with the prediction frozen at e983b84 before the source was read.
| artifact | source | sites | D | P | S | F4 first-and-deciding |
|---|---|---|---|---|---|---|
T-osso-di-morto-R07-v1 |
Italian, 1869 | 44 | 13.6% | 68.2% | 18.2% | 0 of 44 |
T-petits-poemes-R07-v1 |
French, 1869 | 36 | 36.1% | 58.3% | 5.6% | 2 of 36 (5.6%) |
T-postmaster-R07-v1 |
Bengali, 1891 | 30 | 66.7% | 23.3% | 10.0% | 12 of 30 (40.0%) |
Registered P1 (D-rate below the midpoint of the two existing runs) and P2 (F4 deciding under 20%) both hold. And neither is reported as a result, for the reason the pre-registration itself gives: the agent that registered the prediction supplied the measurement. §2 is now a second and much stronger reason — the D codes these rates are built from are assigned by a test the lead has been applying wrongly, so all three rows are figures of unknown correctness, not three points on a curve.
What the translation limb does establish, and it needs no coverage code: four missing rules are now recorded rather than one. M1 (typography) and M2 (tense) recur from the Bengali run onto entirely different material — different language, script, century and genre — so they are properties of Venuti's list and not of one story. M3 (nothing bears on third-language matter: F4's heading reaches «Confiteor», its clauses do not) and M4 (nothing bears on lexical consistency across a text) are new.
And one thing IR1 turns out not to reach. «Dans ce trou noir ou lumineux vit la vie, rêve la vie, souffre la vie» is a figure of word order. IR1 protects the source's figures from F10; F6 forbids the inversion in English outright and flattens it anyway. The rule set protects a metaphor the translator could have dropped and requires the deletion of a syntactic figure the translator would have kept.
4. Nobody defended the ruling both R07 translations were executed under
C2, a separate condition in an independent context: F10's text verbatim, four figurative sites, two renderings each, three raters.
| S1 (Bengali) | S2 (Bengali) | S3 (French) | S4 (French) | |
|---|---|---|---|---|
P1 gpt-5.6-terra |
UNDECIDED | UNDECIDED | UNDECIDED | UNDECIDED |
P3 grok-4.5 |
FLATTEN | FLATTEN | FLATTEN | FLATTEN |
P5 deepseek-v4-pro |
FLATTEN | FLATTEN | FLATTEN | FLATTEN |
8 of 12 cells FLATTEN, 4 of 12 UNDECIDED. P3 and P5 each cited the same clause — "no word choice whose effect is to be noticed as a word choice" — at all four sites.
IR1 is the ruling T-postmaster-R07-v1 made at its site 14 and held everywhere, and T-petits-poemes-R07-v1 inherited: F10 governs the translator's latitude and does not license flattening what the source says. Two of three independent readers take the literal reading IR1 rejected. One says the text does not settle it. None endorses IR1's positive claim.
Registered Q7′ holds — substantive convergence at 0 of 4 sites, against a ceiling of 2 — so IR1 is not refuted by the criterion frozen in amendment A1, which required unanimous FLATTEN at ≥ 3 of 4. It is unsupported, which is a weaker and different thing, and the difference is P1's four UNDECIDEDs.
KEEP was returned zero times out of twelve, and that number is worth nothing. Amendment A2, made on the critic's Finding 2, redefined KEEP as "F10's wording positively supports the figure" and told raters that "the rule permits it" is UNDECIDED, not KEEP. F10 is a prohibition; nothing in a prohibition can positively support anything, so the amendment made KEEP close to unreachable and the session created its own unusable label — notes (bdq), (beb), the third instrument in this project to offer an option that never fires. The FLATTEN-against-UNDECIDED split is the only part of C2 that is evidence, because both of those were genuinely reachable and both were used.
C2 has no repeat control, and §5 is a live reason to want one: the same panel was 24% unstable on a byte-identical prompt in C1.
5. A byte-identical prompt, the same day, moved a quarter of the answers
P1's C1 pass was re-sent verbatim. Prompt token counts 2,698 and 2,698 — identical to the digit, same slug, same provider (OpenAI), same UTC day.
Self-agreement 0.757. Nine of 37 codes flipped, in every direction: S→D three times, D→S once, P→D twice, S→P once, D→P once, P→S once.
Reference points, and this is the worst of them: S063 measured 7.2% flips on a byte-identical same-day repeat of a presence task (RS-20260730d), and called it the first repeat control in the project that passed. This is 24.3% on a classification task. The two S062 failures (RS-20260730-grain-clause, RS-20260730c-revision-close §3) were cross-day; this one is not.
So the S062 backlog row is answered in the direction that costs the most: no reader-panel figure this project holds has a same-day control, and the first classification instrument to get one fails it inside a single day, on a prompt that did not change by one byte. Note (bev).
This is the third consecutive session in which a control decided what could be said, and the ninth running (S056–S064).
6. The C1 figures, published and carrying no weight
Under F1 and Q8′, none of the following is a measurement of the rule set. They are recorded because withholding a computed number is worse than publishing it with its warrant stated.
- Krippendorff's α (nominal, 3 raters × 37 items) = 0.386, against a registered 0.40 — recomputed by two independent routes agreeing to 1e-12. Pairwise: P1~P2 0.730, P1~P3 0.622, P3~P2 0.568.
- Mean re-weighted agreement with the lead 0.654 (raw on the balanced pool 0.577; the re-weighting to the natural marginals D 39 / P 58 / S 13 of 110 moves it +0.077).
- Within stratum, the asymmetry runs the way §2 predicts: raters confirm the lead's P codes at 0.75 / 0.67 / 1.00, and its D codes at 0.42 / 0.83 / 0.25.
- Q4 holds: raters return S far less often than the lead — 0.243 / 0.189 / 0.108 against the lead's 0.351.
- Q6 reachability passes: all three labels used by all three raters.
- Rules cited across 111 codings: F8 35, F6 26, F4 25, F2 15, F10 11, F7 8, F1 7, F5 6, F9 4, F3 3. Mean Jaccard between raters' cited rule-sets per item 0.489 — when they agree on the code they often disagree about which rule produced it.
- By language (confounded four ways over — language, code distribution, whether the site was glossed, which text — and reported as description only): α French 0.415, Bengali 0.336, Italian 0.302.
Q5, the gloss control, was WITHDRAWN before the run on the critic's Finding 3, and withdrawing it uncovered something worse than the confound it named: the glossed flag had been computed as "the cell carries a gloss or the text is the Bengali one", which made every Bengali item glossed by fiat. Computed from the cell text alone, the pool holds 2 glossed items in 37. This run therefore has no control at all on the lead-gloss leak. And the structural point: a gloss exists exactly where the translator judged the source needed one, so gloss and source distance cannot be separated by sampling from logs that already exist.
7. Blinding, again, and the rate is now published
Note (bes) asked for the cheap version: ask every juror whether it recognised the material. Two output lines per call. Four of seven passes recognised it.
| pass | recognised | named |
|---|---|---|
| C1-P1, C1-P3, C1R-P1 | NO | — |
C1-P5 (reserve, gemini-3.6-flash) |
YES | Tagore The Postmaster; Baudelaire Le Spleen de Paris; Tarchetti «Un osso di morto» |
| C2-P1, C2-P3, C2-P5 | YES (3 of 3) | Tagore; Baudelaire |
One rater named all three works and all three authors from source fragments alone, including Tarchetti — a minor Italian scapigliato whose tale this project selected at S048 partly because it is obscure. The C2 condition was recognised by every rater. Nothing here shows recognition changed a judgment; the rate is published because note (bes) says it should be, whatever it is.
8. Instrument and provenance
- Seats are not the same across conditions and it must not be read as if they were. Note (b) fired for the twelfth time: P5
deepseek-v4-proreturnedfinish_reason: lengthwith an empty body at 8,000 tokens on C1 (GMICloud, 196 s, $0.0127 wasted), and the declared reserve P2gemini-3.6-flashtook the seat. So C1's three raters are P1 / P3 / P2 and C2's are P1 / P3 / P5. - Note (bdt) is live in this run rather than latent — the first time in this project.
C1-P5.rawholds the rejected body; the accepted one isC1-P5-reserve1.raw, andanalysis/verify.pyreads the accepted body through the attempt chain rather than by filename. - Pre-run critic
moonshotai/kimi-k3(P4 — not a rater, not the raters' reserve; the S053 role-collision fix, twelfth session running), one call,$0.05598,Fireworks. VerdictNEEDS-AMENDMENT, six findings, two BLOCKING, all six accepted, one discharged by withdrawal rather than amendment. Note (rr), twenty-second consecutive session. Its Finding 1 established that Q7 counted convergence and called it silence — unanimous FLATTEN would have scored as the prediction failing — and §4 above is readable only because that was fixed before dispatch. - Verification.
analysis/verify.pyimports nothing fromanalyse.py, re-parses the stored.rawbytes, recomputes α by the coincidence-matrix route, and asserts every figure above: 248 checks, 0 failures, plus three mutation tests, all three caught. One check in the verifier was itself wrong and is corrected in place rather than deleted — it asserted the chosen rendering is absent from the prompt, which is false by construction since the chosen rendering is one of the live options. - Errata found and applied this session:
T-osso-di-morto-R07-v1has 44 logged sites, not the 21 the pre-registration quoted from a truncated view;T-postmaster-R07-v1's own tally line read P 8 / S 2 and is P 7 / S 3. Neither touches a D count. - Cost: $0.202411981 — critic $0.05598, raters $0.146431981 across 8 dispatches (7 accepted, 1 wasted). 17.4% of the declared worst case of $1.16 (cap-literal $0.623). Key-usage delta $0.1033 against a per-request sum of $0.1464 for the rater block; the per-request sum is what is ledgered, per
config/budget.md's stated method, and the shortfall is the settling lag note (bcx) names.
9. What this leaves owed
- Re-code all 39 published D codes against
R07§5's text — how many leave exactly one option standing? Free, local, no call.ARM-rule-coveragestep 2, and the arm's condition 2 stays open until an instrument that passes its own control exists. - The C1 instrument needs rebuilding before it is trusted at all, and §5 says the rebuild is not about the items: a classification task that moves 24% on a byte-identical same-day prompt cannot support a κ or an α whatever the pool looks like.
- A gloss control that can actually run — which needs items built for it, not sampled from logs.
- C2 has no repeat. One more call per seat would say whether §4 reproduces.