Repository path: workshop/experiments/E-20260728i-coverage-replication/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260728i-coverage-replication |
| status | frozen |
| created | 2026-07-28 |
| updated | 2026-07-28 |
| senses | accuracy, naturalness, style-correspondence, voice, cultural-mediation, consistency, purpose-fit |
| provisional | true |
| links | framework/traceability-inventory.md, wiki/findings/results/RS-20260726e-framework-coverage.md, wiki/arms/ARM-framework.md, workshop/translations/mare-au-diable/R04-v1/translation.md, workshop/translations/takasebune/R04-v1/translation.md, config/models.md, config/budget.md |
Design — does the framework's coverage figure survive an independent reader, and does it replicate off the pair it was measured on?
Frozen 2026-07-28 (S051) before any reader call and before the lead's mapping of Log B was written. ARM-framework step 4. Charter §8 discipline: frozen design → independent pre-run critic → run with raw outputs preserved → post-run verification recomputing every reported number.
1. Question
RS-20260726e-framework-coverage reports that the project's candidate recommendations cover 10 of 21 decisions in a real translation and decide 0 of 21. ARM-framework's closure statement rests on that figure. The figure has three known weaknesses, two of them named by the page itself:
- One reader. The mapping from decision to candidate is the lead's, unchecked. The page calls it "the weakest link in this page and the cheapest thing to check" and makes it revision trigger 2: a second reader re-maps the twenty-one decisions and disagrees on more than two.
ARM-frameworkstep 4 absorbed this check into the arm's completion criterion at S046. - One pair, and the wrong one. The measured translation is FR→EN, and the inventory's §3 item 4 records that zero of fourteen candidates are evidenced on French→English. Every hit was therefore an application outside its evidenced pairs, carrying
untestedunderD-20260724-04. The figure has never been measured where the inventory is strongest. - One text, one translator, 542 words.
This design attacks (1) and (2). It cannot attack (3).
2. Materials
| id | what | frozen at |
|---|---|---|
| Log A | the 21-decision translator's log of T-mare-au-diable-R04-v1 (George Sand, La Mare au Diable ch. II, FR→EN) |
commit 9079930, 2026-07-26, before the inventory existed |
| Log B | the translator's log of T-takasebune-R04-v1 (Mori Ōgai 高瀬舟, JA→EN, 1,506 source characters) |
this session, committed before the lead's mapping of it was written and before any reader call |
| the candidates | all fourteen rows of framework/traceability-inventory.md §2, each with the operative sentence of its claim page where one exists |
2026-07-28 |
Why JA→EN, and it is the choice that matters. JA→EN is the pair the inventory is best evidenced on: 4 of the 14 candidates name it in their evidenced-pairs column (#2 fluency cost, #3 handling set, #5 grammar transcoding, #7 forked-class non-uniformity), against 0 of 14 for FR→EN. If coverage does not rise on the inventory's best pair, the emptiness is not a pair-coverage artefact. The design is built to give the inventory its best case.
Candidate #12 is included in the list given to the readers, with its text as written and no note that it is inadmissible. Withholding it would rig the prescriptive count to zero. It is scored separately, because what it would contribute if Tier D ever passed is a number the closure statement needs.
Contamination. Measured before the study limb was designed, per the standing rule (CLAUDE.md, note (bcd)) — reported on the translation artifact. The relevant threat here is not baseline inflation but authenticity of the log: a translator reproducing a remembered English is not deciding.
3. The scheme
Each reader labels every decision in both logs on a three-way scheme, and names the candidate id(s) invoked:
- DECIDES — a candidate determines what to write at this decision. Not "is relevant to"; determines.
- INFORMS — a candidate bears on the decision (names an option that was live, supplies a diagnostic question, supplies vocabulary) without determining it.
- NONE — no candidate bears on it.
Coverage = DECIDES + INFORMS. Prescriptive coverage = DECIDES. These reproduce RS-20260726e §1's two figures: its "covered by a surviving claim" is coverage, its "covered by something that would have changed what was written: 0 of 21" is prescriptive coverage.
4. Procedure
- Freeze this design. Done before step 2.
- Translate 高瀬舟 spans A and B under R06 (draft, frozen as its own artifact) then R04 (self-revision); write Log B at translation time; commit. The commit is the freeze.
- The lead writes its own mapping of Log B on the scheme in §3 and commits it before any reader call. The lead's Log A mapping already exists, published, in
RS-20260726e§1. - Independent pre-run critic — P2
google/gemini-3.6-flash, one call. P2 is neither reader. Findings are dispositioned in writing before dispatch. - Two readers, one call each, blind to each other and to both lead mappings — P1
openai/gpt-5.6-terraand P3x-ai/grok-4.5. Identical prompt, identical materials,temperature: 0,reasoning: {"effort":"low"}, brevity instructed (note (b), note (abc)). Raw request/response JSON preserved underruns/. - Analysis in
analysis/analyse.py; independent verification inanalysis/verify.py, which imports nothing fromanalysis/analyse.pyand recomputes every reported number from the stored raw bodies.
Blinding, stated exactly. The readers see the fourteen candidates and the two decision lists. They do not see RS-20260726e, the lead's mapping, this design, or each other's output. They are not told which log is which pair beyond what the decision texts themselves reveal — which is a good deal, and is not controllable.
5. Registered predictions
Registered before any reader call. The lead's Log B mapping is frozen first so that it cannot be adjusted toward these.
| # | prediction | falsified if |
|---|---|---|
| P1 | On Log A, RS-20260726e's revision trigger 2 fires: at least one reader disagrees with the published mapping on more than 2 of 21 decisions, on the binary covered/not |
both readers disagree on ≤ 2 |
| P2 | Reader-vs-reader agreement on the binary is no higher than the mean of the two lead-vs-reader agreements, +0.10 tolerance | R1~R2 exceeds mean(lead~R) by more than 0.10 |
| P3 | DECIDES ≤ 1 per reader per log | any reader returns ≥ 2 DECIDES on either log |
| P4 | Log B coverage (mean over readers) does not exceed Log A coverage (mean over readers) by more than 0.15, despite JA→EN being the best-evidenced pair | the gap exceeds +0.15 |
| P5 | ≥ 50% of Log B's decisions that both readers label NONE are about a single word or a single figure (Log A: 9 of 11 = 0.818, RS-20260726e §3) |
below 50% |
P3 is close to entailed and is registered anyway, because the only candidate shaped do X rather than Y addressed to a translator is #12, so a DECIDES can essentially only come through it. What is not entailed is which decisions #12 reaches, and that is the number ARM-framework's closure statement needs.
P4 is the one this session exists to test. The prediction is a null: that giving the inventory its best pair buys nothing.
6. Failure criteria, pre-committed
- A reader returns fewer labels than decisions, labels outside the scheme, or a malformed body → the call is void; rerun once; a second failure drops that reader and the run reports on one, saying so.
- Reader-vs-reader agreement on the binary below 0.50 → no coverage figure from this run is reportable as an estimate. The run's sole output is then the reliability finding, and the closure statement must say that the project cannot measure its own coverage rather than quoting a number. This is pre-committed and is the outcome the lead considers most likely after P1.
- A reader's call costs more than $0.40 → the second reader is dropped and the run reports on one.
finish_reason: lengthwith empty content → drop that slug rather than retry it (note (b), seven sessions of evidence). Declared reserve for either reader:qwen/qwen3.7-max(config/models.md, first reserve). A reserve substitution is reported as such.
7. What this design cannot establish
- Anything about translators other than the lead. Both logs are the lead's self-report, and R04's own limitations section says decisions made without noticing do not appear. Both denominators are under-counts, in the same direction as
RS-20260726ealready declares. - Anything about the quality of either translation. No judging happens here, and the lead never judges its own translation (charter §5).
- A general JA→EN claim. One text, two spans, one translator.
- That the readers are right. A panel model is
NOT CALIBRATEDfor quality judgment; this is not a quality judgment but a classification against a written list, which is the task shape S015 found the panel strong on in the failing direction — it may be trusted to reject a mapping, not to certify one. Disagreement is therefore informative and agreement is weak evidence. Stated before the run because it constrains how the result may be read afterwards. - The ordering property
RS-20260726ehad. Log A was written before the inventory existed. Log B was written by a lead that had read the inventory that morning. The bias runs toward coverage — a lead primed on the inventory would notice inventory-shaped decisions — so P4's null is the hard direction and a rise in coverage is uninterpretable. Declared here rather than discovered later.
8. Cost
Pre-flight, built from max_tokens at list out-price plus the prompt at list in-price (note (abc)):
| call | model | max_tokens | worst case |
|---|---|---|---|
| pre-run critic | google/gemini-3.6-flash (P2) |
8,000 | $0.072 |
| reader 1 | openai/gpt-5.6-terra (P1) |
12,000 | $0.203 |
| reader 2 | x-ai/grok-4.5 (P3) |
12,000 | $0.090 |
| total | $0.365 |
Day headroom at design time: $3.493220 of $5.00 (2026-07-28, seven prior sessions). Fits. The translation limb costs $0 and is never ledgered.
9. Amendments after the pre-run critic pass (2026-07-28, before dispatch)
Critic P2 google/gemini-3.6-flash, verdict NEEDS-REDESIGN, nine findings, seven accepted, one accepted-as-limitation, one declined. Full dispositions with the reasoning: critic/dispositions.md. Sections 1–8 above are preserved verbatim, including the sentences withdrawn below, so what was withdrawn is visible.
- §5 P4 is evaluated on three Log B subsets — all 42, draft-only B1–B27, revision-only B28–B42 — because the two logs differ in size and composition and no subset is a clean counterpart to Log A's 21.
- §7's "the bias runs toward coverage … P4's null is the hard direction and a rise in coverage is uninterpretable" is WITHDRAWN. Priming changes not only which decisions are noticed as inventory-shaped but which decisions get recorded at all; the two effects push the proportion in opposite directions and neither is measured, so the null is as uninterpretable as the rise. What survives: the reliability measurement on Log A (published mapping, log written before the inventory existed) and the DECIDES count. Log B's coverage proportions are description of one log, not a replication of
RS-20260726e's figure, and the word replication is withdrawn from this design's claim about itself. - §6 failure criterion 2 is restated on Cohen's κ: coverage figures are unreportable as estimates at κ ≤ 0.20 on the binary. Raw agreement and the majority-class baseline (0.524 on the lead's Log B base rates) are reported alongside. The original 0.50 raw-agreement criterion sat at chance and is withdrawn.
- New P6, registered before dispatch and capable of defeating the closure this arm intends: if either reader returns DECIDES ≥ 2 on Log B through a candidate other than C12, the inventory holds prescriptive-about-translating content the lead's mapping missed, and
ARM-frameworkmust be extended, not closed. - The reader instruction gains one clause (rule 2): a candidate that merely lists options is INFORMS even where one listed option fits; DECIDES requires the candidate to do the ranking.
- §4's blinding claim is corrected. There is no pair blinding: the source strings identify both pairs outright.
- P3 is restated as a check on reader discipline, not on the inventory.
- P5 is WITHDRAWN. "About a single word or a single figure" has no operational definition, no assigned classifier and no independent check.
RS-20260726e§3's grain finding is not re-tested here. - Analysis and verifier join on an explicit
D<n>↔B<n>map with assertions on both key sets, and report three-way agreement alongside the binary. - §6 criterion 3 rises to $0.50 per reader call; a
finish_reason: lengthbody containing all 63 parseable lines is accepted rather than voided.