Repository path: workshop/experiments/E-20260730-grain-clause/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260730-grain-clause |
| status | frozen |
| created | 2026-07-30 |
| updated | 2026-07-30 |
| senses | cultural-mediation, style-correspondence, consistency |
| provisional | true |
| internal-judgment-only | true |
| links | workshop/experiments/E-20260730-grain-clause/variants.md, workshop/experiments/E-20260730-grain-clause/census.json, workshop/experiments/E-20260729d-decision-grain/design.md, wiki/findings/results/RS-20260729d-decision-grain.md, workshop/translations/kiseru/R04-v1/translation.md, wiki/arms/ARM-decision-grain.md, framework/closure.md, framework/traceability-inventory.md, config/models.md |
E-20260730-grain-clause — is C15's irreproducibility a property of one clause's wording, or of warrant as such?
ARM-decision-grain step 2, and the arm's second and last declared session.
1. The question, and why the arm's own prescription is not enough on its own
RS-20260729d-decision-grain measured a trade: a warranted decision-grain rule (C15) is applied identically by two readers at 14 of 23 sites (κ 0.452), a groundless mechanical one (C16) at 22 of 23 (κ 0.933). §2 then localised eight of the nine disagreements in one clause — test 3's "an exact equivalent from a practice the two cultures share" — where one reader sets the bar loosely and the other tightly.
ARM-decision-grain step 2 prescribes: operationalise that clause and re-run the applicability pass. Taken literally, that experiment cannot answer the question it is asked for, and saying so is the first thing this design does.
- A rise in agreement is already predicted by S056's own result.
C16showed that making test conditions mechanically checkable raises agreement to near ceiling. Any rewrite of test 3 that removes a judgment will therefore raise κ, whether or not it keeps the clause's warrant. A one-condition re-run would measure the thing S056 already measured. - A fall or a null is uninterpretable without knowing whether the pass repeats.
RS-20260728f(S049) found a byte-identical repeat of one reader condition moving agreement by +0.175 — larger than the effect this session is looking for. Without a test–retest control, any difference is unattributable.
So the step is run with two additions that make it answerable, and one that widens it.
2. The three arms
Arm R — Russian, the primary. S056's 23 frozen Gogol sites, unchanged, under three rules:
| condition | what it is |
|---|---|
| C15 | S056's prompt file copied byte-for-byte, asserted identical in build_prompts.py. The test–retest control, and after critic finding 2 it is dispatched twice to each primary reader, so that three observations of one identical request exist: S056's and two of this session's. |
| C15′ | the same file with test 3's second clause replaced by the operationalised condition ARM-decision-grain step 2 specified at S056. |
| C15″ | the same file with test 3's second clause replaced by a matched groundless condition — same dictionary lookup, criterion nothing in this project's evidence connects to handling, matched to within one word in length. |
Readers P1, P3 (the same two as S056 and E-20260728i, so the baseline is the same instrument) and P5 as a third voice, which discharges RS-20260729d revision trigger 1. The primary statistic is P1~P3, because that is the pair 0.609 / 0.452 was computed on; P5's numbers are new data, not a retest, and are reported separately.
Arm A — the anchor sites, and it is what makes the word "warranted" checkable. 17 culture-bound items taken from the cultural-mediation tables of the two second-read precedent anchors (A-garnett-vanka, A-shaw-spider-thread) — the pages test 3 was derived from — restricted to sites where neither test 1 nor test 2 applies, so each is decided by test 3 against test 4 and nothing else. Three sites are excluded on that ground and named in build_anchor.py.
The measured quantity needs no key, and after critic finding 4 it is not set overlap. All three rules share test 3's first limb verbatim, so every established-borrowing site falls in every rule's firing set and any Jaccard between those sets is inflated toward 1 whatever the second limb does. The quantity is instead the set of sites at which a variant flips C15's decision:
D(rule) = { site : the variant fired test 3 where C15 did not, or C15 fired test 3 where the variant did not }, computed per reader from the readers' own outputs.
D is zero when a rewrite preserves the clause's extension and grows as it departs from it, and the shared limb contributes nothing to it. C15′ is a warrant-preserving rewrite only if D(C15′) is small; C15″ is a control only if D(C15″) is not. Three rules × P1, P3. Every arm-A prediction is assessed per reader and holds only if it holds on both (finding 5) — averaging would let one reader carry a prediction the other refutes.
Arm J — Japanese, and it is the translation limb's wire. 34 culture-bound sites from T-kiseru-R04-v1's frozen census (Akutagawa 「煙管」 一–二), under C15 and C15′, P1 and P3. RS-20260729d §6 says explicitly that its finding rests on "one rule, on one passage, in one pair". This is a second passage in the second pair C15 declares itself evidenced on, chosen for a specific reason:
English holds many established borrowings from Japanese, and Russian and English share a great deal of ordinary practice. Test 3 has two limbs — an established borrowing (checkable) and an equivalent from shared practice (the one that carries the disagreements) — and a Japanese passage should load them in different proportions.
What arm J can and cannot do, rewritten after critic finding 6. The arm changes the pair, the passage, the site count, the label distribution and the instruction wording at once. No pair-attributive claim is available from it, and the design pre-registers that none will be made; the design's first draft said "arm J is where it shows" and that sentence is withdrawn here rather than after the fact. What survives is the part that is internally valid and is the part worth having: C15 against C15′ on the same prompt shell, with only the rule differing — an independent replication of the treatment effect on a second site list in a second pair. Q6 and Q7 are demoted to descriptive.
The wire, in one sentence. The Japanese passage was translated so that its site census could supply a second, independent site list on which the repair to test 3 either replicates or does not.
3. Materials, and the freeze chain
| commit | what was frozen |
|---|---|
9830845 |
the source text of 煙管 一–二 with its one gaiji emendation, and variants.md — C15′ and C15″ |
4f0134d |
census.json — the 34 Japanese sites, glosses written before any of them was translated |
6168fc2 |
T-kiseru-R06-v1, the single-pass draft, and its contamination gate |
0dd4071 |
T-kiseru-R04-v1, the revision and its 31-point log |
| this commit | this design, build_prompts.py, build_anchor.py, and all eight prompt files |
The order matters and is the design's protection against two different kinds of tailoring. The rule variants precede the census, so the census cannot have been built to make C15′ resolve crisply. The census precedes the translation, so the translation cannot have chosen its own sites. C15′'s wording is S056's, not this session's (variants.md §C15′), which is the only protection available against tailoring the fix to the nine disagreements this session has read.
Enumeration rule for the Japanese census, fixed before the list was written: every item satisfying C15's own definition of a culture-bound item, and nothing else. Nine items are excluded by that definition and are listed in build_census.py so the boundary is auditable. Glosses are descriptions, never candidate renderings, and none contains a romanisation of its own item — the S056 confound (RS-20260729d §7), here checked mechanically rather than declared.
4. Procedure
Roles (config/models.md): readers P1 openai/gpt-5.6-terra, P3 x-ai/grok-4.5, P5 deepseek/deepseek-v4-pro; critic P4 moonshotai/kimi-k3, a subject in nothing here — a reader may not critique the instrument it is about to be, which is the S053 role-collision fix. Declared reserve for every seat: P2 google/gemini-3.6-flash. Anthropic models excluded from every role (charter §4).
- Independent pre-run critic — one call: this design,
variants.md, and all six distinct materials files. Findings dispositioned in writing incritic/dispositions.mdbefore any reader call. - Arm R — 9 calls (3 rules × 3 readers),
max_tokens9,000, one rule per call so no reader ever sees two rules together. - Arm A — 6 calls (3 rules × 2 readers),
max_tokens6,000. - Arm J — 4 calls (2 rules × 2 readers),
max_tokens6,000.
temperature: 0, reasoning: {"effort": "low"} — the same settings as S056, because the arm-R C15 condition is a retest and a changed setting would not be one. Every call is stateless and independent; raw bodies preserved; judgment is not parallelised across a single decision.
5. Predictions, registered
| prediction | |
|---|---|
| Q1 | The pass repeats, on two counts. On each of the two byte-identical C15 dispatches, P1~P3 raw agreement is within 0.09 of S056's 0.609 (within 2 sites of 23), and κ is within 0.15 of S056's 0.452, and the two in-session dispatches agree with each other to within 0.10 of κ. |
| Q2 | C15′'s κ on arm R exceeds C15's by ≥ 0.15. Assessed only if F1 passes; the threshold is smaller than the +0.175 a byte-identical repeat moved agreement at S049, which is why this session measures its own retest variance rather than asserting the threshold (finding 9). |
| Q3 | The control fires: C15″'s κ on arm R is within 0.10 of C15′'s, or higher. If mechanisation rather than aptness buys reproducibility, the sham clause should do as well as the operationalised one. |
| Q4 | On arm A, ** |
| Q5 | On arm A, C15″ fires test 3 at ≥ 1 of the four proper-name sites (A9, A12, A13, A17) at which C15 fires at none. A rule that prescribes an English equivalent for a dog's name because the equivalent is spelled shorter is doing something the anchored evidence never did. |
| Q6 | (descriptive, no prediction assessed — finding 6.) Arm J's C15 raw agreement and κ are reported beside arm R's, with no pair attribution. |
| Q7 | (descriptive, no prediction assessed — finding 6.) The proportion of sites at which each test is cited, per arm, per rule. |
| Q9 | New, and it is arm J's actual prediction: on arm J the direction of the C15 → C15′ change in κ is the same as on arm R. A replication, inside one instrument, of whatever arm R finds. |
| Q8 | No reader flags UNSURE anywhere, replicating RS-20260729d §1's completeness finding. The NORULE half of this prediction is struck (finding 1): test 4's furniture branch returns a handling for anything, so NORULE is structurally unreachable and could only have fired on reader non-compliance, which F5 already covers. |
The session's own expectation is Q2 and Q3 both holding — a κ rise that the sham clause matches, i.e. RS-20260729d §6 item 3 surviving in a stronger form. F3 below is the outcome that would cost the most to admit and it is registered for that reason.
And a registered discount on the expected outcome (finding 10). The treatment's wording was fixed by S056; the control's was written this session, by an author who had already read the nine sites C15 disagrees on. Q3 holding is therefore weak evidence for the mechanisation reading, because the session had the freedom to make C15″ crisp exactly where C15 is loose. Any closure under F2 must carry that sentence.
6. Failure criteria, registered
- F1 — the retest floor, and it is the gate on everything else. If any clause of Q1 fails — raw agreement outside ±0.09 of 0.609 on either dispatch, κ outside ±0.15 of 0.452 on either dispatch, or the two in-session dispatches more than 0.10 apart in κ — the applicability pass is not stable enough to support a between-rule comparison of the size being looked for, and no κ difference between rules is reportable as an estimate. The instrument instability is then the session's result, and it is a hard one: it would mean
RS-20260729d§1's headline numbers are not repeatable either. Precedent that this can happen:RS-20260728f, +0.175 on a byte-identical repeat. Any reserve substitution in arm R voids retest status for that reader's conditions (finding 8), which are then reported as new data like P5's. - F2 — the control fires (the expected outcome). If Q3 holds, reproducibility is bought by mechanisation and not by the operationalisation's aptness. C15′ does not enter
framework/traceability-inventory.md,RS-20260729d§6 item 3 stands in a stronger form, and the arm closesresolvedon it. - F3 — the arm-defeating outcome. If all three of: Q2 holds; Q3 fails (C15″ more than 0.10 below C15′); and Q4 holds with |D(C15′)| ≤ 2 of 17 on both readers while |D(C15″)| ≥ 4 on both — then a decision-grain rule can be made materially more reproducible without giving up the clause's warrant.
RS-20260729d§6 item 3's trade would then be a property of one rule's wording rather than of warrant as such, the arm produces a claim page rather than a null, and entry of C15′ into the inventory becomes a decision routed through the ratification protocol, not this session's call (ARM-decision-grainconstraint 5). And even then it carries a caveat: a C15″ κ below C15′'s is partly explicable by confound 9 (romanisation schemes differ, so C15″ has an irreproducibility source of its own that has nothing to do with warrant), so F3 firing licenses the claim only with that caveat stated in the same sentence (finding 7). - F4 — completeness, re-scoped (finding 1). If either primary reader flags
UNSUREat more than 2 sites in any condition, that rule's wording is not decidable on that site list and its agreement figures are not comparable with the others'; the condition is reported and excluded from the comparison. (As first registered this criterion countedNORULE, which test 4 makes structurally unreachable — it could not fire, and the critic said so.) - F5 — parse floor. A call not returning exactly the required number of lines, with labels drawn from the eight plus
NORULE, falls through once to the declared reserve P2 and is never retried on the same slug (notes (b), (bdl)). A condition left short of a reader is dropped with the count stated. - F6 — no silent arm. Every arm's numbers are reported whether or not they support the session's expectation, and any arm withheld under F1/F4 is named in the result page's headline rather than in a footnote.
7. What this cannot establish, declared before the run
- Nothing about quality. No jury scores anything, Tier D is NOT PASSED, and the lead never judges its own translation. Whether a translation made under C15′ would be better than one made under C15 is not asked here and could not be answered.
- Nothing about translators other than the lead. Arm A's site list, its glosses, and the anchors it is drawn from are all the lead's work; arm J's census and translation likewise.
- The readers are uncalibrated. They are used for classification, the task shape
config/models.mdrecords the panel as strong on in the failing direction. Their disagreement is informative; their agreement is weak evidence. - Whether C15′ is a good rule. It is
untestedin the charter's sense (rule 4). It is not more warranted than the clause it replaces — both rest on the same single anchored site — and arm A measures only whether it reaches the same places. - Whether any of this generalises past three site lists in two pairs.
8. Declared confounds
- The manipulation is incomplete, and this is the sharpest one. The word exact survives in the EQUIVALENT handling label's own definition — "use an established English borrowing of the item, or an exact English counterpart from a practice both cultures share" — which is identical in all conditions because changing it would have altered more than one clause. A null on Q2 is therefore consistent with the residual rather than with the clause being innocent — and, after critic finding 11, so is an attenuated positive: the caveat is owed at every Q2 outcome, not only at a null, and must be offered wherever Q2 is reported.
- Arm J's instruction block is not byte-identical to arm R's: it names a different passage, a different language and a different site count. Q6 compares two instruments that differ in those respects.
- The three site lists differ in size (23 / 17 / 34) and in label distribution. κ is chance-corrected, raw agreement is not, and both are reported.
- The lead read
C15andRS-20260729d's site-level disagreement table before writing anything in this experiment. C15′'s wording is S056's; C15″'s is this session's, so the control is the arm the lead had freedom over, which is the direction that makes F2 easier to obtain. Named because it cuts against the expected outcome. - Two arm-A glosses approximate the published rendering (A11, A16) because a calque's description is the calque. Both are sites where the comparison is between rules rather than against a key.
- Arm A's published-handling key is the lead's coding of the lead's own anchor pages, is
internal-judgment-only, and no primary statistic uses it. - P5 was not a reader at S056, so its arm-R numbers are new data and not part of the retest.
- C15″ has an irreproducibility source of its own, and the design had not seen it (critic finding 7). Its test 3 asks the reader to count the letters of "a romanisation of the source item", and romanisation schemes differ — хата as khata or hata, 煙管 as kiseru. Two readers can therefore disagree under C15″ for a reason that has nothing to do with mechanisation versus aptness, which is exactly the F3 pattern. Pre-committed: a C15″ κ below C15′'s carries this caveat in the same sentence and does not on its own license F3.
- Carry-over between calls is not controlled by design but by statelessness: each call is a fresh request with one rule; there is no conversation state to carry. The same model answers three conditions, which is what makes the retest possible and also means a model-specific reading of test 4 is shared across conditions.
9. Budget
Worst case built from max_tokens at list out-price plus prompts at list in-price (note (abc)), with P5 priced at the worst plausible routed provider rather than at list (config/models.md pricing caution, ~3.8× on one measured call):
| stage | calls | worst case |
|---|---|---|
| pre-run critic (P4, 16,000; +P2 reserve) | 1 | ≈ $0.45 |
| arm R (3 × P1/P3/P5, 9,000) plus the second C15 dispatch to P1 and P3 (finding 2) | 11 | ≈ $0.95 |
| arm A (3 × P1/P3, 6,000) | 6 | ≈ $0.41 |
| arm J (2 × P1/P3, 6,000) | 4 | ≈ $0.30 |
| total | 22 | ≈ $2.11 |
UTC day 2026-07-30 stands at $0.00 of $5.00 before this session — a fresh day; headroom $5.00. The historical band for output-dominated runs here is 15–34% of worst case, which puts the expected actual at $0.32–$0.72. The critic's own line came in at $0.139212, 31% of its declared chain worst case. Recorded because the estimate is the thing note (abc) fires on, and it has fired on an estimate rather than a spend twice (S056, S060).
10. Amendments, 2026-07-30, after the pre-run critic and before any reader call
Full record and reasoning: critic/dispositions.md. Verdict NEEDS-AMENDMENT, eleven findings — one BLOCKING, seven MANDATORY, three ADVISORY — all eleven accepted, one narrowed in writing (finding 2, two extra calls rather than six, with the reason given).
- Arm A's decisive measurement was rebuilt (finding 4, the BLOCKING one). Jaccard between test-3 firing sets is inflated toward 1 by the limb all three rules share, so
F3'sJaccard ≥ 0.80clause was close to unfailable. Replaced by D(rule), the set of sites where a variant flips C15's test-3 decision. Q4 and F3 rewritten. - The C15 condition is dispatched twice per primary reader (finding 2), so F1 gates on two in-session observations rather than one.
- F1 gains a κ clause (finding 3), because Q2 is a κ claim and F1 was gating raw agreement.
- F4 and Q8 are re-scoped from
NORULEtoUNSURE(finding 1):NORULEis structurally unreachable, so F4 as registered could not fire. - Arm A predictions are assessed per reader and hold only if they hold on both (finding 5), which is stricter than the critic's own averaging fix.
- Q6 and Q7 are demoted to descriptive and a pair-attributive claim is pre-registered as unavailable (finding 6); §2's "arm J is where it shows" is withdrawn, and new Q9 — that arm J replicates arm R's direction — is what arm J is actually predicted to deliver.
- Confound 9 is added (finding 7): C15″'s letter-count criterion depends on a romanisation choice, so the control has an irreproducibility source of its own; F3 now requires the caveat.
- A reserve substitution in arm R voids retest status for that reader (finding 8).
- Q2 is assessed only conditional on the strengthened F1 (finding 9).
- A registered discount on the expected outcome (finding 10): Q3 holding is weak evidence for the mechanisation reading because the control was the arm the lead had freedom over.
- Confound 1's caveat extends to every Q2 outcome, not only a null (finding 11).