Repository path: workshop/experiments/E-20260729d-decision-grain/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260729d-decision-grain |
| status | frozen |
| created | 2026-07-29 |
| updated | 2026-07-29 |
| senses | cultural-mediation, style-correspondence, consistency |
| provisional | true |
| internal-judgment-only | true |
| links | workshop/experiments/E-20260729d-decision-grain/candidates.md, workshop/translations/patsyuk/R04-v1/translation.md, framework/closure.md, framework/traceability-inventory.md, wiki/findings/results/RS-20260728i-coverage-independent.md, wiki/arms/ARM-decision-grain.md, config/models.md |
E-20260729d — is the framework's zero a fact about its candidates, or a fact about translation decisions?
Frozen before dispatch. C15 and C16 were committed at 0579229; the source and census at 4b4a647; the translation and its log at 152e849. This design was written after all three and before any call.
1. The question
framework/closure.md §1.4 holds the sharpest thing the project knows about its own framework: DECIDES 0 of 126. Two independent readers, 63 logged translation decisions from two language pairs, all fourteen candidate recommendations — and not one classification says any candidate determines what to write. Two explanations are live and this project has never separated them.
- (i) The candidates are the wrong shape. They are taxonomic, diagnostic, descriptive, or about a whole revision pass. Write one at the grain of a single decision, in the imperative, with checkable conditions, and it will decide.
- (ii) A translation decision is not the kind of thing a written rule decides. The zero is a property of the object, and no rule at any grain would move it.
If (i) is right, a release has a route that does not wait on Tier D. If (ii) is right, framework/closure.md §6 item 3 — "a prescriptive candidate that is not about a whole pass" — is not a gap to be filled but a category error, and the closure statement gets stronger and bleaker.
2. What is new, and the control that makes it interpretable
C15 is the treatment: a four-test ordered procedure for handling a culture-bound item, every clause traceable to a site in a second-read Tier 1 precedent anchor (X1a), written to be as decidable as the evidence permits. Its warrant table is in candidates.md and is not shown to readers.
C16 is a sham, and its existence is the design's central protection. Testing (i) by writing a good rule and finding that readers now say DECIDES would prove nothing: any procedure decides, including a worthless one, because deciding is a property of the form. C16 has C15's surface — ordered imperative tests over the same handling set, each with a reason clause — and is keyed to properties (paragraph position, syllable count, order within a sentence) that no evidence in this project connects to anything. It is fully checkable and its last test catches everything, so it never returns "no rule fits".
C16 establishes the ceiling. The interesting quantity is not C15's DECIDES count but C15's count against C16's.
3. Materials
- Classification pass.
runs/materials.md, built bybuild_materials.pyfrom stored files:E-20260728i's materials verbatim for C1–C14, Log A (21 decisions, FR→EN,T-mare-au-diable-R04-v1) and Log B (42 decisions, JA→EN,T-takasebune), plus C15 and C16. The only difference fromE-20260728iis the two new rows, so this pass is a direct replication with one variable changed, and the 63 decisions were all written before either candidate existed. - Applicability pass.
runs/prescribe-C15.mdandruns/prescribe-C16.md, built bybuild_prescription.pyfrom the frozen census (census.json, 23 culture-bound item types, 50 occurrences) and the frozen source. Each gives, per site: source form, literal gloss, paragraph, ordinal position among that paragraph's sites, occurrence count, and the Russian sentence. No English rendering of any site appears and the lead's translation is not in these materials. The two files are identical apart from the rule.
A design change made before dispatch and recorded rather than hidden. T-patsyuk-R04-v1's log was going to be a third log in the classification pass. It cannot be: the translation was made under C15, and the log names C15 by id at fifteen of its twenty-nine entries. Including it would have told the readers which candidate was the treatment. It is excluded, and the classification pass therefore runs on pre-existing decisions only — which is strictly better, because it removes the compliance confound from the primary measurement altogether.
4. Procedure
Roles (config/models.md): readers P1 openai/gpt-5.6-terra and P3 x-ai/grok-4.5 — the same two readers as E-20260728i, deliberately, so the baseline is the same instrument. Critic P2 google/gemini-3.6-flash, a subject in nothing here. Anthropic models are excluded from every role (charter §4).
- Pre-run critic — one call, P2, this design plus both candidates plus both materials sets. Findings dispositioned in writing in
critic/dispositions.mdbefore any reader call. - Classification pass — 2 calls (one per reader),
runs/materials.md,E-20260728i's reader prompt with the candidate count changed from fourteen to sixteen and nothing else. 63 lines each. - Applicability pass — 4 calls (2 readers × 2 rules), separate calls so no reader ever sees the two rules together; the contrast between a warranted rule and a groundless one is exactly what would reveal the sham. 23 lines each: prescribed handling, and which numbered test returned it.
temperature: 0, reasoning: {"effort": "low"}, generous max_tokens per note (bdl). Raw bodies preserved. Judgment is not parallelised across a single decision.
5. Predictions, registered
| prediction | |
|---|---|
| P1 | At least one reader assigns DECIDES to at least one of the 63 entries via C15. The project's standing count is 0 of 126. |
| P2 | C16's DECIDES count is greater than or equal to C15's. If deciding is bought by form, the sham should do at least as well, since it reaches every item and never abstains. |
| P3 | On the applicability pass, inter-reader agreement is higher for C16 than for C15. C16's tests are mechanical; C15's require judging "load-bearing or furniture" and "exact equivalent", which two readers can read differently. |
| P4 | C1–C14 receive DECIDES zero times, replicating RS-20260728i on the same 63 entries with the same two readers. |
| P5 | C15's prescriptions match the translator's actual handling on more than 17 of 23 sites — the translator complied by construction and the log records only 3 overrides and 3 unfollowable sites, so anything below that means the readers cannot reproduce the rule's own output. |
P5 is a compliance check, not an agreement measurement, and is labelled that way wherever it is reported.
6. Failure criteria, registered
- F1 — reliability floor. If reader-vs-reader agreement on the three-way classification label falls below κ 0.20, no count from the classification pass is reportable as an estimate. Same threshold as
RS-20260728i. - F2 — the uncomfortable outcome. If both C15 and C16 receive zero
DECIDESfrom both readers, explanation (ii) gains and explanation (i) loses: procedural form is not the binding constraint on the framework's zero. C15 does not enterframework/traceability-inventory.md, andframework/closure.md§6 item 3 is rewritten to say so. - F3 — decidable and wrong. If C15 receives
DECIDESbut its applicability agreement with itself (P3) is at or below C16's, C15 is a procedure and not a recommendation, and does not enter the inventory. - F4 — the closure-defeating outcome, registered so this design is not arranged to confirm what the session expects. If C15 receives
DECIDESfrom both readers on ≥ 2 entries and out-agrees C16 on the applicability pass, thenframework/closure.md§1.4's zero is not a property of translation decisions, the arm must produce a claim page rather than a null, and the closure statement's "the cheapest route to a release is Tier D" is wrong for the first time since it was written.
The session's own expectation is F2 or a split. F4 is reachable and is what would cost the most to admit; it is registered for that reason (the RS-20260728i P6 precedent).
7. What this cannot establish, declared before the run
- Nothing about quality. No jury scores anything; Tier D is NOT PASSED; the lead never judges its own translation. Whether following C15 produces a better translation is not asked and could not be answered here.
- Nothing about translators other than the lead. Every log in the classification pass is the lead's self-report, as
RS-20260726eandRS-20260728iboth already declare. - The readers are uncalibrated for quality and are used for classification, the task shape
config/models.mdrecords the panel as strong on in the failing direction. Their disagreement is informative; their agreement is weak evidence. - C15 is the lead's synthesis of anchors the lead built. Its warrant is
X1a, but the reading that produced it is not independent of the reading that produced the anchors. - The sham is a knowingly false statement placed in a prompt. It stays inside this experiment, is declared in
candidates.md, and must never enter the inventory.
8. Budget
Worst case built from max_tokens at list out-price plus prompts at list in-price (note (abc)): critic ≈ $0.11, classification ≈ $0.23, applicability ≈ $0.15. Worst case ≈ $0.49. UTC day 2026-07-29 stands at $0.899855719 of $5.00 before this session; headroom $4.100144. Fits with wide margin.
9. Amendments, 2026-07-29, after the pre-run critic and before any reader call
Full reasoning in critic/dispositions.md. Verdict NEEDS-REDESIGN, five findings, four accepted and one declined in writing.
C17, a positive control, is added to the classification pass (finding A4). The design had no evidence that theDECIDESlabel can be returned at all by this instrument, and the entire session rests on reading a zero.C17decides with no judgment at any entry recording live alternatives, and is obviously bad advice. New prediction P6:C17receivesDECIDESfrom both readers on ≥ 5 of the 63 entries. New failure criterion F5, which supersedes F2: below that on either reader, theDECIDESlabel is not a functioning measurement, no inference from any zero here or inRS-20260728iis available, and the instrument failure is the session's result.- §2's ceiling claim is narrowed (A1). C16 is the ceiling for a mechanically decidable procedure, not for any procedure; nothing here separates procedural form from algorithmic determinism, and P2's reading is restricted accordingly.
- The applicability pass's primary statistic is reader-versus-reader agreement (A2), which is not circular. Reader-versus-translator is a descriptive third number, reported with the circularity named in the same sentence.
- P5 is relabelled (A3): reader-versus-translator reproduction of the rule's output, not a compliance check.
build_prescription.pyno longer prints(the FIRST in its paragraph)(B). It pre-computed C16's test-1 trigger and gave C15 nothing comparable, tilting the very agreement figure P3 compares. The bare ordinal stays.- F4's second conjunct becomes an absolute threshold (C): C15's reader-versus-reader agreement ≥ 0.60 raw, rather than out-agreeing a near-ceiling mechanical rule.
verify.pygains a test-to-handling coherence check (D), with the admissible pairs fixed incritic/dispositions.mdbefore any output exists; incoherent rows are excluded with the count stated.- Finding E is declined, with the reason written: the sequence is enforced by public git commit order, which is this project's mechanism throughout and is adequate against the threat model (the lead's own optimism).
Budget, restated. One further candidate lengthens the classification prompt by ~90 words. Worst case ≈ $0.50. Critic actual: $0.0334875.