Repository path: workshop/experiments/E-20260729d-decision-grain/critic/dispositions.md · rendered 2026-09-09
Page metadata (front matter)
| type | note |
|---|---|
| id | E-20260729d-dispositions |
| status | frozen |
| created | 2026-07-29 |
| updated | 2026-07-29 |
| links | workshop/experiments/E-20260729d-decision-grain/design.md, workshop/experiments/E-20260729d-decision-grain/candidates.md |
Pre-run critic dispositions — E-20260729d
Critic: P2 google/gemini-3.6-flash (Google AI Studio), 9,535 in / 2,558 out, stop, 15s, $0.0334875. Verdict NEEDS-REDESIGN, five findings. Four accepted, one declined in writing. Every amendment below was made before any reader call was dispatched, and the ordering is checkable in git: this file and the amended design are committed before run_readers.py and run_prescribe.py are run.
A4 — ACCEPTED, and it is the finding that changes the experiment
"Zero
DECIDESacross 126 prior calls and this run could reflect a strong model prior against assigningDECIDESunder the classification prompt instructions, rather than a fact about rule form or translation decisions. A positive control … is strictly required to prove the reader models are capable of returningDECIDESin this harness."
The design's F2 was unavailable as written, and this is the second time in this project's history that a design's central inference has been withdrawn by its own critic before a call (RS-20260728i §3 is the first). The whole session rests on reading a zero, and the design had no evidence that the instrument can return anything else. Neither did RS-20260728i, whose zero this session set out to explain — so the defect is inherited, not introduced.
Amendment 1. A positive control, C17, is added to the classification pass.
C17. Where more than one rendering is live, take the shorter one senses: naturalness, style-correspondence · evidenced on: all full statement: At any decision where the translator has more than one candidate rendering in view, count the English words in each candidate and write the one with the fewest words. Where two candidates have the same number of words, count the characters and write the shorter. Where they are the same length in both, write the one whose first word comes earlier in alphabetical order. Apply this at every such decision without exception.
C17 determines the answer at any entry that records live alternatives, with no judgment of any kind, and it is obviously bad advice — which is why it is safe: it can never be mistaken for framework content. It was written after the critic pass and before the logs were re-read, so it is not tuned to them.
Amendment 2. A new prediction and a new failure criterion, which supersedes F2.
- P6 —
C17receivesDECIDESfrom both readers on ≥ 5 of the 63 entries. - F5 — the instrument check, and it outranks every other reading in this design. If
C17receivesDECIDESon fewer than 5 entries from either reader, theDECIDESlabel is not a functioning measurement, and no inference from any zero in this design or inRS-20260728iis available. F2's reading is withdrawn in advance in that case, and the session reports the instrument failure as its result.
A1 — ACCEPTED
"If C16 gets more
DECIDESvotes than C15, the design cannot separate whether form drives the rating or whether algorithmic determinism drives it."
Correct, and §2's "C16 establishes the ceiling: what DECIDES rate is achievable by writing any procedure at all" is false as written. Amendment 3: C16 establishes the ceiling for a mechanically decidable procedure, not for any procedure, and nothing in this design separates procedural form from algorithmic determinism. P2's reading is narrowed to match: a C16 ≥ C15 result licenses "a rule that needs no judgment decides more than a rule that needs judgment" and does not license "deciding is bought by form."
A2 — ACCEPTED
"The applicability pass measures reader-author agreement on rule interpretation, not objective rule validity."
Correct. Amendment 4: the applicability pass's primary statistic is reader-versus-reader agreement on the prescribed handling — a rule two readers apply differently does not decide anything, whatever label a classification pass returns, and that number is not circular. Reader-versus-translator is demoted to a descriptive third number reported with the circularity named in the same sentence.
A3 — ACCEPTED
"P5 … is a reproducibility measurement between model readers and the translator, not a 'compliance check'."
Correct; the design's own label was wrong. Amendment 5: P5 is restated as reader-versus-translator reproduction of the rule's output, and is reported under §Amendment 4's demotion.
B — ACCEPTED
"The site prompt pre-computes and spoon-feeds the exact trigger condition for C16 Test 1 (
the FIRST in its paragraph)."
Correct and material: it tilts inter-reader agreement toward C16, which is the very quantity P3 compares. Amendment 6: the parenthetical (the FIRST in its paragraph) is removed from build_prescription.py. The bare ordinal number stays, because both rules can use it and neither is told what to do with it.
C / F4 — ACCEPTED
"Requiring C15 … to exceed C16's applicability agreement makes F4 structurally impossible to trigger."
Correct. Amendment 7: F4's second conjunct becomes an absolute threshold — C15's reader-versus-reader agreement on the prescribed handling is ≥ 0.60 raw — rather than a comparison with a mechanical rule that is near-ceiling by construction.
D — ACCEPTED
"A reader may select a valid handling string that is logically contradicted by the test number it cites."
Amendment 8: verify.py gains a cross-field consistency check. The admissible (test → handling) pairs are fixed in advance here, before any output exists:
| rule | test | admissible handlings |
|---|---|---|
| C15 | 1 | RETAIN |
| C15 | 2 | CONVERT |
| C15 | 3 | EQUIVALENT |
| C15 | 4 | SCAFFOLD, RETAIN, OMIT |
| C16 | 1 | RETAIN |
| C16 | 2 | GLOSS |
| C16 | 3 | SUBSTITUTE |
Rows violating this are counted and reported as incoherent, and are excluded from the agreement statistics with the exclusion count stated.
E — DECLINED, with the reason
"Unenforceable manual sequence constraint … no programmatic gate or cryptographic timestamp."
Declined. The sequence is enforced by git commit order and is checkable after the fact by anyone reading the repository: candidates.md at 0579229, source and census at 4b4a647, the translation and its frozen log at 152e849, and this file plus the amended design committed before either runner exists in a run state. That is the mechanism this project uses throughout and it has caught real ordering defects. A cryptographic timestamp would add assurance against an adversary; the threat model here is the lead's own optimism, against which a public commit order is adequate. Recorded as declined rather than silently ignored.