Repository path: framework/tierD-repaired-rules.md · rendered 2026-09-09
Page metadata (front matter)
| type | note |
|---|---|
| id | tierD-repaired-rules |
| status | active |
| created | 2026-07-29 |
| updated | 2026-07-29 |
| senses | accuracy, cultural-mediation, literary-quality, naturalness, style-correspondence, voice |
| internal-judgment-only | true |
| links | wiki/arms/ARM-tierD-repair.md, wiki/arms/ARM-tierD.md, wiki/findings/results/RS-20260727b-tierD-rules.md, wiki/findings/results/RS-20260728g-scale-usage.md, wiki/findings/results/RS-20260729c-neutral-summary.md, wiki/decisions/resolved/D-20260725-07-athenaeum-1906-condition-ii.md, config/models.md, framework/closure.md, framework/control-arm-spec.md |
The repaired Tier D rules — import this block; do not re-derive it
ARM-tierD-repair step 5, the arm's closing artifact. This page exists so that the next
Tier D design does not have to reconstruct three sessions of repair work from three result
pages. It is a specification, not a result, and it calibrates nothing: Tier D is NOT PASSED
and this page does not change that.
Charter §5.5 wording discipline binds everything here, in the form S050's critic tightened it to: every statement must be expressible as "on these stored scores, statistic S at threshold T would have decided D." A validation-shaped sentence that avoids the banned words is still not licensed.
R1 — the sham band (condition (i), repaired S040)
Do not import §6.5 as written. Under the design's own null its lower branch fires with probability 0.533936 against the upper branch's 0.004639 — a 115-fold asymmetry, and a branch that fires more often than not on a jury doing nothing.
Repaired form. State the lower branch on N₋ ≥ 5 of 6, matched at 0.004639 by construction. Do not collapse the two branches into one conclusion: an upper firing means detection is confounded with edit-presence; a lower firing means the sham materials were not neutral and the false-alarm rate is unmeasured. These license different sentences.
And size the sham for power, which nobody had computed. 5-of-6 on 6 units catches a genuine 70% edit-presence bias 0.42 of the time. A working sham needs ≈15 units (5 items), not 6. S020, S033's critic and S034's 71-check verifier all asked whether the rule fires wrongly; none asked whether it fires at all.
R2 — the sham must be the same kind of edit as the operator (S040, defect (iii))
0 of 16 sham sites change the word count against 10 of 24 targeted sites (+44 words, ≈4% per item), and 6 of 16 sham sites change no lexeme at all against 0 of 24 targeted. A floor measured on edits that cannot change length does not bound the false-alarm rate of edits that do. Any future sham must be matched to the operator on edit kind, and the match must be declared in the frozen design.
Open and returned to the backlog, not answered here: whether the length difference is a cue
a juror actually uses (ARM-tierD-repair steps 3 and 3b, never run).
R3 — the stage-1 scale-usage gate (condition (ii), UNREPAIRABLE, S050)
There is no threshold. Do not import §10's gate with a tuned number.
All six (run × juror) cells clear the 0.75 criterion margin on their targeted arm (realised margins 1.500 to 4.083); the lowest observed sham dispersion is 0.143. Any threshold above it fails a demonstrably capable juror; any threshold at or below it passes everybody. The ordering by sham dispersion is not monotone in the realised margin, in either run. One cell in six passes the inherited 0.75, and it is the cell most inflated by sense-level offset (0.866 → 0.617 centred; under the centred statistic none of the six passes).
A gate with no discriminating content is a decision-shaped object, not a strict gate.
Its replacement, settled here because step 5 owed a decision and leaving it open would leave
the arm open. S050 named two candidates: a prior positive control on a known-difference
pair, or a posterior max-abs-d / mean-abs-D threshold on the arms that matter. Take the
prior positive control. The gate's entire function was to be prior — to establish before
stage 2 that the instrument can separate anything at all — and a posterior threshold cannot do
that however well it performs. The materials exist: T-bargamot-R04-v1 Unit B plus
variant-F1, eight sites, two of each of O4's four documented failure types, matched to the
operator at +4.54% length and 4 of 8 word-count-changing sites.
This is a specification choice, not a ratified decision. It is recorded with its reasoning
so that a future design can adopt it by citation or overturn it by argument; if a design leans
on it hard it should go through the ratification protocol first, alongside
framework/control-arm-spec.md, whose backlog row names "the next design that builds a paired
comparison" as its trigger. E-20260729c built a paired comparison and did not route that
ratification — declared here rather than left for a reader to notice.
R4 — the held-out arm's materials (condition (iii), DISCHARGED S055)
The Garnett/Hapgood pair on the Memoirs of a Sportsman cycle is the project's only pair clearing conditions (i), (ii) and (iii) on any sense. What binds:
- Admissible for
accuracyandcultural-mediationonly. Reproduced three times out of three on evidence the lead did not write (RS-20260729c). This is the robust clause. - Inadmissible for
naturalness,literary-quality,style-correspondence,voice. Same reproduction. - Scoped to the Memoirs cycle — and this clause is the weak one. It reproduced on one of two independently authored non-lead summaries; on the other, the routed vote licensed A House of Gentlefolk as well. A design that needs clause 3 should re-open it as a decision rather than cite it.
- The evidence base for clauses 1–2 is thinner than the decision page shows, in one specific way: all eighteen paired extracts the 1904 review prints are cited to A Nobleman's Nest, and the review prints no style exhibit from the Memoirs at all. The exclusion in clause 2 rests, in printed exhibits, entirely on the work clause 3 refuses to license. This was drawn by no non-lead voice and by exactly one lead-written one; it is recorded here because a future design will otherwise inherit clause 2 as though its evidence were work-matched.
R5 — three constraints that survive from the closed arms and are easy to lose
- The S034 verdict is not reopened. Repaired rules apply to future runs only; retrospective application is a labelled diagnostic with no verdict authority.
- A gate is prior; a detection result is posterior. "The gate would have blocked a run that
then worked" shows conservatism, not invalidity.
ARM-tierD-repairmade this mistake once, at S050, and had it refuted before running. lowgrowingin Hapgood HB is an unrepaired OCR artefact, deliberately, and is a named confound onH-HB. A future design may authorise the repair by name, before the freeze.- The dose finding is live and cheap: the accuracy margin was +2.04 at both 3 and 8 sites
and only
naturalnessseparated them; an independently drawn 3-site set would turn one deterministic subset into a sample at roughly $0.35. It belongs to the run, not to the repair.
What the repair does and does not change
Changes: the sham's decision rule, its size, its edit-kind matching, the stage-1 gate's status (from "a threshold to be tuned" to "no threshold exists"), and the licence attached to the only qualified materials.
Does not change: the calibration state. TIER D REMAINS NOT PASSED. Nothing on this page is evidence that any jury detects anything. What it buys is that the next run's failures will be new ones.
Required pre-check, absorbed from wiki/backlog.md at S067
Before any future Tier D run, establish whether the sham/operator LENGTH MISMATCH is a cue that is
present, and then whether it is a cue a juror uses. S040 measured the mismatch — 0 of 16 sham sites
change the word count against 10 of 24 targeted sites (+44 words, ≈4% per item) — and
ARM-tierD-repair returned the question unrun when it closed.
- (a) Free, local, one function call. Run
tools/metric_a.py's Metric A over the Tier D sham and targeted materials, all five cues, to establish whether the length cue is there. This is a precondition for (b) and there is no reason to defer it. - (b) Costs money, one cell. A panel probe asking whether a juror separates two texts differing only by ≈4% length with no semantic change.
S041 showed that there and used are independent properties, which is why (a) does not answer (b) and (b) without (a) is uninterpretable. This block lives here rather than in the backlog because the design that needs it will read this page and may never read that one.