Repository path: workshop/experiments/E-20260727b-arm-identifiability/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260727b-arm-identifiability |
| status | frozen |
| created | 2026-07-27 |
| updated | 2026-07-27 |
| provisional | true |
| senses | accuracy, naturalness, voice, style-correspondence, literary-quality |
| links | workshop/experiments/E-20260724-r01r02-selfrevise/design.md, workshop/experiments/E-20260727b-arm-identifiability/critic.md, wiki/findings/results/RS-20260724-selfrevise-first.md, wiki/findings/results/RS-20260727b-tierD-rules.md, framework/traceability-inventory.md, wiki/arms/ARM-framework.md, workshop/regimes/R06-lead-single-pass.md, workshop/regimes/R04-lead-close.md |
Frozen design v2 — arm identifiability: can a regime comparison's two arms be told apart without reading?
Status: v1 written 2026-07-27 (S041) before the session's translation limb existed and before any number below was computed, except as disclosed in §0. Independent adversarial critic pass (critic.md, openai/gpt-5.6-terra, P1, non-Anthropic) returned FREEZE-AFTER-FIXES with six mandatory fixes; all six are incorporated here. v2 is FROZEN for the run.
0. Disclosure — what was already seen when this design was written
The project's rule is that a design is frozen before its numbers are seen. Part of this one was not, and the honest thing is to say exactly which part.
While checking whether E-20260724-r01r02-selfrevise had already answered this question — the standing critic disposition is that re-deriving an established finding is a process failure, so the check was obligatory — the lead printed two blocks of that experiment's runs/analysis.json:
M5_length_leakage, in full (both comparison sets, all five senses).M4_revision_magnitude, whoseper_repblock carriesd7_wordsandrev_words; nine of the twelve rep-triples were visible before the console output was truncated.
The reported preference rates (naturalness 0.76, accuracy 0.43, style-correspondence 0.39) were already known from RS-20260724-selfrevise-first, which every session reads.
Therefore, binding on this page:
- Nothing about MAIN word-count deltas is pre-registered. Every MAIN word-count finding below is reported as post hoc, and none is counted as a met prediction.
- The critic's fix 5 extends this further than v1 did. A prediction whose rationale depends on the seen pattern is not pre-registered either, even if its own number is unseen. P2 and P4 are relabelled accordingly. A rule written after seeing the result that motivated it cannot make that result confirmatory.
- Pre-registered and genuinely unseen: the whole TEMP set (D4 word counts appear nowhere in what was printed); the lead pair, which did not exist when this was written; and the demonstration in Metric C.
1. Question
In a paired regime comparison, is the treatment arm identifiable from read-free surface properties alone — and if so, at what rate compared with the preference rates the comparison reports as findings?
A read-free cue is any property of a text computable without understanding it: how many words it has, how many sentences, how many commas.
2. Why this, now
ARM-framework's declared next step is the second regime comparison — one of the two release gates in framework/README.md, and the project has had exactly one comparison, at S010. This design is the gate the second one is built through, and it is aimed at the first one.
framework/traceability-inventory.md sorts twelve candidate recommendations. One is prescriptive about translating — #12, one self-revision pass buys naturalness without moving accuracy — and it rests entirely on E-20260724-r01r02-selfrevise. The inventory records it as inadmissible pending calibration, and §4 says re-running it "needs a calibrated jury, so it is downstream of the two rebuilt Tier D controls … not of more translating." That sentence assumes the only thing wrong with the comparison is the jury.
S040 (RS-20260727b-tierD-rules §3) found something that is not about the jury. In the Tier D instrument, 0 of 16 sham sites changed the word count and 10 of 24 targeted sites did, by ≈4% per item — so the two arms of that comparison were separable by a property visible without reading either text. That was a finding about a control. The same question has never been asked of a regime comparison, and it is a different question from anything Tier D repair can answer: a perfectly calibrated jury judging arms that are surface-separable still yields a number nobody can interpret.
3. The gap this is aimed at — stated as narrowly as it is true
(Rewritten under critic fix 1. v1 stated this one step too strongly, and the narrower statement is the true one.)
E-20260724-r01r02-selfrevise did pre-register a length check. Its M5 computes, per sense, the Pearson correlation between an item's preference and that item's D7→REV word delta, and flags the sense length-cued when |r| ≥ 0.5. Its §10 threat 3 names the risk in the right words: "a juror preferring the longer/smoother passage would show revision-preference > 0.5 on every sense."
The correct claim, and the whole claim:
M5 cannot detect a purely directional (sign-only) confound. If
delta_ihas the same sign for every item, and the juror's preference is driven by that sign rather than by the magnitude, then preference is constant across items. A constant has zero variance, so Pearson's r is undefined — a 0/0 form — and in practice is computed from whatever residual variation the ties and near-misses supply. Meanwhile the arm is identifiable with certainty from the sign alone.
Three qualifications the critic forced, all load-bearing:
- r in that case is undefined, not zero. Reporting it as zero would misdescribe the arithmetic.
- Same-sign deltas can produce a large, correctly-flagged r — if preference varies with the magnitude of the delta. M5 is a valid instrument for a graded length effect and this design does not say otherwise.
- What M5 is blind to is therefore a specific thing: sign-only identifiability with preference roughly constant across items. That is the case measured here, and nothing broader.
Correlation measures a graded association. Constant-sign identifiability is a categorical property. Nothing in the S010 design, its critic pass, or its verification distinguishes them.
4. Materials
workshop/experiments/E-20260724-r01r02-selfrevise/runs/— stored and immutable: 12 rep-triples × {d7.txt,revision.txt,d4.txt},blinding-key.json,analysis.json,jury/. Two comparison sets, as that design defines them: - MAIN = REV vs D7 (treatment = the revision), n = 12. - TEMP = D4 vs D7 (treatment = the temperature control), n = 12. Cells: 2 works (kumonoito,yumejuya-1-2) × 2 translators (P1,P5) × 3 reps.- The session's own lead pair, produced under the translation limb and frozen before this analysis runs:
-
T-kusamakura-R06-v1— lead single-pass draft (regime R06, the lead analogue of R01). -T-kusamakura-R04-v1— the lead's self-revision of that draft (regime R04, whose draft is R06 by construction, exactly as R02's draft call is R01). This pair is n = 1. It is never pooled with the twelve model pairs and no rate is computed from it.
5. Instrument — frozen verbatim
Text extraction: the stored .txt files, read as UTF-8, with leading/trailing whitespace stripped and nothing else changed. A byte-level integrity check (§8 step 1) runs before any cue is computed.
Five cue families. C1 words is the single pre-specified primary cue (critic fix 3); the other four are secondary and descriptive.
Regexes are stated here outside any table, unescaped, because v1's C4 was frozen with a markdown table-escape inside it and was consequently not the expression it claimed to be (critic fix 6):
- C1
words— count of maximal runs of non-whitespace characters. Regex:\S+ - C2
chars— count of non-whitespace characters. Regex:\S - C3
paras— count of lines that are non-empty afterstrip(). - C4
sents— count of matches of the regex:[.!?]+(?=\s|$) - C5
commas— count of the literal character,
C4's splitter is crude and will miscount abbreviations. It is frozen crude deliberately: it is applied identically to both arms of every pair, so its error is symmetric, and a splitter tuned after seeing the data is what disposition 9 forbids.
For a pair (treatment T, baseline B) and cue c: delta_c = c(T) − c(B).
Metric A — arm identifiability
For a set of N pairs and cue c:
n_pos = #{i : delta_c,i > 0}
n_neg = #{i : delta_c,i < 0}
n_tie = #{i : delta_c,i = 0}
p_c = (n_pos + 0.5·n_tie) / N
A_c = max(p_c, 1 − p_c)
A_c is the accuracy of the better of the two constant read-free rules — always pick the larger, always pick the smaller — at identifying which arm is the treatment. A_c = 0.5 is unidentifiable; A_c = 1.0 is perfectly identifiable without reading a word.
Reporting, under critic fixes 2 and 3:
- Per cell (4 cells × 3 reps) is the primary form; the pooled 12-pair rate is secondary. The twelve pairs are 2 works × 2 translators × 3 reps and are not twelve independent units.
- The significance statement is a sign test conditional on non-ties —
n_possuccesses inn_pos + n_negtrials at p = 0.5, two-sided, doubled for the max over the two rules — computed only for the primary cue C1, and labelled descriptive, because the rep-triples are dependent. No p-value is attached toA*(the max over five cues) at all: maximising over cues is a selection the doubling does not cover. n_tieis reported separately. A delta of 0 is not a cue in any perceptual sense; half-credit is a statistical convention for the rate and is not a claim that anything was perceived (critic recommendation 4).
Metric B — the aggregate ceiling (a necessary condition only)
For each sense s, compare the comparison's per-sense preference rate M1_s — recomputed from runs/jury/ by this session's own code path, not read from analysis.json — with A_c for the primary cue and with A*.
M1_s ≤ A is a necessary condition for the cue rule to reproduce the jury's behaviour and is nothing more than that: it compares two aggregate rates and says nothing about which items each got right. It is reported as a screen, and the inference is carried by Metric B′.
Metric B′ — item-level agreement (critic fix 4, and it is the load-bearing one)
For each cue rule R (larger, smaller), each sense s and each juror j: the fraction of items on which R's predicted winner equals the winner j actually chose, ties in the juror's choice scored 0.5. Reported per juror and per cell.
What this licenses, exactly. High item-level agreement shows that a rule reading nothing reproduces the jury's choices, item by item — so the comparison does not identify whether the jury's preference tracks the prose or the surface. It does not show that any juror used the cue; no juror was asked, and this analysis sees only recorded choices. Any sentence of the result claiming a juror did use a cue is a failure of this design, not a finding of it (§7).
Metric C — the blindness, demonstrated numerically
Construct a worked case with the frozen arithmetic: N = 12 pairs, delta_i < 0 for all i with varying magnitude, and a juror preferring the treatment arm on every item because it is the shorter one. Report M5's r and Metric A's A_words on the same constructed data — including the case where preference is exactly constant, so that r's undefinedness (§3 qualification 1) is shown rather than asserted. The point is arithmetic, not empirical, and is reported as such.
Metric D — the minority-sign split
Partition the pairs by sign(delta_words). Report the preference rate per sense within each cell. Pre-committed: if either cell has fewer than 3 pairs, the test is declared powerless and reported as such — a null in a 1-pair cell is not evidence that length is irrelevant. (Disposition 5, applied in advance: check the test can fire before pre-registering it.)
6. Predictions, with failure conditions
- P1
[PRE]— TEMP is not read-free identifiable.A_words(TEMP) < 0.75. D4 and D7 differ only in sampling temperature; temperature is not a length instruction, so the arms should be near-unidentifiable. Falsified ifA_words(TEMP) ≥ 0.75— which would mean the S010 design's control arm is itself surface-separable, damaging the "it is the revision act, not temperature" conclusion independently of anything about MAIN. - P2
[POST-DERIVED]— the cue families agree in MAIN.A_chars(MAIN) ≥ 0.75. The number is unseen but the rationale is not: it is derived from the seen word-count pattern. Reported as a post-hoc consistency check and never as a met prediction (critic fix 5). Its content: if revision systematically shortens in words it should shorten in characters, and a divergence would mean the pattern is an artefact of tokenisation. - P3
[PRE, single case]— the lead's self-revision lengthens. For the single lead pair,delta_words > 0. Written before translating: the lead's revision habit is to add explicitness where the draft compressed, so the prediction runs opposite to whatever the model revisions do. Falsified ifdelta_words ≤ 0. Reported under its own heading, outside the prediction arithmetic; n = 1 documents a direction and estimates no rate (critic recommendation 3). - P4
[POST]— the downgrade rule. IfA(MAIN) ≥ 0.76on the primary cue — the naturalness preference rateRS-20260724-selfrevise-firstreports — and Metric B′ shows item-level agreement materially above chance, then candidate #12 inframework/traceability-inventory.mdis recorded as confounded as well as inadmissible, with the second defect noted as not discharged by passing Tier D. The threshold was chosen around a known result, so this is a decision rule, not a prediction, and is not counted as met or unmet. - P5
[PRE]— the null is reachable. IfA_c < 0.6for all five cues on both sets, the finding is that the S010 arms are not read-free identifiable, M5 was adequate, and #12's only problem is calibration. Written down before the run so it cannot be quietly discarded.
7. Failure criteria
- Instrument failure. Any cue whose value differs between arms for a reason other than the arm — trailing metadata, a truncated generation, a heading one arm kept — invalidates that cue for that pair. The §8 step-1 integrity check reports every such case; affected pairs are excluded per-cue with the exclusion named, not silently dropped.
- Insufficient data. If fewer than 10 of the 12 rep-triples have all three arms intact, no rate is reported for the affected set.
- Over-reach. Any sentence claiming a juror used a read-free cue is a failure of this design. The licensed claim is about identifiability and about what the comparison can support.
- Honest null is first-class (P5).
8. Procedure
- Integrity check. For all 36 stored translation files: byte length, UTF-8 decodability, presence of a leading heading line,
finish_reasonfrom the sibling.json. Any anomaly reported before cues are computed. - Cue extraction for all 36 files and the 2 lead files →
runs/cues.json. - Metric A for MAIN and TEMP, primary cue then secondaries, per cell then pooled; Metric B and Metric B′ against per-sense preferences recomputed from
runs/jury/by this session's own code path. - Metric C on constructed data; Metric D with its power declaration.
- The lead pair reported separately, n = 1, direction only.
- Verification — a fresh recomputation of every reported number by an independent path (
analysis/verify.py, importing nothing fromanalysis/identifiability.pyor fromtools/), discrepancies reported and not smoothed.
9. Budget
$0 for the analysis and the translation limb — everything is computed locally from stored files, and lead translation is free and never ledgered (charter §3, A4). The one spend was the independent pre-run critic pass: pre-flight worst case $0.191 built from the max_tokens cap actually sent (12,000 out at $15.00/M = $0.180, plus 4.5k in at $2.50/M = $0.011); actual $0.065876875, 34% of the worst case.
10. Threats to validity
- Identifiability is computable, not necessarily perceptible. Every number here is about what a cue can separate, not about what a reader notices. The critic asked for a skim-only task establishing perceptual availability; there is no human in this loop, and the lead cannot be the perceiver for a measurement about whether the project's jurors perceive something. Perceptual availability is unmeasured and is not claimed.
- Identifiability is not use. Every number bounds what the design licenses. It does not show any juror counted words, and no design here could.
- Post hoc status of the MAIN word-count numbers and of P2/P4, per §0 and critic fix 5.
- The cue set is not exhaustive. Five families are five. A null on these five is a null on these five.
- Whole-work lengths are dominated by the source. Only within-pair deltas are used, which is why the metric is a sign count and not a length comparison.
- The 12 pairs are not 12 independent units. Per-cell reporting is primary; every p-value is descriptive.
- The lead pair is n = 1, on different material, by a different translator, under a regime frozen this session.
Change log
- 2026-07-27 (S041) — v1 drafted, with §0 disclosure written before any new computation.
- 2026-07-27 (S041) — v2, FROZEN. Independent critic pass (
critic.md, P1, $0.065877) returned FREEZE-AFTER-FIXES. All six mandatory fixes incorporated: §3's Pearson claim narrowed to sign-only confounds with the undefined-not-zero correction (1); sign test conditional on non-ties, descriptive, per-cell primary (2); C1wordspre-specified as the single primary cue,A*demoted to descriptive with no p-value (3); Metric B′ item-level agreement added and Metric B demoted to a necessary condition (4); P2 relabelled[POST-DERIVED]and P4 relabelled a decision rule rather than a prediction (5); C4's regex corrected — v1 froze[.!?]+(?=\s\|$), a markdown-table escape that requires a literal pipe and does not count sentences (6). Recommended 2, 3, 4 incorporated; recommended 1 recorded as threat 1.