Repository path: workshop/experiments/E-20260801-strata/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260801-strata |
| status | frozen |
| created | 2026-08-01 |
| updated | 2026-08-01 |
| track | T2 |
| senses | naturalness, style-correspondence, cultural-mediation, accuracy |
| provisional | true |
| links | wiki/arms/ARM-panel-strata.md, wiki/findings/results/RS-20260730h-strict-coverage.md, workshop/experiments/E-20260730h-strict-coverage/design.md, workshop/regimes/R07-fluency.md, workshop/canon/max-havelaar-i/manifest.md, config/models.md, config/budget.md |
E-20260801 — the two codes nobody has checked, and what the cited-rule scaffold was doing
ARM-panel-strata step 1. Session S077. Frozen before the P and S rows of any published
R07 log were opened, and before the Dutch source was read for translation.
0. The wire, in one sentence
The study limb asks whether
R07'sPandScodes are reachable by a reader who did not make the choice — which forces an instrument that shows all ten rules rather than the cited ones, and so also re-measures theDstratum whose published agreement was obtained with the cited-rule scaffold in place; the translation limb is a fifthR07run whosePandSsites are held out from every published figure, so the answer does not rest only on material the same translator coded three months of sessions ago.
1. The question
R07 §5 defines three codes, assigned by the translator at the moment of each decision:
- D — a numbered rule names the feature at issue and only one live option satisfies it.
- P — one or more rules bear, and two or more live options satisfy all of them.
- S — no rule bears on the feature at issue at all.
RS-20260730h §3 put D to three independent seats and it survived: 4/4 positive control per
rater, 0.9216 byte-identical repeat, 0.8438 agreement with the lead. Its stated limit: "the panel saw
only sites the lead coded D. Nothing here bears on the project's P and S codes, and no claim is
made about them."
Published corpus: 110 sites — 39 D, 58 P, 13 S. The majority code and the silent code are
both unchecked, and three coverage rates are computed over all three.
Q1 — is P reachable? Do independent readers, given the live options and the whole rule set,
derive P where the lead coded P?
Q2 — is S reachable? Same, for S.
Q3 — what was the scaffold doing? E-20260730h showed each rater only the rules the site's log
cites. On D sites drawn from that run's own item set, does removing the scaffold — showing all ten
rules — change the derived code, and by how much? This bounds the published 0.8438.
Q4 — is S reachable in principle, or is it excluded by construction? S requires every one
of F1–F10 to be judged not to bear. F10 (nothing that calls attention to the language) and F6
(idiomatic syntax before close syntax) are stated broadly enough that a rater could hold them to bear
everywhere. If any rater judges a rule to bear at ≥ 90% of sites, then for that rater S is
unreachable whatever the lead coded, and Q2's answer is about the rule set rather than about the
rater. This is measured directly and reported whether or not Q2 succeeds.
Q5 (translation limb). On a fresh R07 run with the live options enumerated exhaustively and both
tests applied at the moment of decision, what is the strict-test D rate, and how large is the gap to
the exclusion-test rate? RS-20260730h §4 measured 12.7% and a 69.1-point gap on Russian; three
published strict rates are 11.4% / 5.6% / 6.7%.
2. Priming, declared before anything else
RS-20260730hhas been read in full during this session's orientation, including its §2 recount table, its §3 panel figures and its §2.1/§2.2 defect list. So the lead knows the strict-test outcome at the three replication sites («খোল-করতাল» P, «শ্মশান» P, «les gargoulettes» D) and knows the aggregate shape of theDrecount. TheDstratum is therefore declaredprimedat the aggregate level, which is one reason it is used here as an anchor rather than as a headline.- The
PandSrows of the fourR07logs have NOT been opened. What has been done to them is a regular-expression count of code letters in table cells, whose entire output was three numbers per file. No option set, no site description and no cited rule from anyPorSrow has been read. They are the blind strata and are opened only after this file is committed. - The Dutch source has not been read for translation. Its extent was fixed mechanically (first four
body paragraphs of chapter I, 716 words) by a script that printed word counts and 60-character
prefixes. No published English rendering of it has been opened and none will be before the
translation and its log are frozen (
R07§Procedure 1). - The
R07rule set isfrozenv1.0 and was frozen at S048, long before this design.
Consequence, registered: the D stratum may not carry any headline. Every Q1/Q2 figure is reported
over the blind strata alone, and the D stratum appears only in Q3, where its role is to be compared
against E-20260730h's own cells on the same sites.
3. Materials and sampling — the procedure is frozen and seeded
3.1 The four strata
| stratum | source | n | blind? |
|---|---|---|---|
| D-anchor | sites that E-20260730h scored (its 36 primary + 3 replication items), drawn by seeded sample |
10 | primed (aggregate) |
| P-published | all P rows across the three S064-era logs (Italian, French, Bengali) — 58 sites — drawn by seeded sample |
12 | blind |
| S-published | every S row across the same three logs — 13 sites, the whole stratum |
13 | blind |
| fresh | the translation limb's own log, seeded sample stratified 3/3/3 across the lead's D/P/S codes (fewer if a code is short; the shortfall is reported) |
≤ 9 | held out |
| controls | synthetic, built by the lead (§3.3) | 4 | — |
The Russian run (may-night-golova, S067) is excluded from P-published and S-published, because
its codes were assigned under both tests by a translator who already knew the distinction, which makes
it a different object from the three runs the published rates come from. It is named here so the
exclusion is a decision and not an omission.
Sampling procedure, executed by analysis/build_items.py with random.Random(20260801):
- Parse all four
R07logs into rows: site id, source expression, the log's own site description, the live-options cell, the cited-rules cell, the lead's code. - Filter to the stratum's population as above; sort by
(work, site_number)for determinism. sample(population, n).- Nothing is hand-picked and nothing is replaced. A drawn site that turns out to be one of the
four malformed rows
RS-20260730h§2.1 identified (two across-textF5rows, two multi-site rows) is kept and reported as drawn, and its derived code is reported asMALFORMEDrather than silently recoded. Excluding them after the draw would be fitting the sample to the answer.
3.2 The instrument — all ten rules, self-gated, per option
Each rater receives, per item:
- the source expression, in its own script, with the log's own one-line statement of what the decision is about;
- the live options verbatim from the log, in an order permuted per rater by a seeded shuffle, with the chosen rendering not marked;
- the whole rule set F1–F10, verbatim from
R07v1.0.
Each rater returns, per item:
bears— the subset of F1–F10 that bears on this decision at all;judgments— for each rule inbearsand each option,S(satisfies) orV(violates).
Raters are never shown the labels D, P, S, are never told what a coverage rate is, and are
never asked to classify a site. The code is derived arithmetically:
bearsempty → S- else survivors = options with no
Von any borne rule; D iff survivors == 1, P iff survivors ≥ 2, ANOMALY iff survivors == 0.
This is E-20260730h §3 steps 4–5 unchanged. The single difference from that instrument is which
rules the rater is shown, which is what Q3 measures.
3.3 The four synthetic controls
Built by the lead, in English only, with no source-language material, and placed in the item stream by the same seeded shuffle so they are not identifiable by position.
- X1 — an unambiguous
D. Two options for one feature, one of them plainly dated: F1 bears and exactly one option survives. Correct answer:bearscontains F1; derived codeD. - X2 — an unambiguous
P. Two options, both current, both standard, both idiomatic; F1 bears (the feature is currency of usage) and both survive. Correct answer:bearsnon-empty; derived codeP. - X3 — an unambiguous
S. Two renderings of a neutral narrative verb, identical in currency, register, dialect, syntax, continuity, determinacy, cadence and salience. Correct answer:bearsempty; derived codeS. - X4 — an unambiguous
Don a different rule. F4 bears (one option retains an untranslated source-language word) and exactly one option survives. Correct answer:bearscontains F4; derived codeD.
Scoring is on the derived code only, not on the exact bears set, because a rater that adds a
second borne rule which excludes nothing still derives the right code, and excluding it would be the
S064/S067 defect a third time.
X3 is the control this design is most likely to have mis-built, and it is named as such for the
pre-run critic. If X3 fails across all three raters, Q2 is void and the finding is Q4's — that
S is not expressible in this instrument — not a finding about the lead's S codes.
4. Conditions
C1 — the panel (three seats, one pass each)
Seats P1 openai/gpt-5.6-terra, P3 x-ai/grok-4.5, P2 google/gemini-3.6-flash — the same
three seats RS-20260730h §3 used, deliberately, so Q3's cross-instrument comparison is within-seat
and not confounded by seat identity. Declared reserve for a rater seat: P5
deepseek/deepseek-v4-pro. The critic seat and its reserve (§4.4) are excluded from the rater
reserve table by construction — note (bgj), the S053 role collision.
Items are dispatched in three batches per seat, split by the seeded shuffle, to keep each response
inside its cap. max_tokens 8,000 per batch.
C2 — the byte-identical repeat
One seat (P1) receives batch 1 a second time, byte-identically, in a separate dispatch. Prompt
bytes and prompt-token counts are asserted equal by the verifier. Reported at cell level and at derived-
code level, as RS-20260730h §3 reported them.
C3 — the prose-only leak null (registered, not a diagnostic)
One seat (P2) receives, for every item, only the log's one-line site description — no
options, no rules, no source expression — plus the three code definitions verbatim, and is asked to
name the code. Note (bfx): E-20260731d cleared its leak null by 0.035, and this design's items
carry the same lead-written prose.
Registered reading, written before the run: if C3's agreement with the lead reaches within 0.05 of C1's majority agreement on the same items, no reachability claim is made for the affected stratum, and the session reports the null as its result.
C4 — the permutation null (free, local)
For each stratum, the agreement between the panel majority and the lead is compared against the
distribution obtained by permuting the lead's code labels across the sampled sites, 20,000 draws,
random.Random(20260801). Reported as a p-value beside every agreement figure. This is not a
substitute for C3 — it prices sampling, not leakage.
C5 — the trivial-majority baseline (free, local, registered now)
RS-20260730h §3's registered 0.75 threshold "turned out, by accident, to be exactly the no-effort
baseline", and the same trap is live here. Computed and printed for every stratum before any
agreement figure is read: the score a rater would get by always returning the stratum's most common
lead code. On the pooled blind sample (12 P + 13 S + ≤ 9 fresh) the majority-class rate is
approximately 0.5 and is reported exactly.
C6 — the translation limb (free, no call)
A fifth R07 run: Multatuli, «Max Havelaar» (1860), chapter I, first four body paragraphs, 716
Dutch words — the project's first Dutch and its sixteenth source language. Both the original
and a published English translation (Nahuÿs 1868) are freely reachable, which is the materials
preference CLAUDE.md states.
Executed under R07 v1.0 with two additions, both from RS-20260730h §4's own procedure:
- Live options enumerated exhaustively before any code is assigned at each site.
- Both tests applied at the moment of each decision —
excl(a cited rule excludes ≥ 1 live option) andstrict(§5's exactly one survivor) — recorded as separate columns.
Contamination. tools/dependence_check.py, unmodified, against the Nahuÿs text, which is extracted
mechanically to a scratch file and not opened by the translator. Per CLAUDE.md's standing rule the
gate runs on paragraph 1 before paragraphs 2–4 are drafted, and per note (bfd) it is re-run on
the whole extent at the end, because a unit-1 gate does not bound a work.
Freeze order, which is the limb's whole evidential value: this design → the translation and its log
with codes → analysis/build_items.py. The lead's fresh codes are committed before the item builder
exists, so they cannot be fitted to anything.
4.4 The pre-run critic
qwen/qwen3.7-max, max_tokens 12,000, declared reserve moonshotai/kimi-k3 at the same cap
(note (b), twenty-second firing at S076; note (bfc)'s one-line-per-stage rule). Neither is a
rater seat and neither appears in the rater reserve table.
The critic is shown this design and analysis/build_items.py, analysis/prompts.py and
analysis/run.py — note (bgj): S076's critic was shown design.md and not the runner, and the
defect was in the runner. It is given the commit hashes of the freeze chain — note (bff).
It is pointed at §3.3 first, in the prompt, because a mis-built control is this design's most likely defect and has been the defect twice running in this line of work.
5. Predictions, registered
Registered before any log row of the blind strata was opened and before the source was read.
- P1 —
Sis the least reachable of the three codes. Panel-majority agreement with the lead is lower on the S-published stratum than on the D-anchor stratum. - P2 —
Pis reachable above the trivial baseline. Panel-majority agreement on the P-published stratum exceeds C5's majority-class rate for that stratum by ≥ 0.10. - P3 — the repeat holds. C2 cell-level self-agreement ≥ 0.85 (the precedent is 0.9216 on a smaller cell count and an easier task).
- P4 — the scaffold was doing work. On the D-anchor stratum, the all-ten-rules instrument agrees
with the lead less than
E-20260730h's cited-rule instrument did on the same sites — i.e. below its within-seat rate there. Direction only; no size is predicted. - P5 — at least one rater treats at least one rule as near-universal. Some (rater, rule) pair has
a
bearsrate ≥ 0.90 across all scored items. Named candidates, in advance: F10, then F6. - P6 — the fresh run replicates the gap. The translation limb's
exclrate exceeds itsstrictrate by ≥ 15 percentage points (RS-20260730h§4's P5 held at 69.1). - P7 — the fresh run's strict rate lands in the published band. Its strict
Drate is between 4% and 15% (published: 11.4%, 5.6%, 6.7%; Russian: 12.7%).
A null is a result. If P and S both come back unreachable, that is the finding, and it is a
finding about a scheme this project publishes rates over.
6. Failure criteria — what voids what
- F1. If X1, X2 and X4 (the three non-
Scontrols) are not all derived correctly by at least 2 of 3 raters, the whole run is void and no figure is reported. - F2. If X3 fails for all three raters, Q1 and Q2 are void as reachability claims and the
session reports Q4 instead:
Sis not expressible in this instrument. Q3 survives, because it does not involveS. - F3. If C2 cell-level self-agreement falls below 0.80, no agreement figure is reported as a measurement; all are reported as a single unreplicated observation with the instability stated.
- F4. If C3 comes within 0.05 of C1 on a stratum, that stratum reports no reachability claim (§4 C3).
- F5. A seat returning fewer than all items in a batch is re-dispatched once; if it fails again
the declared reserve takes the seat and the substitution is reported. A batch answered by two
different models is not a seat and is reported
VOIDfor that batch — note (bgl). - F6. If
ANOMALY(zero survivors) exceeds 20% of a rater's items, that rater's derivations are reported as anomalous rather than pooled, because a zero means the rater's own judgments are internally inconsistent.
7. Verification
analysis/verify.py, importing nothing from analyse.py:
- re-parses every stored
.rawbody and re-sumsusage.costindependently; - asserts
finish_reason == "stop"and a stored content length for every accepted body — the repair named by thewiki/backlog.mdrow "A verifier must be shown to FAIL when the body it reads is wrong" (S075, age 1), which found 4 of 18 verifiers blind to their own data; - re-derives every code from the raters' raw
bears/judgmentsby an independent implementation; - re-parses the four
R07logs from markdown and re-counts the published strata; - re-computes every agreement, baseline, permutation p-value and rate reported;
- runs mutation tests, including at least one that truncates an accepted body to 40% and asserts the verifier fails.
8. Budget
Declared worst case $1.10, built from max_tokens per note (abc) and not from assumed output
length:
| dispatch | seats × calls | cap | worst case |
|---|---|---|---|
| pre-run critic | 1 (+1 reserve) | 12,000 | $0.36 |
| C1 rater batches | 3 seats × 3 batches | 8,000 | $0.54 |
| C2 repeat | 1 | 8,000 | $0.06 |
| C3 prose-only null | 1 | 4,000 | $0.04 |
| headroom for one re-dispatch (F5) | 1 | 8,000 | $0.10 |
UTC day 2026-08-01 opens at $0.00 of $5.00. Note (bgk): x-ai/grok-4.5 did not enforce
max_tokens at S076, so the worst case is not in fact an upper bound for P3 and this is recorded
before the run rather than after it.
The translation limb costs $0.00 and is never ledgered (charter §3, A4).