Repository path: workshop/experiments/E-20260731e-option-census/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260731e-option-census |
| status | frozen |
| created | 2026-07-31 |
| updated | 2026-07-31 |
| senses | cultural-mediation, style-correspondence, voice, naturalness, consistency |
| links | wiki/arms/ARM-option-census.md, framework/closure.md, framework/traceability-inventory.md, wiki/findings/results/RS-20260730i-candidate-reach.md, wiki/findings/results/RS-20260730h-strict-coverage.md, workshop/experiments/E-20260731e-option-census/census-rule.md, workshop/experiments/E-20260731e-option-census/census.json, workshop/translations/smierc/R04-v1/translation.md, workshop/experiments/README.md, config/models.md |
E-20260731e — whose option list is it?
Re-frozen 2026-07-31 after the independent pre-run critic pass (critic.md, qwen/qwen3.7-max,
NEEDS-AMENDMENT, three findings — one BLOCKING, one MANDATORY, one ADVISORY, all three
accepted; all eight control items ENDORSED sight-unseen). Nothing was dispatched to a rater
before this amendment.
Frozen 2026-07-31 (S073) before any call was dispatched. The census instrument was frozen
earlier still, at 8c12508, before the source span existed (census-rule.md).
1. Question
Every framework-reach figure this project publishes is computed over an option census written by
the lead. framework/closure.md §1.4 reports DECIDES 0 of 126 and k = 0.14; §1.4's boxed
correction (S068) reports 0.000 at four options against 0.154 at two;
framework/traceability-inventory.md §3 reports which of the fourteen candidates a decision reaches
and how the invocations concentrate. All of it is a function of a list of live renderings, and
every such list in the repository was written by one translator who has read the fourteen
candidates. ARM-candidate-reach said so in its own closure: "It never coded a non-lead option
census, so §2's reach figures are still one translator."
RS-20260730h (S067) then showed that the project's other coverage statistic was rank-identical to
the mean number of live options the translator wrote down, and RS-20260730i (S068) showed that
DECIDES doubles-to-appearing when the option list is halved while k does not move.
So the denominator is doing work, and nobody has asked where the denominator comes from.
Does exposure to the framework change the option list?
If it does, the object every reach statistic is computed over is endogenous to the thing the statistic is measuring.
2. Why this is T5
The object under audit is the synthesis object's own denominator — not a jury, not a
calibration, not a typology. It is the S068 shape: auditing what the synthesis object already
claims, which wiki/tracks.md records as the one kind of T5 work that is neither downstream of
Tier D nor a fourth container. It is not a release push and ARM-option-census's completion
criterion forbids it becoming one. No candidate is added, removed or licensed by this design.
3. Materials
- Source. Reymont, «Śmierć» (1893), ¶0–51, 904 Polish words. Project's first Polish.
- Translation limb.
T-smierc-R04-v1, withT-smierc-R06-v1as its frozen draft. Contamination clean, measured: whole-artifact longest shared run 10 tokens, 0 shared 12-grams, against Benecke & Busch 1921. The gate was run prospectively on ¶0–9 before the span was fixed; it had already fired at 18 tokens on a different work, which was abandoned (workshop/translations/janko/R06-v1/unitA.md). - Census.
census.json: 40 sites, of which the first 24 in source order are the primary items by the rule frozen at8c12508, plus 8 constructed control items (C1–C4CONSTRAINED, C5–C8OPEN). - Item payload. 32 items. Each gives the source paragraph with the locus marked by
«...», and the locus repeated on its own line. The lead's own options are never shown. - Item order. One frozen shuffle, identical across every arm and every seat, produced by
random.Random(20260731).shuffle(ids)on the sorted id list. Stated here so it is reproducible and so the order cannot be a per-arm variable.
4. Arms
Three preambles. The preamble is the only thing that differs between arms; task text, items and order are byte-identical.
| arm | preamble |
|---|---|
| NONE | none |
| SHAM | 14 lines of generic translation advice with no warrant of any kind, matched to FRAMEWORK on line count and word count (prompts.py, and the realised counts are reported) |
| FRAMEWORK | the fourteen candidate recommendations of framework/traceability-inventory.md §2, verbatim, row text only, as RS-20260728i and RS-20260730i supplied them |
The sham is built and measured, not asserted (standing disposition 6). It is matched on shape (14 numbered imperative-or-declarative lines) and on length; its realised word count against FRAMEWORK's is reported in the analysis, and a mismatch above 15% is declared as a confound.
Nothing in either preamble mentions counting, enumerating, listing more, or option lists. The draft sham line "the first rendering that comes to mind is rarely the last one to consider" was removed before freezing because it encourages enumeration directly and would have confounded the sham with the hypothesis; it is replaced by "where the source is ambiguous, the rendering may be ambiguous too." Recorded because a removed item is part of the design.
5. Task, identical in all arms
For each item below you are shown a paragraph of Polish with one locus marked
«…». List every distinct English rendering of the marked locus that you would consider live — one you would actually weigh before choosing. Do not choose one. Do not rank them. Give each as a short phrase.Answer with one JSON object per line:
{"id": "...", "options": ["...", "..."]}and nothing else.
6. Seats
Raters P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5
(config/models.md). Reserve declared before dispatch (note (bfc)): deepseek/deepseek-v4-pro.
Pre-run critic: qwen/qwen3.7-max, probed-but-not-selected, not a rater and not the raters'
reserve — the S053 role-collision fix, seventeenth consecutive session. max_tokens 6,000 on
raters (note (b), which has fired on empty length returns eighteen times), 8,000 on the critic.
Judgment is not parallelised and none is asked for. Enumerating renderings one would weigh is not a quality judgment about any of them; no seat is asked which is better, and no seat sees the lead's translation or its census.
7. Dispatch plan
| round | arm | seats | purpose |
|---|---|---|---|
| 1 | NONE | P1, P2, P3 | baseline |
| 2 | SHAM | P1, P2, P3 | preamble-presence control |
| 3 | FRAMEWORK | P1, P2, P3 | the manipulation |
| 4 | NONE, byte-identical repeat | P1, P2, P3 | the same-day noise floor |
12 rater calls + 1 critic call. Round 4 is workshop/experiments/README.md's cross-day drift
gate, option 1, taken by choice — the first design to take it rather than the exemption.
Completion length per seat is reported for every call, as that gate also requires.
8. Statistics
n(seat, arm, item)= number of options returned.N(seat, arm)= meannover the 24 primary items.N_open(seat, arm),N_con(seat, arm)= meannover the 4OPENand 4CONSTRAINEDcontrols.frac2(seat, arm)= fraction of the 24 primary items withn ≤ 2.containment(seat, arm)= mean over primary items of |lead options matched by some seat option| / |lead options|, under the matcher frozen in §9. A lower bound, and reported as one.
9. The matcher, frozen verbatim (standing disposition 9)
import re
STOP = r"\b(the|a|an|to|of|my|his|her|your|it|is|are|i|me|we|they|he|she|was|were|be|been|that|and)\b"
def norm(s):
s = s.lower().strip()
s = s.replace("’", "'")
s = re.sub(r"[^a-z0-9' ]+", " ", s)
s = re.sub(STOP, " ", s)
s = re.sub(r"\s+", " ", s).strip()
return s
Two renderings match iff norm() is equal. Nothing fuzzier; no stemming, no synonymy.
Amended after the critic pass (ADVISORY finding 3): the stop list gained the personal pronouns and the copula/auxiliary forms, which it had omitted. The change loosens matching in both arms equally, so it cannot move the Pred5 contrast; it raises the absolute containment figures, which were always declared a lower bound.
10. Failure criteria — registered, and two of them can void the run
F1 — the positive control, VOIDING. Pooled over the three arms, for each seat,
N_open − N_con ≥ 1.00. Required in at least 2 of 3 seats. If fewer than 2 seats separate the
control classes, the seats are not tracking option availability at all and every primary number in
§8 is VOID, reported as such and not reinterpreted.
F2 — the noise floor. δ(seat) = |N(seat, NONE, round 1) − N(seat, NONE, round 4)|. Any arm
difference whose magnitude is below max_seat δ is reported as inside the noise floor and is not
an estimate. This is RS-20260730-grain-clause's F1 imported as a standing condition rather than
rediscovered.
F3 — coverage. A seat returning fewer than all 32 item ids in an arm is excluded from that arm, and the exclusion is reported. If fewer than 2 seats survive in any arm, the primary is VOID.
F4 — sham length. If the realised word counts of the SHAM and FRAMEWORK preambles differ by more than 15%, the sham is declared a confounded control and P2's result is reported with that caveat attached. It does not void anything.
11. Predictions, each with its failure condition
| # | prediction | fails if |
|---|---|---|
| Pred1 | N(FRAMEWORK) > N(NONE) in ≥2 of 3 seats, by more than the F2 floor |
the direction reverses in ≥2 seats, or the difference lies inside the floor |
| Pred2 | \|N(SHAM) − N(NONE)\| < \|N(FRAMEWORK) − N(NONE)\| in ≥2 of 3 seats |
otherwise — which would say the effect is the presence of a preamble, not its content |
| Pred3 | the lead's own mean n over the 24 primary items, 4.750, frozen at 6b9b447, exceeds N(NONE) averaged over seats |
it is lower. Descriptive: one subject, a different elicitation mode |
| Pred4 | frac2 is 0.000 for the lead and > 0 for at least one seat in the NONE arm |
every arm also returns 0.000 |
| Pred5 | containment(FRAMEWORK) > containment(NONE) in ≥2 of 3 seats |
otherwise. Descriptive, and a lower bound by construction |
Pred4 is the prediction that touches a published figure. RS-20260730i measured DECIDES at
0.000 with four live options and 0.154 with two. frac2 is therefore the lever on
framework/closure.md §1.4's headline, and the lead's census supplies 0 of 24 sites at n ≤ 2 —
which is the third independent time a lead census has produced no n = 2 sites (S068's pre-run
critic found the second and it changed that design).
12. What this design does not claim
It does not claim any rendering is better than any other, and no seat is asked. It does not claim the lead's census is wrong. It does not test whether shared runs are forced. And it cannot establish that a non-lead census is the right denominator — three language models enumerating renderings on demand are not a translator writing a log at the moment of choosing, and the two elicitation modes differ in more ways than the manipulation. What it can establish is whether the denominator moves under a manipulation of exactly one variable, which is the only question §1 asks.
13. Pre-flight cost
Worst case $1.20, built from max_tokens (note (abc)) at list out-price × 1.5 provider premium
(S022 routing caution), summed over 12 rater calls at 6,000 and 1 critic call at 8,000, with input
priced at the realised payload size. Today's headroom before this run: $4.293746048.