Repository path: workshop/experiments/E-20260801b-census-author/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260801b-census-author |
| status | frozen |
| created | 2026-08-01 |
| updated | 2026-08-01 |
| senses | accuracy, naturalness, style-correspondence, voice, cultural-mediation, consistency, purpose-fit |
| provisional | true |
| links | wiki/arms/ARM-option-census.md, workshop/experiments/E-20260801b-census-author/mode-rule.md, workshop/experiments/E-20260801b-census-author/register.md, workshop/experiments/E-20260731e-option-census/design.md, wiki/findings/results/RS-20260731e-option-census.md, wiki/findings/results/RS-20260730i-candidate-reach.md, framework/closure.md, framework/traceability-inventory.md, workshop/translations/kowalski/R04-v1/passageA.md, workshop/translations/kowalski/R04-v1/passageB.md, config/models.md, config/budget.md |
E-20260801b — does k survive a change of census author, and is a change of author the same thing as a change of elicitation mode?
ARM-option-census step 2. Frozen before any rater was dispatched. The arm page's step-2 text is
the design's starting point and is quoted rather than paraphrased:
"
RS-20260730iestablishedk(renderings ruled out) as the only option-count-invariant statistic and this project now recommends quoting it. Step 1 showed the option contents do not reproduce between subjects (three-seat Jaccard 0.170) and never coded a single candidate against a non-lead census. Sokhas been shown invariant to how long the list is and never to whose list it is."
1. The question, and the one the arm page did not ask
Primary. framework/closure.md §1.4 tells the project to report k — the absolute number of
live renderings a bearing candidate rules out — at 0.14, on the ground that it did not move when
the option list was halved (RS-20260730i: 0.1395 at four options, 0.1538 at two). Every census that
figure was ever computed over was written by the lead.
Does
kmove when the census author changes?
Secondary, and it is why this session has a translation limb. The four author arms are not
author-pure. The lead's census is a moment-of-decision log; the three seats' censuses are
on-demand reconstructions by subjects who did not translate. RS-20260731e §7 named that
difference and did not measure it. So a LEAD-vs-seats gap in k is author confounded with
elicitation mode, and the size of the mode effect is unknown.
Two things follow, and the design uses both:
- The among-seats spread (P1 vs P2 vs P3) is mode-held-constant and is a clean author contrast on its own.
- The mode arms — the same translator, the same regime, the same work, two adjacent passages, one census written at the moment of choosing and one reconstructed afterwards — estimate the mode effect directly, on prose translated this session for the purpose.
2. Materials
The author arms. The same 24 Reymont «Śmierć» loci (S01–S24) frozen at S073, under four
censuses:
| arm | census | mode | provenance |
|---|---|---|---|
LEAD |
the lead's own | MOMENT |
E-20260731e/census.json, frozen S073 |
P1 |
openai/gpt-5.6-terra |
ONDEMAND |
E-20260731e/runs/r1_NONE_P1.txt |
P2 |
google/gemini-3.6-flash |
ONDEMAND |
E-20260731e/runs/r1_NONE_P2.txt |
P3 |
x-ai/grok-4.5 |
ONDEMAND |
E-20260731e/runs/r1_NONE_P3.txt |
The seats' UNPRIMED (NONE) round is used and the FRAMEWORK round is not. A census written
after reading the fourteen candidates cannot then be coded against them without circularity. Stated
here so the choice is not read as a selection among available arms.
The mode arms. Adam Szymański, «Stolarz Kowalski» (1911), Polish Wikisource, opening paragraph,
translated this session under R04 and split at a sentence boundary:
| arm | passage | Polish → English | census | sites |
|---|---|---|---|---|
ONDEMAND |
A, opening → "…grają na słońcu cudownymi blaski." | 393 → 520 | reconstructed after the translation froze | O01–O20 |
MOMENT |
B, "Ale dłużej trochę pobyć tu potrzeba…" → end of ¶1 | 366 → 525 | written at the moment of decision | M01–M20 |
Contamination, measured before any locus was selected and before B existed
(tools/dependence_check.py, unmodified, against Benecke & Busch 1921, Project Gutenberg 8378, a
comparator located and written to file programmatically and never read):
| unit | longest shared run | 7-grams | 12-grams | 15-grams | verdict |
|---|---|---|---|---|---|
| passage A | 6 | 0 | 0 | 0 | clean |
| passage B | 9 | 6 | 0 | 0 | clean |
| whole span | 9 | 6 | 0 | 0 | clean |
The A/B pair is itself a free datum for a question NEXT.md has been carrying: two adjacent
~400-word passages of one work, one translator, one comparator, one session — 6 tokens and 9.
That is a within-condition spread of the longest-run statistic measured for the first time.
The register declaration is register.md, frozen in commit 1bc4b41 before the first English
word of either passage. It is a genuine advance declaration for the Szymański passages and a
declaration supplied by this design for the Reymont sites, which had none. §7 carries the cost.
3. The coding instrument, and it is not new
The rater prompt is E-20260730i's condition-C prompt, reused with three edits, all recorded here:
- the standing register fact reads British rather than American (
register.md); - each item tags its source language, including the six controls, so that "no language tag" is not a marker identifying the control block;
- each site prints its source sentence as
CONTEXT, truncated at 240 characters.
Nothing else changes: the fourteen candidate rows are read verbatim out of the frozen S068 prompt
file by build_items.py, the task wording is unchanged, and the output format is unchanged.
Derived quantities, unchanged from E-20260730i §5, with n = options and k = options
excluded, over bearing (site, candidate) pairs by three-rater majority:
| code | condition |
|---|---|
INVOKED |
0 < k < n |
DECIDES |
k = n − 1 |
VACUOUS |
k = n |
k |
the primary statistic: mean options excluded per bearing (site, candidate) pair |
E = k/n is not computed. RS-20260730i §1.1 measured it to be option-count-dependent and
framework/closure.md forbids it.
3.1 The control block — data-qualified, not endorsed
Note (bfy): "a control item is qualified by DATA, not by a second opinion — either by a prior run in which it behaved as declared, or by building more control items than the criterion needs and reporting which ones behaved." Two consecutive sessions had a critic's blanket sight-unseen endorsement as their weak point.
So this design builds no new control items. All six are reproduced verbatim from
E-20260730i's condition-C payload, where P1 passed 6 of 6, P3 passed 6 of 6 and P5 passed 3 of
6 — recorded in that run's analysis/results.json and re-read by build_items.py from the
frozen prompt file rather than retyped.
| id | 730i item | kind | declared expectation | prior behaviour |
|---|---|---|---|---|
X1 |
I06 | CTRL-POS |
C1 bears, excludes exactly 3 of 4 | P1 ✓ P3 ✓ P5 ✗ |
X2 |
I19 | CTRL-POS |
C1 bears, excludes exactly 3 of 4 | P1 ✓ P3 ✓ P5 ✗ |
X3 |
I30 | CTRL-VAC |
C1 bears, excludes all 4 | P1 ✓ P3 ✓ P5 ✓ |
X4 |
I62 | CTRL-NEG |
no candidate bears | P1 ✓ P3 ✓ P5 ✗ |
X5 |
I43 | CTRL-BEAR |
C3 bears | P1 ✓ P3 ✓ P5 ✓ |
X6 |
I54 | CTRL-BEAR |
C1 bears | P1 ✓ P3 ✓ P5 ✓ |
Only X6 reads the standing register fact; the other five carry their own note or are
register-independent. Under British dialect X6's four options still span four registers, so C1
bears on it exactly as before.
Every control appears in all three payloads, which buys a within-rater consistency check on the control block itself at no extra dispatch.
Note (bgo) is applied literally: the pre-run critic is pointed at the control block first, in its prompt, because "quoting a defect back at yourself does not prevent it — pointing the critic at the control block first does."
3.2 Payloads — a rotation, one arm per site per payload
⚠ AMENDED before any rater call, on the pre-run critic's BLOCKING finding 2 (
critic.md, amendment A2). The structure below replaces a three-payload version in which every Reymont site appeared four times inside one payload, once per arm. The critic's objection is accepted in full: ratings 2–4 would have anchored on rating 1, and a rater could have identified theLEADarm from list length alone. The superseded version is in commitd57ca16.
| payload | contents | items |
|---|---|---|
R1 |
all 24 Reymont sites, arm = AUTHOR[(i + 0) mod 4] + 6 controls |
30 |
R2 |
all 24 Reymont sites, arm = AUTHOR[(i + 1) mod 4] + 6 controls |
30 |
R3 |
all 24 Reymont sites, arm = AUTHOR[(i + 2) mod 4] + 6 controls |
30 |
R4 |
all 24 Reymont sites, arm = AUTHOR[(i + 3) mod 4] + 6 controls |
30 |
B |
Szymański O01–O20 and M01–M20 + 6 controls |
46 |
AUTHOR = (LEAD, P1, P2, P3) and i is the site's index in source order. The rotation gives:
- each Reymont site appears exactly ONCE per payload, so within-payload anchoring on a repeated site is structurally impossible and list length is never comparable across arms at one site;
- each (site, arm) cell appears exactly once across the four payloads — 96 cells, no repeats;
- each payload carries six sites from each of the four arms, so a payload effect (a different provider, a tired seat, a different completion length) still cannot be confounded with an arm.
All four properties are asserted by build_items.py at build time and re-asserted independently by
analysis/verify.py. Item order inside each payload is shuffled with the fixed seed 20260801
declared here and nowhere else; the keymap is written to disk before dispatch.
What the rotation does not fix. In payload B the twenty ONDEMAND items all come from passage
A and the twenty MOMENT items all from passage B, so a rater attending to subject matter could
separate them. Declared, not repaired.
3.3 Raters and the repeat
Three seats, P1 / P2 / P3 of config/models.md, one call per payload. Reserve declared
before dispatch, per note (bfc): deepseek/deepseek-v4-pro.
The drift gate is taken by choice (workshop/experiments/README.md, option 1, and note
(bfz)): payload A1 is re-issued byte-identically to one seat in the same session, and the
δ it measures is this run's noise floor. Completion length per seat is reported beside it because it
is free and (bfz) asks for it.
Any arm difference smaller than the measured floor is reported as inside the floor and not as an estimate. S073 published a null instead of a two-of-three-seat effect on exactly this ground.
4. Registered predictions
Registered before any rater call. The primary threshold is RS-20260730i's own: that run scored
its length manipulation on Δ mean k ≤ 0.25 and reported −0.0605. Scoring the author
manipulation on the same bar is the only way the two are comparable.
⚠ AMENDED before any rater call, on the pre-run critic's BLOCKING finding 1 (amendment A1). The four-arm range confounds author with elicitation mode, because
LEADis a moment-of-decision census and the three seats are on-demand reconstructions. It is renamed author-or-mode and may never be reported as "the author effect". The clean author contrast — three different authors, elicitation mode held constant — is the three-seat range, and it is promoted to its own prediction P1b.
| # | prediction | what falsifies it |
|---|---|---|
| P1 | author-or-mode. Range of k across the four author arms at the same 24 loci is ≤ 0.25 |
a range above 0.25 — k is not invariant to who wrote the census and how, and framework/closure.md's recommendation to quote it must be qualified. A large range is a FALSIFICATION under every reading; there is no wording of P1 on which it is a pass |
| P1b | the clean author contrast. The three-seat range of k — three different authors, all ONDEMAND — is ≤ 0.25 |
a range above 0.25 — author alone moves k, with mode held constant, and this is the finding that needs no confound argument at all |
| P2 | the decomposition. The three-seat range is smaller than the four-arm range | equal or larger. UNRESOLVED, not falsified, if the four-arm range is at or inside the repeat's measured δ (amendment A5): comparing two quantities that are both inside the noise floor is not a result in either direction |
| P3 | mode. |k(MOMENT) − k(ONDEMAND)| on the Szymański loci is ≤ 0.25, the same bar. It may not be subtracted from P1 (amendment A3): the mode arms are on different text, are confounded with passage, and bound an order of magnitude, nothing more |
above 0.25 — elicitation mode moves k by more than option-list length did |
| P4 | (descriptive) mean n is highest in LEAD |
any arm above it |
| P5 | frac2 (sites with ≤ 2 live renderings) is 0.000 in every arm, as it was in nine of nine cells at S073 and in the lead's own census |
any arm above 0 — and DECIDES becomes arithmetically reachable there |
| P6 | closure-defeating, registered as such. If a candidate outside {C1, C2, C3, C5} is INVOKED at ≥ 3 sites by majority in any arm, framework/closure.md §1.4's reach sentence must be restated, not annotated, and this arm may not close without doing so |
— (a trigger, not a scored prediction) |
5. Failure criteria
| # | fires when | consequence |
|---|---|---|
| F1 | fewer than 2 of 3 raters pass all six controls | the run is descriptive only; no k comparison is reported as an estimate |
| F2 | the byte-identical repeat's δ on k exceeds the four-arm range of P1 |
P1 is unresolved; the instrument cannot see an effect of the size present |
| F3 | any arm returns fewer than 8 bearing (site, candidate) pairs by majority | k for that arm is not reported as an estimate and the arm is described only |
| F4 | a rater returns fewer lines than its payload has items, after one retry and one fall-through to the declared reserve | that seat's payload is excluded whole and every figure is recomputed on the remaining seats, declared post hoc |
| F5 | the control block's within-rater consistency across the three payloads is below 5 of 6 for a rater that F1 passed | the control pass is reported as unstable and F1's verdict is repeated with the instability named |
F1 is written to cover the case the null wins. NEXT.md records that S077's registered F4 "as
registered does not cover the null winning" and that the stronger reading had to be applied after
the fact (note (bgn)). Here: if the range in P1 exceeds 0.25 in the direction that qualifies
framework/closure.md, that is a FALSIFICATION and is reported as one — there is no reading of P1
under which a large range is a pass.
6. Procedure
- Instrument
mode-rule.mdfrozen (a448ed1) before the span was chosen. ✔ - Register declaration frozen (
1bc4b41) before the first English word. ✔ - Passage A
R06draft (af9c8dd) →R04(b29062e) → contamination gate →ONDEMANDcensus (41a01ef) → passage BR06(81b1af0) →R04+MOMENTcensus (ef64a59). ✔ - Items built by
build_items.pyfrom the frozen censuses and the frozen S068 prompt file; keymap written to disk. ✔ - This design frozen.
- Independent pre-run critic, pointed at the control block first (note (bgo)); findings
recorded in
critic.mdwith accept/decline for each; amendments committed before any rater call. - Three raters × three payloads, raw bodies preserved before parsing (note (bdt)); one byte-identical repeat.
analyse.py, thenanalysis/verify.py, which imports nothing fromanalyse.py, re-parses every answer from the stored.rawbytes, and carries mutation tests.
7. What this cannot establish, written before the run
- It cannot separate author from mode for the LEAD arm. That is the whole reason the mode arms exist, and the mode arms are on a different passage, so the decomposition is by order-of-magnitude comparison, not by subtraction.
- The mode arms are confounded with passage. A and B are adjacent but different text. Per-arm
nand class distributions are reported so a reader can see how far apart they are. ONDEMANDis biased towardMOMENT. The translator knew a census would be wanted, so the reconstruction is by someone who had recently been attending to choices — which shrinks any mode effect. Conservative for P3, not for a null reading of it.- The register given to raters for the Reymont sites was supplied by this design, not declared
by the S073 translator (
register.md). C1's behaviour on those 24 loci is therefore conditional on a declaration made after the fact. - Three models enumerating on demand are not three translators. Unchanged from
RS-20260731e§7 and not repaired here. - It judges nothing. No rater is asked whether any rendering is good; the lead never judges its
own translation (charter §5). The panel is NOT CALIBRATED and every figure is
provisional. - One work per limb, one language pair, 24 + 40 sites, three raters.
8. Pre-flight cost estimate
Worst case is built from the max_tokens cap the request permits, not from an expected length —
note (abc).
| stage | calls | max_tokens |
worst case |
|---|---|---|---|
pre-run critic (qwen/qwen3.7-max) — SPENT, $0.071769075 |
1 | 12,000 | $0.20 |
payloads R1–R4 × 3 seats |
12 | 6,000 | $1.10 |
payload B × 3 seats |
3 | 6,000 | $0.28 |
byte-identical repeat of R1, one seat |
1 | 6,000 | $0.10 |
| declared worst case | 17 | $1.68 |
P3 is priced at the worst plausible provider per the S022 routing caution. Today's ledger (UTC
2026-08-01) opened at $0.580589442 of $5.00 with $4.419410558 headroom, so the worst case is
38% of headroom. The worst case was raised from $1.14 to $1.68 by amendment A2, which
doubled the author-arm dispatches from six to twelve — the cost of the critic's BLOCKING finding,
recorded as such rather than absorbed silently. A stage that will not fit is dropped whole, in the
order B → repeat → R4, and the deferral is written into NEXT.md.