Repository path: workshop/experiments/E-20260820b-device-function-2/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260820b-device-function-2 |
| status | frozen |
| created | 2026-08-20 |
| updated | 2026-08-20 |
| senses | style-correspondence, perceived-source-carriage, accuracy |
| provisional | true |
| links | wiki/arms/ARM-device-function.md, workshop/experiments/E-20260816g-device-function/design.md, wiki/findings/results/RS-20260816g-device-function.md, workshop/translations/zhongli/R04-v1/translation.md, framework/v0.2/README.md, wiki/method-notes.md, config/models.md, config/budget.md |
E-20260820b — the same twelve stretches, replaced by a hand that did not write the labels
Design v4, frozen and DISPATCHED. ARM-device-function step 2 (T5), and the whole of it: the arm
closes on this run either way.
v1, v2 and v3 were each reviewed by an independent adversarial critic and NONE OF THEM WAS
DISPATCHED. v1: NEEDS REDESIGN, 9 findings, 3 BLOCKING ($0.063292). v2: NEEDS REDESIGN, 7
findings, 3 BLOCKING ($0.067593) — none a restatement; the pass reviewed the machinery v2 had added
and broke it: v2's noise band was anti-conservative in exactly the direction the arm's payload
needed. v3: NEEDS REDESIGN, 7 findings, 2 BLOCKING ($0.085206) — both new, both fixable free, both
fixed here. $0.216091 on three critiques, and each one removed a claim the design was not entitled
to. §9 is the full record. Every change across the three revisions is a narrowing: weaker
claims, stricter gates, Q2 rebuilt twice, and Q1 — what step 2 was constituted to buy — reduced
to arithmetic that licenses nothing (§4).
Pre-commitment against regress, written BEFORE pass 3 was bought and applied as written. "v3 is dispatched after pass 3 unless pass 3 returns a BLOCKING finding that is both new and fixable without a new purchase. Anything else goes into §7's limits verbatim and the run proceeds." Pass 3 returned two such findings; both are fixed and v4 dispatches. No fourth pass is bought. An adversarial critic asked to attack will always return findings, so the stopping rule was fixed before they could be read.
1. What is being bought, and only that
E-20260816g (S201) asked whether a translator's declared account of what a device is doing predicts
which reader-side property blind reading seats report losing when it goes. It got an answer with a
hole in it, and the design said so before the run: one hand wrote the translation, the loci, the
labels and the plain alternatives. The pre-run critic's BLOCKING 1 and 2 both named that
dependency; RS-20260816g §7 limit 1 repeats it; the arm page named the remedy as the whole of step 2.
So the primary T of E-20260816g — which held at exact P = 0.00952 — licensed nothing on its
own, because the author who assigned the label at each locus also chose what replaced the stretch.
This run buys one thing: the alternative text is written by a hand that has never seen the labels.
The translation, the loci, the labels, the reading seats, the probes, the temperature, the caps, the
replicate count, the parse rule and the estimator are held identical to E-20260816g.
What that is NOT. The lead still chose which stretches to mark, still marked them, still wrote the rewriting instruction, and still assigned the labels. The only thing excluded is that the lead wrote the alternative text, and every outcome sentence in §7 says exactly that and no more (the pass-2 critic's finding 7, taken in full).
The one sentence (
continue-prompt.md§4.5). This unit measures whether a translator's stated reason for a choice tracks what that choice does to a reader, when the replacement text is written by someone who does not know the reason. That is a fact about translating. It is not about the project's statistics, raters, verifiers or published figures.
2. Materials
T-zhongli-R04-v1 — 蒲松齡〈種梨〉 ("Planting Pears"), Liaozhai zhiyi juan 1, 575 Chinese
characters, translated whole by the lead under R04, frozen 2026-08-16 at 7ace478, 686 English words.
Contamination against Giles 1880, measured after the freeze: 0 / 0 / 0 shared 7-, 12-, 15-grams,
longest common run 6 tokens, clean. Nothing in it is rewritten for this run — that is the point
of the step.
The declaration is §Device declaration of that file, committed 2026-08-16 before any design
existed: fourteen stretches, each assigned exactly one job from a closed vocabulary — MANNER,
STANCE, ORNAMENT, NONE — or marked SHAM. Three per job, three NONE, two SHAM. Unchanged.
The Chinese is materials/zhongli-zh.txt, byte-identical to E-20260816g's copy.
The arms. FULL is E-20260816g's materials/FULL.txt, byte-identical, rebuilt by the same
logic from the same frozen string and asserted equal to it in the runner. FLAT2 is
byte-identical to FULL outside the twelve non-SHAM markers, where each stretch is replaced by
the independent hand's alternative. The two SHAM markers are identical in both arms, as in
step 1.
Markers are M01–M14 in text order; the declaration ids L01–L14 are in declaration order and
appear nowhere a reading seat can see. Mapping in materials/map.json.
2a. What "an independent hand" means here, exactly
The hand is P1 openai/gpt-5.6-terra, which supplies no scored body in this run. It
receives the Chinese tale whole, the English translation with the fourteen stretches marked, and the
rule the lead wrote under, transcribed verbatim from the frozen declaration:
replace the marked stretch with the plainest English rendering of the same events that I can write; change nothing outside the stretch; add and remove no event.
It is told nothing about MANNER, STANCE, ORNAMENT or NONE, nothing about which stretches are
SHAM, nothing about a second version being read by anyone, and nothing about any hypothesis. It is
asked for all fourteen, so that its blindness to which two are SHAM is real; the two SHAM
alternatives are recorded and discarded, and the SHAM stretches stay identical in both arms.
This is a translation act by a panel model as a budgeted contrast subject (charter §3,
continue-prompt.md §5.1) and it is the only translation this design can contain: the lead's
translation is the thing being held constant, and new lead prose anywhere in these fourteen stretches
would destroy the comparison with E-20260816g that the whole step exists to make.
What this hand is NOT blind to. P1 sees the marked stretch and can infer for itself what is
salient about it — that craning his neck and staring is about manner, that your worship is
deferential. Blinding it to the four label names does not blind it to the properties those names
pick out, and the lead's own act of marking tells it where to look. So: the lead did not write the
alternative text. Nothing stronger is claimed anywhere in this design. The residual pathway — that
the label and an independent rewriter's sense of salience are two readings of the same visible
surface — is not excluded and cannot be excluded here.
2b. The lead's craft attempt, frozen before Stage A is dispatched
RS-20260816g §7 limit 3 says the stance-locus alternatives "may simply not be plain enough … a
different hand might have got closer." Before buying the different hand, the lead records one
good-faith attempt to defeat that limit himself — to write an attitude-free English of each of the
three stance propositions, keeping the events. Craft on the record, not a test.
| locus | Chinese | the device | the lead's step-1 plain version | an attitude-free attempt | verdict |
|---|---|---|---|---|---|
L04 |
於居士亦無大損 | It is no great loss to your worship** | It is no great loss to you | You will lose little. | available. The attitude at L04 is deference, carried by a form of address attached to a proposition that is not itself attitudinal. Strip the address and the assessment and a bare quantity is left. |
L05 |
良朋乞米,則怫然 | they go sour in the face** | they are displeased | their faces change | not available without dropping an event. 怫然 is resentfully. "Their faces change" is attitude-free and no longer says which way; the resentment is the event. |
L06 |
蠢爾鄉人,又何足怪 | That the countryman was a fool — what is there in that to wonder at?** | The countryman was stupid, and that is not surprising | (none found) | not available. The proposition is the evaluation. There is no English that says he was a fool and no wonder without saying he was a fool. |
Q6is EXPLORATORY and is labelled so here, before dispatch (pass-1 critic's MAJOR 7, correct). The address-form/predicate distinction was devised after seeingE-20260816g's three stance outcomes and is operationalised by one favoured locus against two. Registering it before this dispatch does not make it independent confirmation. It will not be written intoframework/v0.2whatever it does.
3. Procedure
Seats. P1 openai/gpt-5.6-terra writes and scores nothing. P2 google/gemini-3.6-flash,
P3 x-ai/grok-4.5, QR qwen/qwen3.7-max read — the three seats that supplied every scored body
in E-20260816g. P4 and P5 are out, notes (bps), (bne); GL is out on long prompts.
There is no fourth reading seat available to this project, which is the honest answer to the
pass-1 critic's MAJOR 6 and is why §7's scoping paragraph is written as hard as it is.
Every call records provider, the response's resolved model string, and a UTC dispatch
timestamp, so routing and order are auditable (pass-2 critic's finding 6).
Stage A — the independent alternatives. One call to P1, as §2a. max_tokens 3000, reasoning cap
1200 (ratio per note (bpv)). On finish_reason: length or a parse failure, one re-dispatch at
6000/2400, budgeted in advance at the dearest rate per note (bqk). A second failure is a dead run
under F3. temperature 0. The output is committed to git before anything else is bought.
Stage A′ — the mechanical pre-dispatch screen. For each of the twelve non-SHAM loci, compare the
hand's alternative with the FULL stretch after casefolding, stripping punctuation and collapsing
whitespace. A locus where they are identical is a non-subtraction: the arms do not differ there
and no Δ at it can be read. Named and excluded. Arithmetic on strings; no judgment; not adjustable
after the fact.
Stage A″ — the hand's own account, PROCESS DOCUMENTATION ONLY. A second P1 call, made after the
alternatives are frozen and committed, shows P1 its own twelve pairs and asks, for each, in ≤ 12
words, what the second version stops saying. v2 claimed this measured the residual shared-salience
pathway. That claim is WITHDRAWN (pass-2 critic's finding 4, taken in full): it is a post-hoc
self-description by the model that made the rewrite, and coding it against the lead's labels would be
the label-knowing lead scoring it unblinded. No agreement rate is computed. No coding is done. No
test is conditioned on it. The twelve replies are published verbatim in the result page's limits as
a record of what the writing hand said it was doing, and nothing is inferred from them.
Stage B — the gates and the measurements. P2 and P3 are shown, for each of the twelve altered
loci, the Chinese stretch and the two English renderings labelled A and B in an order fixed by the
locus index (odd index → A is FULL), and answer:
| field | question | status |
|---|---|---|
not_english |
is either A or B ungrammatical or not English prose? | GATE — locus excluded if flagged by both seats |
adds |
does either rendering state or imply an event, participant, object or number that the other does not state or imply at all? | GATE — excluded if flagged by both seats |
drops_named |
does either rendering omit a named participant, object, number or action that the other states? | GATE — excluded if flagged by both seats |
beyond_vividness |
beyond one being more vivid or more specific, does either change polarity, agency, evaluation, modality, or the social relation between speaker and hearer? — plus ≤ 15 words naming which | MEASUREMENT — published per locus in §7's limits, gates nothing |
same_events |
do A and B report the same events? | MEASUREMENT — kept only so the number is comparable with step 1's |
Why the global parity bar is not the gate, and what is bought instead. Notes (bpu) and
(bkr): a global content-parity bar at device-marked sites fails at those sites for every hand,
published ones included — ARM-low-pole was blocked in both its sessions by exactly this gate under
two phrasings, and four published translations of 1890–1918 were flagged by the same instrument as
omitting or misstating at 0.68 and as supplying at 0.77; E-20260816e returned SAME on 6 of 11 and
lost its design. A depiction commits to particulars a statement leaves open and a statement commits to
particulars a depiction leaves open, so plain can never pass equivalent.
The pass-2 critic is nonetheless right that adds and drops_named do not cover polarity, agency,
evaluation, modality or social relation, and that at the stance loci those are exactly what may
change. Its own stated fallback is taken, in full: the estimand is restated as the total effect of
these twelve particular rewrites, and no sentence anywhere calls the contrast "removing a device"
(§4). beyond_vividness is bought so that the threat is measured per locus and published beside the
result, rather than argued about.
A locus excluded here is not rewritten: rewriting the independent hand's English would put the lead's hand back inside the replacement, which is the one thing this run exists to keep out.
P2 and P3 also read in stage C. The calls are stateless and the stage-B prompt shows neither the
passage nor the probes; but this is the same seat doing both jobs, exactly as in E-20260816g's
stage 0, and it is held identical rather than improved, because the value of this run is the
comparison with step 1. Named in §7's limits.
The exclusion list from A′ and B is written to materials/exclusions.json and committed to git
BEFORE stage C is dispatched. After that commit no locus may be added to or removed from it except
by F2, which is mechanical.
Stage C — the reading run. Identical to E-20260816g stages 1–2 except that the second arm is
FLAT2. Each body receives one arm only, blind: no mention of a second version, of the Chinese, of
the translator, or of any hypothesis. Three yes/no probes at each of the fourteen markers, wording
transcribed unchanged:
| probe | wording put to the seat |
|---|---|
manner |
Does the marked stretch tell you anything about how something was done, or what it looked or sounded like, beyond the bare fact that it happened? |
stance |
Does the marked stretch convey an attitude — the narrator's, or a character's — toward what it describes? |
ornament |
Does the marked stretch make you suppose the original had a figure of speech, a set phrase, or sound-play at this point? |
plus one whole-body item: list any marker that is ungrammatical or does not read as English.
Three replicates × three seats × two arms = 18 bodies. max_tokens 2000, reasoning cap 900,
temperature 0.8 — replicates must be genuine samples and the seat, not the body, is the unit
of the estimator. Strict JSON; E-20260816g §3a's parse rule transcribed unchanged. An
unparseable body is dead, reported dead, counted in nothing, never re-rolled. No fourth replicate is
bought under any circumstance.
Arm order. The two arms of a (seat, replicate) cell are dispatched back-to-back, so drift
within a cell is minimal, and which arm goes first alternates on a rule fixed here: for replicate r
(1–3) and seat index k (P2=0, P3=1, QR=2), FULL first iff (r + k) is even — giving 4
FULL-first cells of 9, and 2:1 within each seat. Perfect within-seat balance is impossible at
three replicates and is not bought, because moving to two or four replicates would break the
identity with step 1 that this whole step exploits (pass-2 critic's finding 6, taken in part). The
residual imbalance and the timestamps are published.
Why FULL is re-bought rather than reused. E-20260816g's FULL bodies are four days old and
OpenRouter routes a slug to whichever provider it picks; comparing a fresh FLAT2 against stale
FULL bodies would confound the replacement with whatever moved in between. Re-buying costs nine
bodies and buys a cross-run comparability check with a prespecified margin (Q5).
4. The statistic, the estimand, and what the numbers are not
Contributing seats (pass-3 finding 4). A seat contributes only if it has at least one usable body in both arms. Every Δ below is computed over contributing seats and no others; a seat live in one arm alone is named and used in nothing.
For contributing seat s, locus i, probe q, let y_s(i,q) be that seat's mean yes-rate over its
usable bodies in an arm. Then
Δ(i,q) = mean over contributing seats of [ y_s^FULL(i,q) − y_s^FLAT2(i,q) ]
so a seat contributes once however many bodies it supplied. For the surviving job loci, d(i) is the
declared job's probe, and
T = mean over surviving job loci of [ Δ(i, d(i)) − mean of Δ(i,q) over the other two probes ]
The estimand. The contrast is between two English strings. adds and drops_named catch gross
damage; nothing gates polarity, evaluation, modality, agency or social relation, and at the stance
loci those may be exactly what differs. So:
The estimand is the total effect of these twelve particular rewrites on three probe answers. It is not "the effect of removing a device", and the phrase is not used of this contrast anywhere in this design or in the result page.
This is E-20260816g §4's own formulation, restored after v2 dropped it.
band is a reference value, not a validated detection threshold (pass-3 finding 2, taken as far
as it can be bought). Define band = max over the two SHAM loci and the three probes of |Δ(i,q)| —
six values, on the only two stretches whose text is identical in both arms.
- It is used one-sidedly and only one-sidedly. v2 used it as an equivalence margin, which is anti-conservative; v3 and v4 do not, and no registered prediction asserts equivalence.
- It is not a calibrated upper bound on null variability. Two loci is a thin base and the realized maximum of six values is unstable. It is not enlarged, because enlarging it means changing the stimulus and losing the identity with step 1 that the whole step exploits. Every outcome sentence therefore says "exceeds the sham reference", never "is detected" or "is real."
- One thing the critic reads as a defect is why it is the right reference here. The sham stretches are text-identical locally but sit in bodies that differ elsewhere, so their Δ absorbs arm-wide context and spillover. For a per-locus contrast between two passages that differ at eleven other places, a reference that includes that spillover is the appropriate one; a purely local noise estimate would be too small.
floor_q(mean over the twoSHAMloci of|Δ(i,q)|, per probe) andfloor_step1(step 1'smean over SHAM of max_q |Δ|) are reported for comparability and used in nothing.
What the 1,680 enumeration is — all test language removed (pass-3 finding 3, taken in full). v3 still called it a conditional randomisation test of association under a sharp null. The labels were chosen from the visible wording for properties aligned with the probes; they are not exchangeable, and calling the enumeration a test does not make it one.
Tand the relabeling count are an ALGORITHMIC SENSITIVITY SUMMARY and nothing else.Tis computed on the observed labelling, and the same statistic is computed on all 1,680 relabelings of the same multiset; the count reaching or exceedingTis reported. There is no null, no test, no P value, no threshold, and no evidential sentence anywhere in this design or the result page rests on that count. An inferential claim would need independently blinded labels under a prespecified rubric on held-out loci —E-20260816g's BLOCKING 1, a larger experiment, not bought here.
This is a loss and it is reported as one. Step 2 was constituted to identify step 1's primary. It
now cannot: three passes of adversarial review have established that the primary was never
identifiable from a label set the translator wrote, whoever writes the alternatives. What step 2
still does is re-test §7.16's per-locus claims — Q2 and Q3 — which are directional, registered,
and do not depend on the enumeration at all. That is the arm's payload and it was the informative half
of step 1 too.
5. Predictions, registered
Q1(algorithmic sensitivity summary; carries no evidential weight).T, and the exhaustive count of relabelings reaching or exceeding it (1,680 if all nine job loci survive), reported as arithmetic under §4.Q2— §7.16's exception, re-tested. Two components, neither an equivalence claim.Q2a— the attitude does not fall. At all threeSTANCEloci, both (i) Δstance≤ 0 — one-sided, no margin, falsified by any fall at any of the three; and (ii) theFLAT2stance yes-rate ≥ 0.50, which is a degeneracy guard, not a level claim: without it, (i) is trivially satisfied at a locus where no seat reports attitude in either arm. (The 0.50 figure is acknowledged to sit near step 1's observedFLATrates of 0.67, 1.00, 1.00; the weight ofQ2ais carried by (i), which no prior number informed.)Q2b— the source-belief probe rises, absolutely and relatively. At at least two of the threeSTANCEloci, both Δornament>band(a positive directional movement, per pass-3 finding 1 — v3 required only the difference, which a negative Δornamentcould satisfy) and Δornament− Δstance>band.Q2holds iffQ2aandQ2bboth hold, andQ2is UNAVAILABLE if either component is unavailable for any reason (pass-3 finding 7); no failure language is used in that case.Q3— a single 6-of-6 prediction (pass-3 finding 5, first option). At each of the threeMANNERand threeORNAMENTloci, the declared probe's Δ is the largest of the three at that locus and exceedsband.Q3is UNAVAILABLE if any one of the six loci is excluded; there is no per-class partial verdict.Q4(algorithmic sensitivity summary, asQ1).mean over NONE of max_q |Δ| < mean over job loci of max_q |Δ|, with the exhaustive count over which three of the surviving non-SHAMloci carryNONE(C(12,3) = 220if all survive). Requires all threeNONEloci to survive, else UNAVAILABLE.Q5— purely descriptive; establishes nothing (pass-3 finding 6, taken in full). The mean and the maximum, over all 14 loci × 3 probes, of|this run's FULL yes-rate − E-20260816g's FULL yes-rate|, computed over seats contributing to this run'sFULLarm that were also live in step 1, and reported per locus and probe. No margin, no threshold, no acceptance decision. Cross-run comparison is UNCONDITIONALLY PROHIBITED: neither the result page nor the framework will say that step 1's pattern reproduced or failed to reproduce, whateverQ5shows.Q2andQ3re-test §7.16's claims on this run's own data, which needs no cross-run comparison at all. UNAVAILABLE if fewer than two seats are eligible.Q6(EXPLORATORY; will not enter the framework in any outcome). Δstance(L04) > Δstance(L05) and Δstance(L04) > Δstance(L06).
6. The decision algorithm
Executed in this order, in analyse.py:
0. Contributing seat := a seat with >= 1 usable body in BOTH arms.
F3 — fewer than 12 usable bodies, or fewer than 2 CONTRIBUTING seats, or a second
Stage A failure -> DEAD RUN. Bodies and reason reported; every Q is UNAVAILABLE.
1. E := Stage A' non-subtractions
u Stage B loci flagged not_english / adds / drops_named by BOTH seats
u F2: loci flagged not-English by >= 3 of the 18 bodies
(E is fixed by 1; nothing may be added to it thereafter.)
2. Survivors per class computed from E.
3. |E n job loci| > 3 -> Q1 UNAVAILABLE.
4. any STANCE locus in E -> Q2 and Q6 UNAVAILABLE.
5. any of the six MANNER/ORNAMENT loci in E -> Q3 UNAVAILABLE (no partial verdict).
6. fewer than 3 NONE loci survive -> Q4 UNAVAILABLE.
7. band := max over SHAM loci x probes of |D|.
F1: band >= mean over surviving job loci of max_q |D|
-> Q1, Q2b, Q3, Q4, Q6 UNAVAILABLE (instrument outcome), and THEREFORE Q2
UNAVAILABLE as a composite. Q2a's two quantities are still computed and
reported as raw numbers, with no verdict attached.
8. sep := mean over surviving job loci of [ max_q D - min_q D ].
F4: sep <= band -> Q1 and Q3 UNAVAILABLE (instrument outcome).
9. Q5 computed and reported descriptively. Cross-run language prohibited regardless.
10. Whatever survives 3-9 is tested. EVERYTHING ELSE IS WRITTEN "UNAVAILABLE" — never as a
failed test, never as a null, never with a number presented as an outcome.
Q2 verdict branches, and when they may be used at all. The one-, two- and three-locus branches
below are available only if both Q2a and Q2b were available and evaluated (pass-3 finding 7).
Q2 fails if either component fails. Two or three stance loci failing Q2a(i) → §7.16's exception
is WITHDRAWN. Exactly one failing → RESTATED as locus-specific, naming which proposition's attitude
fell when the stretch was rewritten and which did not. Fixed here so it is not decided after the
numbers.
A run in which the hand largely declined to rewrite (|E ∩ job| > 3) is reported as a
procedural outcome — these materials did not yield an independent replacement set — and not
as a finding about the loci or about the translator.
7. What each outcome licenses
Every line below is about this translator, this tale, these fourteen loci, one independent writing
hand and three language-model reading seats sampled three times each. n at the level that matters
is THREE. Nothing here is a statement about human readers, about translators in general, or about
language models in general, and no sentence will be written as though it were. There is no fourth
reading seat available to buy. Tier D is NOT PASSED; every evaluative sentence carries
internal-judgment-only and every self-assessment provisional: true.
Q1andQ4→ arithmetic, reported, licensing nothing. §4 says why, and the result page will repeat it rather than let a small count do evidential work by implication.Q2aholds at all three loci → in this run, at the three loci the translator declaredSTANCE, the attitude probe did not fall when the stretch was replaced by another hand's plainer rendering, and at least half the reading seats reported attitude in the replaced version. That sentence is the whole claim. It does not establish that the attitude is inseparable from the proposition; it establishes that these rewrites did not separate them.Q2bholds at two or three loci → at those loci the source-belief probe rose by more than the sham reference, and rose by more than the attitude probe did. It does not establish that the translator's device caused the source-belief, only that the two versions differ that way (pass-3 finding 1).Q2holds → §7.16's exception stands as re-tested on an independently written replacement set, in the two narrow sentences above and no others.Q2fails at two or three loci → §7.16's exception is withdrawn: the attitude fell when a hand other than the lead's wrote the plainer version, which is whatRS-20260816glimit 3 allowed for.Q3fails → the manner/ornament half of §7.16 does not survive an independently written replacement, and §7.16 is restated to its floor: a translator can tell when a choice is inert and little else.Q4no longer carries evidential weight for even that, so the floor is stated as an observation.Q6→ recorded on the result page as a conjecture with the pass-1 critic's MAJOR 7 attached. It does not enterframework/v0.2in any outcome.
Published in the limits whatever happens: the per-locus beyond_vividness verdicts; the
same_events rate; the twelve Stage A″ self-reports verbatim; the arm-order imbalance and the
timestamps; the resolved provider and model of every body; every excluded locus with its reason;
every seat that failed to contribute in both arms; and the three critics' unanswerable findings
verbatim (§9, pass 3, findings 2, 3 and 6's residue).
8. Budget
Worst case from the caps the requests permit (note abc), assuming every call needs the
doubled-cap re-dispatch, priced at the dearest seat's rate (note (bqk); P3 $2.00 / $6.00
per M, QR $1.475 / $4.425, P2 $0.75 / $3.75, P1 $1.00 / $6.00).
| stage | calls | caps in / out | worst |
|---|---|---|---|
critic pass 1 (P1) |
1 | 8k / 6k | $0.063292 actual, spent |
critic pass 2 (P1) |
1 | 11k / 6k | $0.067593 actual, spent |
critic pass 3 (P1) |
1 | 15k / 6k | $0.085206 actual, spent |
A — independent alternatives (P1) |
1 | ~2,300 / 3,000 (+6,000) | $0.060 |
A″ — the hand's own account (P1) |
1 | ~1,500 / 2,000 (+4,000) | $0.038 |
B — gates + measurements (P2, P3) |
2 | ~1,800 / 3,000 (+6,000) | $0.130 |
| C — 3 replicates × 3 seats × 2 arms | 18 | ~1,400 / 2,000 (+4,000) | $0.700 |
| total | 25 | $1.144 |
Worst case $1.144 < stop-loss $1.35 < declared ceiling $1.50 (note (bpq): the stop-loss sits above the worst case, because its job is to catch an estimate that is wrong, not one that is right). Enforced inside each runner, each carrying the experiment's running total. No fourth critic pass is bought: the stopping rule at the head of this file is the guard against regress, and it was fixed before pass 3 was read.
The UTC day 2026-08-20 opened for this session at $0.4210325 of $5.00, leaving $4.5789675; the ceiling is 33% of that. Dispatch order is a gate order: A → A′ → A″ → B → commit exclusions → C.
9. Pre-run critics — three passes, no design dispatched until v4
Pass 1, on v1 — NEEDS REDESIGN, 9 findings, 3 BLOCKING, 5 MAJOR, 1 MINOR, $0.063292
| # | sev | what it said | what was done |
|---|---|---|---|
| 1 | BLOCKING | the permutation P has no randomisation basis: declared jobs are not randomly assigned but deterministically related to the wording and the probes | taken in substance in v2, taken further in v3 — see pass 2 finding 1 below, where the inferential claim is removed outright. |
| 2 | BLOCKING | P1 is blind to the label names, not to what they encode; it can remove the same salient property the lead labelled |
taken in substance. §2a says it in the design's own words and every outcome sentence is narrowed to the lead did not write the alternative text. v2's Stage A″ "measurement of the residual" is withdrawn in v3 (pass 2 finding 4). Refused: concealing the marked target or commissioning several alternatives with blind selection is a different, larger design and breaks the identity with step 1. |
| 3 | BLOCKING | event preservation is not a gate, so a probe difference may come from changed content | refused as stated, on notes (bpu) and (bkr), whose evidence is that a global parity bar at device-marked sites fails there for every hand including published ones. The critic's own fallback is taken in full in v3: the estimand is restated as the effect of these twelve rewrites, and the phrase "removing a device" is not used of the contrast. Gates adds / drops_named are bought; beyond_vividness measures the rest per locus. |
| 4 | MAJOR | F5 allows outcome-relevant exclusions after P1's output is seen; Q2/Q3/Q6 undefined when a named locus is excluded |
taken in full. Complete named locus set required; UNAVAILABLE otherwise; exclusion list committed to git before stage C; screens mechanical or two-seat adjudicated; a declining hand is a procedural outcome. |
| 5 | MAJOR | Q2 treats a large negative Δ as support; the floor is max-absolute and not probe-specific; two sham loci cannot estimate it |
taken; then taken again and harder in v3 after pass 2 showed v2's two-sided use of the band was anti-conservative. Q2 is rebuilt with no equivalence claim in it. |
| 6 | MAJOR | three seat clusters; temperature replicates are not independent readers; n = 3 | taken in full by scoping. §7 says n = 3 in those words. Not bought: no fourth reading seat exists for this project. |
| 7 | MAJOR | Q6 was devised after seeing the earlier stance outcomes |
taken in full. Q6 is EXPLORATORY and never enters the framework. |
| 8 | MAJOR | no randomisation or audit of arm order or routing; Q5 has no threshold |
taken in full, and the threshold rebuilt again in v3 (pass 2 finding 6). |
| 9 | MINOR | gates incomplete and inconsistent; a Q1 null cannot establish attribution |
taken in full, and v2's table replaced by v3's ordered algorithm after pass 2 finding 5. |
Pass 2, on v2 — NEEDS REDESIGN, 7 findings, 3 BLOCKING, 4 MAJOR, $0.067593
No finding was a restatement; the pass reviewed the machinery v2 had added.
| # | sev | what it said | what v3 did |
|---|---|---|---|
| 1 | BLOCKING | calling the permutation "a randomisation test of association" does not create a randomisation basis; v2 still used P ≤ 0.05 and still said a hold shows the declaration predicts where loss lands | taken. §4 removes every inferential word from Q1 and Q4: no P threshold, no "predicts". The reported quantity is T and the exhaustive count of relabelings reaching it. Refused: deleting the enumeration entirely — the count with its null stated is the descriptive report the remedy asks for. The full remedy (blinded labels, held-out loci) is named as what an inferential claim would need and is not bought. |
| 2 | BLOCKING | adds/drops_named do not gate polarity, causation, modality, evaluation, intensity, temporal relation or social relation, and at the stance loci those are exactly what may change |
the critic's own fallback taken in full. §4: the estimand is the total effect of these twelve rewrites; "removing a device" is not used of the contrast anywhere. beyond_vividness is added to stage B to measure the threat per locus and publish it. Refused: a full semantic-preservation gate, on (bpu) and (bkr) — it is unpassable at device sites and has already killed two of this project's designs. |
| 3 | BLOCKING | band = max over six sham values is anti-conservative as an equivalence margin, which is exactly what Q2 used it for |
taken in full, and Q2 rebuilt from scratch. band is now used only as a one-sided detection threshold. Q2a is an absolute raw-rate level (≥ 0.50) and Q2b a within-locus probe contrast; no registered prediction in v3 asserts equivalence. |
| 4 | MAJOR | Stage A″ does not measure the residual: it is a post-hoc self-description coded by the label-knowing lead, and its agreement rate will be read as mechanism evidence | taken in full, first remedy. The measurement claim is withdrawn; no agreement rate is computed and no coding is done; the twelve replies are published verbatim as process documentation. |
| 5 | MAJOR | the decision table is incomplete — no F2 row, Q4 undefined below three NONE loci, F1/F4 intersections unresolved, "reported" invites reporting a null |
taken in full. §6 replaces it with an ordered algorithm covering F1–F5 intersections, F2-derived exclusions and minimum surviving counts per class, ending in an explicit rule that an unavailable estimand is written UNAVAILABLE and never as a null. |
| 6 | MAJOR | the alternation does not balance order within seat; Q5 compares a mean against a max on incompatible scales; and §7 still used reproduction language when Q5 is unavailable |
taken; one part in part. Q5's margin is now 0.111, a yes-rate difference compared with a yes-rate difference, calibrated to the instrument's own resolution. Reproduction language is prohibited whenever Q5 is unavailable. Timestamps and resolved model recorded, arms of a cell dispatched back-to-back. In part: perfect within-seat balance needs an even replicate count, which would break the identity with step 1; the 2:1 imbalance is published instead. |
| 7 | MAJOR | "rules out author-engineered subtraction" is too broad: the lead still selected, marked and instructed | taken in full. Every outcome sentence now says only that the lead did not write the alternative text, and §1 and §2a say what remains his. |
Pass 3, on v3 — NEEDS REDESIGN, 7 findings, 2 BLOCKING, 4 MAJOR, 1 MINOR, $0.085206
The stopping rule fixed at the head of this file was applied as written. Both BLOCKING findings are new and both were fixable without a new purchase, so both are fixed and v4 is what dispatches. The three findings whose remedy cannot be bought are taken as far as they go and their residue is carried into §7's limits verbatim, as the rule provides. No fourth pass is bought.
| # | sev | what it said | what v4 did |
|---|---|---|---|
| 1 | BLOCKING | Q2a tests only a raw FLAT2 rate and so does not show the attitude survived; 0.50 is data-informed; Q2b can hold with Δornament negative, so it need not show the source-belief rose at all; yet §7 claimed the exception was "evidenced" |
taken in full, both components rebuilt. Q2a becomes Δstance ≤ 0 — one-sided, no margin, falsified by any fall at any of the three — with the ≥ 0.50 rate demoted to an explicit degeneracy guard (without it, "did not fall" is trivially true where nothing was there), and the data-informed origin of 0.50 stated. Q2b now requires Δornament > band as well as the difference. §7's Q2 licences are rewritten to two flat descriptive sentences, with "does not establish that the device caused the source-belief" written into the licence itself. |
| 2 | MAJOR (restatement; v3's answer judged inadequate) | restricting band to one-sided use does not make a realized maximum over six sham values a validated detection threshold, and sham loci sit in bodies that differ elsewhere, so their Δ absorbs arm-wide spillover |
taken as far as it is buyable; the residue goes to the limits verbatim. §4 now says band is a reference value, not a validated threshold, and every outcome sentence says "exceeds the sham reference" rather than "is detected". More sham loci cannot be bought without changing the stimulus and losing the identity with step 1. On the spillover half the finding is answered rather than conceded: for a per-locus contrast between two passages that differ at eleven other places, a reference that includes arm-wide spillover is the right one; a purely local estimate would be too small. |
| 3 | MAJOR (restatement; v3's answer judged inadequate) | v3 said "every inferential word is removed" and then called the enumeration a conditional randomisation test of association under a sharp null; §7 still turned the count into a conclusion | taken in full. All test, null and association language is gone. T and the 1,680-relabeling count are an algorithmic sensitivity summary, Q1 and Q4 carry no evidential weight, and §4 says in plain words that step 2 therefore cannot identify step 1's primary and that this is a loss, not a result. The arm's payload is Q2 and Q3, which do not touch the enumeration. |
| 4 | BLOCKING | F3 allows an analysis with two live seats per arm without requiring the same seats in both arms; the seat-clustered Δ is undefined for a seat missing an arm |
taken in full. §4 defines a contributing seat as one with at least one usable body in both arms; Δ is computed over contributing seats only; F3 now requires two contributing seats; Q5's eligibility is defined separately and is UNAVAILABLE below two eligible seats. |
| 5 | MAJOR | Q3's "6 of 6" and the algorithm's class-level unavailability contradict each other, letting a favourable three-locus class be reported after the other became unavailable |
taken in full, the critic's first option. Q3 is a single 6-of-6 prediction, UNAVAILABLE if any one of the six loci is excluded. No partial-class verdict exists. |
| 6 | MAJOR (restatement; v3's answer judged inadequate) | d5 ≤ 0.111 is an untested acceptance decision on a point estimate, and a mean over 42 cells can hide large per-locus divergence; passing it should not authorise reproduction language |
taken in full. The margin is removed. Q5 is purely descriptive — mean and maximum, reported per locus and probe — and cross-run comparison is now UNCONDITIONALLY PROHIBITED, whatever it shows. Q2 and Q3 re-test §7.16's claims on this run's own data, which needs no cross-run comparison. |
| 7 | MINOR | when F1 fires, Q2b is unavailable but composite Q2 is not explicitly set unavailable, leaving a route to narrate a Q2 failure that was really an unavailability |
taken in full. §5 and §6 step 7 now say composite Q2 is UNAVAILABLE whenever either component is, and the one-/two-/three-locus branches are usable only if both components were available and evaluated. |