Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260820b-device-function-2/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260820b-device-function-2
statusfrozen
created2026-08-20
updated2026-08-20
sensesstyle-correspondence, perceived-source-carriage, accuracy
provisionaltrue
linkswiki/arms/ARM-device-function.md, workshop/experiments/E-20260816g-device-function/design.md, wiki/findings/results/RS-20260816g-device-function.md, workshop/translations/zhongli/R04-v1/translation.md, framework/v0.2/README.md, wiki/method-notes.md, config/models.md, config/budget.md

E-20260820b — the same twelve stretches, replaced by a hand that did not write the labels

Design v4, frozen and DISPATCHED. ARM-device-function step 2 (T5), and the whole of it: the arm closes on this run either way.

v1, v2 and v3 were each reviewed by an independent adversarial critic and NONE OF THEM WAS DISPATCHED. v1: NEEDS REDESIGN, 9 findings, 3 BLOCKING ($0.063292). v2: NEEDS REDESIGN, 7 findings, 3 BLOCKING ($0.067593) — none a restatement; the pass reviewed the machinery v2 had added and broke it: v2's noise band was anti-conservative in exactly the direction the arm's payload needed. v3: NEEDS REDESIGN, 7 findings, 2 BLOCKING ($0.085206) — both new, both fixable free, both fixed here. $0.216091 on three critiques, and each one removed a claim the design was not entitled to. §9 is the full record. Every change across the three revisions is a narrowing: weaker claims, stricter gates, Q2 rebuilt twice, and Q1 — what step 2 was constituted to buy — reduced to arithmetic that licenses nothing (§4).

Pre-commitment against regress, written BEFORE pass 3 was bought and applied as written. "v3 is dispatched after pass 3 unless pass 3 returns a BLOCKING finding that is both new and fixable without a new purchase. Anything else goes into §7's limits verbatim and the run proceeds." Pass 3 returned two such findings; both are fixed and v4 dispatches. No fourth pass is bought. An adversarial critic asked to attack will always return findings, so the stopping rule was fixed before they could be read.

1. What is being bought, and only that

E-20260816g (S201) asked whether a translator's declared account of what a device is doing predicts which reader-side property blind reading seats report losing when it goes. It got an answer with a hole in it, and the design said so before the run: one hand wrote the translation, the loci, the labels and the plain alternatives. The pre-run critic's BLOCKING 1 and 2 both named that dependency; RS-20260816g §7 limit 1 repeats it; the arm page named the remedy as the whole of step 2.

So the primary T of E-20260816g — which held at exact P = 0.00952 — licensed nothing on its own, because the author who assigned the label at each locus also chose what replaced the stretch.

This run buys one thing: the alternative text is written by a hand that has never seen the labels. The translation, the loci, the labels, the reading seats, the probes, the temperature, the caps, the replicate count, the parse rule and the estimator are held identical to E-20260816g.

What that is NOT. The lead still chose which stretches to mark, still marked them, still wrote the rewriting instruction, and still assigned the labels. The only thing excluded is that the lead wrote the alternative text, and every outcome sentence in §7 says exactly that and no more (the pass-2 critic's finding 7, taken in full).

The one sentence (continue-prompt.md §4.5). This unit measures whether a translator's stated reason for a choice tracks what that choice does to a reader, when the replacement text is written by someone who does not know the reason. That is a fact about translating. It is not about the project's statistics, raters, verifiers or published figures.

2. Materials

T-zhongli-R04-v1 — 蒲松齡〈種梨〉 ("Planting Pears"), Liaozhai zhiyi juan 1, 575 Chinese characters, translated whole by the lead under R04, frozen 2026-08-16 at 7ace478, 686 English words. Contamination against Giles 1880, measured after the freeze: 0 / 0 / 0 shared 7-, 12-, 15-grams, longest common run 6 tokens, clean. Nothing in it is rewritten for this run — that is the point of the step.

The declaration is §Device declaration of that file, committed 2026-08-16 before any design existed: fourteen stretches, each assigned exactly one job from a closed vocabulary — MANNER, STANCE, ORNAMENT, NONE — or marked SHAM. Three per job, three NONE, two SHAM. Unchanged.

The Chinese is materials/zhongli-zh.txt, byte-identical to E-20260816g's copy.

The arms. FULL is E-20260816g's materials/FULL.txt, byte-identical, rebuilt by the same logic from the same frozen string and asserted equal to it in the runner. FLAT2 is byte-identical to FULL outside the twelve non-SHAM markers, where each stretch is replaced by the independent hand's alternative. The two SHAM markers are identical in both arms, as in step 1.

Markers are M01–M14 in text order; the declaration ids L01–L14 are in declaration order and appear nowhere a reading seat can see. Mapping in materials/map.json.

2a. What "an independent hand" means here, exactly

The hand is P1 openai/gpt-5.6-terra, which supplies no scored body in this run. It receives the Chinese tale whole, the English translation with the fourteen stretches marked, and the rule the lead wrote under, transcribed verbatim from the frozen declaration:

replace the marked stretch with the plainest English rendering of the same events that I can write; change nothing outside the stretch; add and remove no event.

It is told nothing about MANNER, STANCE, ORNAMENT or NONE, nothing about which stretches are SHAM, nothing about a second version being read by anyone, and nothing about any hypothesis. It is asked for all fourteen, so that its blindness to which two are SHAM is real; the two SHAM alternatives are recorded and discarded, and the SHAM stretches stay identical in both arms.

This is a translation act by a panel model as a budgeted contrast subject (charter §3, continue-prompt.md §5.1) and it is the only translation this design can contain: the lead's translation is the thing being held constant, and new lead prose anywhere in these fourteen stretches would destroy the comparison with E-20260816g that the whole step exists to make.

What this hand is NOT blind to. P1 sees the marked stretch and can infer for itself what is salient about it — that craning his neck and staring is about manner, that your worship is deferential. Blinding it to the four label names does not blind it to the properties those names pick out, and the lead's own act of marking tells it where to look. So: the lead did not write the alternative text. Nothing stronger is claimed anywhere in this design. The residual pathway — that the label and an independent rewriter's sense of salience are two readings of the same visible surface — is not excluded and cannot be excluded here.

2b. The lead's craft attempt, frozen before Stage A is dispatched

RS-20260816g §7 limit 3 says the stance-locus alternatives "may simply not be plain enough … a different hand might have got closer." Before buying the different hand, the lead records one good-faith attempt to defeat that limit himself — to write an attitude-free English of each of the three stance propositions, keeping the events. Craft on the record, not a test.

locus Chinese the device the lead's step-1 plain version an attitude-free attempt verdict
L04 於居士亦無大損 It is no great loss to your worship** It is no great loss to you You will lose little. available. The attitude at L04 is deference, carried by a form of address attached to a proposition that is not itself attitudinal. Strip the address and the assessment and a bare quantity is left.
L05 良朋乞米,則怫然 they go sour in the face** they are displeased their faces change not available without dropping an event. 怫然 is resentfully. "Their faces change" is attitude-free and no longer says which way; the resentment is the event.
L06 蠢爾鄉人,又何足怪 That the countryman was a fool — what is there in that to wonder at?** The countryman was stupid, and that is not surprising (none found) not available. The proposition is the evaluation. There is no English that says he was a fool and no wonder without saying he was a fool.

Q6 is EXPLORATORY and is labelled so here, before dispatch (pass-1 critic's MAJOR 7, correct). The address-form/predicate distinction was devised after seeing E-20260816g's three stance outcomes and is operationalised by one favoured locus against two. Registering it before this dispatch does not make it independent confirmation. It will not be written into framework/v0.2 whatever it does.

3. Procedure

Seats. P1 openai/gpt-5.6-terra writes and scores nothing. P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5, QR qwen/qwen3.7-max read — the three seats that supplied every scored body in E-20260816g. P4 and P5 are out, notes (bps), (bne); GL is out on long prompts. There is no fourth reading seat available to this project, which is the honest answer to the pass-1 critic's MAJOR 6 and is why §7's scoping paragraph is written as hard as it is.

Every call records provider, the response's resolved model string, and a UTC dispatch timestamp, so routing and order are auditable (pass-2 critic's finding 6).

Stage A — the independent alternatives. One call to P1, as §2a. max_tokens 3000, reasoning cap 1200 (ratio per note (bpv)). On finish_reason: length or a parse failure, one re-dispatch at 6000/2400, budgeted in advance at the dearest rate per note (bqk). A second failure is a dead run under F3. temperature 0. The output is committed to git before anything else is bought.

Stage A′ — the mechanical pre-dispatch screen. For each of the twelve non-SHAM loci, compare the hand's alternative with the FULL stretch after casefolding, stripping punctuation and collapsing whitespace. A locus where they are identical is a non-subtraction: the arms do not differ there and no Δ at it can be read. Named and excluded. Arithmetic on strings; no judgment; not adjustable after the fact.

Stage A″ — the hand's own account, PROCESS DOCUMENTATION ONLY. A second P1 call, made after the alternatives are frozen and committed, shows P1 its own twelve pairs and asks, for each, in ≤ 12 words, what the second version stops saying. v2 claimed this measured the residual shared-salience pathway. That claim is WITHDRAWN (pass-2 critic's finding 4, taken in full): it is a post-hoc self-description by the model that made the rewrite, and coding it against the lead's labels would be the label-knowing lead scoring it unblinded. No agreement rate is computed. No coding is done. No test is conditioned on it. The twelve replies are published verbatim in the result page's limits as a record of what the writing hand said it was doing, and nothing is inferred from them.

Stage B — the gates and the measurements. P2 and P3 are shown, for each of the twelve altered loci, the Chinese stretch and the two English renderings labelled A and B in an order fixed by the locus index (odd index → A is FULL), and answer:

field question status
not_english is either A or B ungrammatical or not English prose? GATE — locus excluded if flagged by both seats
adds does either rendering state or imply an event, participant, object or number that the other does not state or imply at all? GATE — excluded if flagged by both seats
drops_named does either rendering omit a named participant, object, number or action that the other states? GATE — excluded if flagged by both seats
beyond_vividness beyond one being more vivid or more specific, does either change polarity, agency, evaluation, modality, or the social relation between speaker and hearer? — plus ≤ 15 words naming which MEASUREMENT — published per locus in §7's limits, gates nothing
same_events do A and B report the same events? MEASUREMENT — kept only so the number is comparable with step 1's

Why the global parity bar is not the gate, and what is bought instead. Notes (bpu) and (bkr): a global content-parity bar at device-marked sites fails at those sites for every hand, published ones included — ARM-low-pole was blocked in both its sessions by exactly this gate under two phrasings, and four published translations of 1890–1918 were flagged by the same instrument as omitting or misstating at 0.68 and as supplying at 0.77; E-20260816e returned SAME on 6 of 11 and lost its design. A depiction commits to particulars a statement leaves open and a statement commits to particulars a depiction leaves open, so plain can never pass equivalent.

The pass-2 critic is nonetheless right that adds and drops_named do not cover polarity, agency, evaluation, modality or social relation, and that at the stance loci those are exactly what may change. Its own stated fallback is taken, in full: the estimand is restated as the total effect of these twelve particular rewrites, and no sentence anywhere calls the contrast "removing a device" (§4). beyond_vividness is bought so that the threat is measured per locus and published beside the result, rather than argued about.

A locus excluded here is not rewritten: rewriting the independent hand's English would put the lead's hand back inside the replacement, which is the one thing this run exists to keep out.

P2 and P3 also read in stage C. The calls are stateless and the stage-B prompt shows neither the passage nor the probes; but this is the same seat doing both jobs, exactly as in E-20260816g's stage 0, and it is held identical rather than improved, because the value of this run is the comparison with step 1. Named in §7's limits.

The exclusion list from A′ and B is written to materials/exclusions.json and committed to git BEFORE stage C is dispatched. After that commit no locus may be added to or removed from it except by F2, which is mechanical.

Stage C — the reading run. Identical to E-20260816g stages 1–2 except that the second arm is FLAT2. Each body receives one arm only, blind: no mention of a second version, of the Chinese, of the translator, or of any hypothesis. Three yes/no probes at each of the fourteen markers, wording transcribed unchanged:

probe wording put to the seat
manner Does the marked stretch tell you anything about how something was done, or what it looked or sounded like, beyond the bare fact that it happened?
stance Does the marked stretch convey an attitude — the narrator's, or a character's — toward what it describes?
ornament Does the marked stretch make you suppose the original had a figure of speech, a set phrase, or sound-play at this point?

plus one whole-body item: list any marker that is ungrammatical or does not read as English.

Three replicates × three seats × two arms = 18 bodies. max_tokens 2000, reasoning cap 900, temperature 0.8 — replicates must be genuine samples and the seat, not the body, is the unit of the estimator. Strict JSON; E-20260816g §3a's parse rule transcribed unchanged. An unparseable body is dead, reported dead, counted in nothing, never re-rolled. No fourth replicate is bought under any circumstance.

Arm order. The two arms of a (seat, replicate) cell are dispatched back-to-back, so drift within a cell is minimal, and which arm goes first alternates on a rule fixed here: for replicate r (1–3) and seat index k (P2=0, P3=1, QR=2), FULL first iff (r + k) is even — giving 4 FULL-first cells of 9, and 2:1 within each seat. Perfect within-seat balance is impossible at three replicates and is not bought, because moving to two or four replicates would break the identity with step 1 that this whole step exploits (pass-2 critic's finding 6, taken in part). The residual imbalance and the timestamps are published.

Why FULL is re-bought rather than reused. E-20260816g's FULL bodies are four days old and OpenRouter routes a slug to whichever provider it picks; comparing a fresh FLAT2 against stale FULL bodies would confound the replacement with whatever moved in between. Re-buying costs nine bodies and buys a cross-run comparability check with a prespecified margin (Q5).

4. The statistic, the estimand, and what the numbers are not

Contributing seats (pass-3 finding 4). A seat contributes only if it has at least one usable body in both arms. Every Δ below is computed over contributing seats and no others; a seat live in one arm alone is named and used in nothing.

For contributing seat s, locus i, probe q, let y_s(i,q) be that seat's mean yes-rate over its usable bodies in an arm. Then

Δ(i,q) = mean over contributing seats of [ y_s^FULL(i,q) − y_s^FLAT2(i,q) ]

so a seat contributes once however many bodies it supplied. For the surviving job loci, d(i) is the declared job's probe, and

T = mean over surviving job loci of [ Δ(i, d(i)) − mean of Δ(i,q) over the other two probes ]

The estimand. The contrast is between two English strings. adds and drops_named catch gross damage; nothing gates polarity, evaluation, modality, agency or social relation, and at the stance loci those may be exactly what differs. So:

The estimand is the total effect of these twelve particular rewrites on three probe answers. It is not "the effect of removing a device", and the phrase is not used of this contrast anywhere in this design or in the result page.

This is E-20260816g §4's own formulation, restored after v2 dropped it.

band is a reference value, not a validated detection threshold (pass-3 finding 2, taken as far as it can be bought). Define band = max over the two SHAM loci and the three probes of |Δ(i,q)| — six values, on the only two stretches whose text is identical in both arms.

What the 1,680 enumeration is — all test language removed (pass-3 finding 3, taken in full). v3 still called it a conditional randomisation test of association under a sharp null. The labels were chosen from the visible wording for properties aligned with the probes; they are not exchangeable, and calling the enumeration a test does not make it one.

T and the relabeling count are an ALGORITHMIC SENSITIVITY SUMMARY and nothing else. T is computed on the observed labelling, and the same statistic is computed on all 1,680 relabelings of the same multiset; the count reaching or exceeding T is reported. There is no null, no test, no P value, no threshold, and no evidential sentence anywhere in this design or the result page rests on that count. An inferential claim would need independently blinded labels under a prespecified rubric on held-out loci — E-20260816g's BLOCKING 1, a larger experiment, not bought here.

This is a loss and it is reported as one. Step 2 was constituted to identify step 1's primary. It now cannot: three passes of adversarial review have established that the primary was never identifiable from a label set the translator wrote, whoever writes the alternatives. What step 2 still does is re-test §7.16's per-locus claims — Q2 and Q3 — which are directional, registered, and do not depend on the enumeration at all. That is the arm's payload and it was the informative half of step 1 too.

5. Predictions, registered

6. The decision algorithm

Executed in this order, in analyse.py:

0.  Contributing seat := a seat with >= 1 usable body in BOTH arms.
    F3 — fewer than 12 usable bodies, or fewer than 2 CONTRIBUTING seats, or a second
    Stage A failure  ->  DEAD RUN. Bodies and reason reported; every Q is UNAVAILABLE.
1.  E  :=  Stage A' non-subtractions
         u  Stage B loci flagged not_english / adds / drops_named by BOTH seats
         u  F2: loci flagged not-English by >= 3 of the 18 bodies
    (E is fixed by 1; nothing may be added to it thereafter.)
2.  Survivors per class computed from E.
3.  |E n job loci| > 3                      ->  Q1 UNAVAILABLE.
4.  any STANCE locus in E                   ->  Q2 and Q6 UNAVAILABLE.
5.  any of the six MANNER/ORNAMENT loci in E ->  Q3 UNAVAILABLE (no partial verdict).
6.  fewer than 3 NONE loci survive          ->  Q4 UNAVAILABLE.
7.  band := max over SHAM loci x probes of |D|.
    F1: band >= mean over surviving job loci of max_q |D|
        ->  Q1, Q2b, Q3, Q4, Q6 UNAVAILABLE (instrument outcome), and THEREFORE Q2
            UNAVAILABLE as a composite. Q2a's two quantities are still computed and
            reported as raw numbers, with no verdict attached.
8.  sep := mean over surviving job loci of [ max_q D - min_q D ].
    F4: sep <= band  ->  Q1 and Q3 UNAVAILABLE (instrument outcome).
9.  Q5 computed and reported descriptively. Cross-run language prohibited regardless.
10. Whatever survives 3-9 is tested. EVERYTHING ELSE IS WRITTEN "UNAVAILABLE" — never as a
    failed test, never as a null, never with a number presented as an outcome.

Q2 verdict branches, and when they may be used at all. The one-, two- and three-locus branches below are available only if both Q2a and Q2b were available and evaluated (pass-3 finding 7). Q2 fails if either component fails. Two or three stance loci failing Q2a(i) → §7.16's exception is WITHDRAWN. Exactly one failing → RESTATED as locus-specific, naming which proposition's attitude fell when the stretch was rewritten and which did not. Fixed here so it is not decided after the numbers.

A run in which the hand largely declined to rewrite (|E ∩ job| > 3) is reported as a procedural outcome — these materials did not yield an independent replacement set — and not as a finding about the loci or about the translator.

7. What each outcome licenses

Every line below is about this translator, this tale, these fourteen loci, one independent writing hand and three language-model reading seats sampled three times each. n at the level that matters is THREE. Nothing here is a statement about human readers, about translators in general, or about language models in general, and no sentence will be written as though it were. There is no fourth reading seat available to buy. Tier D is NOT PASSED; every evaluative sentence carries internal-judgment-only and every self-assessment provisional: true.

Published in the limits whatever happens: the per-locus beyond_vividness verdicts; the same_events rate; the twelve Stage A″ self-reports verbatim; the arm-order imbalance and the timestamps; the resolved provider and model of every body; every excluded locus with its reason; every seat that failed to contribute in both arms; and the three critics' unanswerable findings verbatim (§9, pass 3, findings 2, 3 and 6's residue).

8. Budget

Worst case from the caps the requests permit (note abc), assuming every call needs the doubled-cap re-dispatch, priced at the dearest seat's rate (note (bqk); P3 $2.00 / $6.00 per M, QR $1.475 / $4.425, P2 $0.75 / $3.75, P1 $1.00 / $6.00).

stage calls caps in / out worst
critic pass 1 (P1) 1 8k / 6k $0.063292 actual, spent
critic pass 2 (P1) 1 11k / 6k $0.067593 actual, spent
critic pass 3 (P1) 1 15k / 6k $0.085206 actual, spent
A — independent alternatives (P1) 1 ~2,300 / 3,000 (+6,000) $0.060
A″ — the hand's own account (P1) 1 ~1,500 / 2,000 (+4,000) $0.038
B — gates + measurements (P2, P3) 2 ~1,800 / 3,000 (+6,000) $0.130
C — 3 replicates × 3 seats × 2 arms 18 ~1,400 / 2,000 (+4,000) $0.700
total 25 $1.144

Worst case $1.144 < stop-loss $1.35 < declared ceiling $1.50 (note (bpq): the stop-loss sits above the worst case, because its job is to catch an estimate that is wrong, not one that is right). Enforced inside each runner, each carrying the experiment's running total. No fourth critic pass is bought: the stopping rule at the head of this file is the guard against regress, and it was fixed before pass 3 was read.

The UTC day 2026-08-20 opened for this session at $0.4210325 of $5.00, leaving $4.5789675; the ceiling is 33% of that. Dispatch order is a gate order: A → A′ → A″ → B → commit exclusions → C.

9. Pre-run critics — three passes, no design dispatched until v4

Pass 1, on v1 — NEEDS REDESIGN, 9 findings, 3 BLOCKING, 5 MAJOR, 1 MINOR, $0.063292

# sev what it said what was done
1 BLOCKING the permutation P has no randomisation basis: declared jobs are not randomly assigned but deterministically related to the wording and the probes taken in substance in v2, taken further in v3 — see pass 2 finding 1 below, where the inferential claim is removed outright.
2 BLOCKING P1 is blind to the label names, not to what they encode; it can remove the same salient property the lead labelled taken in substance. §2a says it in the design's own words and every outcome sentence is narrowed to the lead did not write the alternative text. v2's Stage A″ "measurement of the residual" is withdrawn in v3 (pass 2 finding 4). Refused: concealing the marked target or commissioning several alternatives with blind selection is a different, larger design and breaks the identity with step 1.
3 BLOCKING event preservation is not a gate, so a probe difference may come from changed content refused as stated, on notes (bpu) and (bkr), whose evidence is that a global parity bar at device-marked sites fails there for every hand including published ones. The critic's own fallback is taken in full in v3: the estimand is restated as the effect of these twelve rewrites, and the phrase "removing a device" is not used of the contrast. Gates adds / drops_named are bought; beyond_vividness measures the rest per locus.
4 MAJOR F5 allows outcome-relevant exclusions after P1's output is seen; Q2/Q3/Q6 undefined when a named locus is excluded taken in full. Complete named locus set required; UNAVAILABLE otherwise; exclusion list committed to git before stage C; screens mechanical or two-seat adjudicated; a declining hand is a procedural outcome.
5 MAJOR Q2 treats a large negative Δ as support; the floor is max-absolute and not probe-specific; two sham loci cannot estimate it taken; then taken again and harder in v3 after pass 2 showed v2's two-sided use of the band was anti-conservative. Q2 is rebuilt with no equivalence claim in it.
6 MAJOR three seat clusters; temperature replicates are not independent readers; n = 3 taken in full by scoping. §7 says n = 3 in those words. Not bought: no fourth reading seat exists for this project.
7 MAJOR Q6 was devised after seeing the earlier stance outcomes taken in full. Q6 is EXPLORATORY and never enters the framework.
8 MAJOR no randomisation or audit of arm order or routing; Q5 has no threshold taken in full, and the threshold rebuilt again in v3 (pass 2 finding 6).
9 MINOR gates incomplete and inconsistent; a Q1 null cannot establish attribution taken in full, and v2's table replaced by v3's ordered algorithm after pass 2 finding 5.

Pass 2, on v2 — NEEDS REDESIGN, 7 findings, 3 BLOCKING, 4 MAJOR, $0.067593

No finding was a restatement; the pass reviewed the machinery v2 had added.

# sev what it said what v3 did
1 BLOCKING calling the permutation "a randomisation test of association" does not create a randomisation basis; v2 still used P ≤ 0.05 and still said a hold shows the declaration predicts where loss lands taken. §4 removes every inferential word from Q1 and Q4: no P threshold, no "predicts". The reported quantity is T and the exhaustive count of relabelings reaching it. Refused: deleting the enumeration entirely — the count with its null stated is the descriptive report the remedy asks for. The full remedy (blinded labels, held-out loci) is named as what an inferential claim would need and is not bought.
2 BLOCKING adds/drops_named do not gate polarity, causation, modality, evaluation, intensity, temporal relation or social relation, and at the stance loci those are exactly what may change the critic's own fallback taken in full. §4: the estimand is the total effect of these twelve rewrites; "removing a device" is not used of the contrast anywhere. beyond_vividness is added to stage B to measure the threat per locus and publish it. Refused: a full semantic-preservation gate, on (bpu) and (bkr) — it is unpassable at device sites and has already killed two of this project's designs.
3 BLOCKING band = max over six sham values is anti-conservative as an equivalence margin, which is exactly what Q2 used it for taken in full, and Q2 rebuilt from scratch. band is now used only as a one-sided detection threshold. Q2a is an absolute raw-rate level (≥ 0.50) and Q2b a within-locus probe contrast; no registered prediction in v3 asserts equivalence.
4 MAJOR Stage A″ does not measure the residual: it is a post-hoc self-description coded by the label-knowing lead, and its agreement rate will be read as mechanism evidence taken in full, first remedy. The measurement claim is withdrawn; no agreement rate is computed and no coding is done; the twelve replies are published verbatim as process documentation.
5 MAJOR the decision table is incomplete — no F2 row, Q4 undefined below three NONE loci, F1/F4 intersections unresolved, "reported" invites reporting a null taken in full. §6 replaces it with an ordered algorithm covering F1–F5 intersections, F2-derived exclusions and minimum surviving counts per class, ending in an explicit rule that an unavailable estimand is written UNAVAILABLE and never as a null.
6 MAJOR the alternation does not balance order within seat; Q5 compares a mean against a max on incompatible scales; and §7 still used reproduction language when Q5 is unavailable taken; one part in part. Q5's margin is now 0.111, a yes-rate difference compared with a yes-rate difference, calibrated to the instrument's own resolution. Reproduction language is prohibited whenever Q5 is unavailable. Timestamps and resolved model recorded, arms of a cell dispatched back-to-back. In part: perfect within-seat balance needs an even replicate count, which would break the identity with step 1; the 2:1 imbalance is published instead.
7 MAJOR "rules out author-engineered subtraction" is too broad: the lead still selected, marked and instructed taken in full. Every outcome sentence now says only that the lead did not write the alternative text, and §1 and §2a say what remains his.

Pass 3, on v3 — NEEDS REDESIGN, 7 findings, 2 BLOCKING, 4 MAJOR, 1 MINOR, $0.085206

The stopping rule fixed at the head of this file was applied as written. Both BLOCKING findings are new and both were fixable without a new purchase, so both are fixed and v4 is what dispatches. The three findings whose remedy cannot be bought are taken as far as they go and their residue is carried into §7's limits verbatim, as the rule provides. No fourth pass is bought.

# sev what it said what v4 did
1 BLOCKING Q2a tests only a raw FLAT2 rate and so does not show the attitude survived; 0.50 is data-informed; Q2b can hold with Δornament negative, so it need not show the source-belief rose at all; yet §7 claimed the exception was "evidenced" taken in full, both components rebuilt. Q2a becomes Δstance ≤ 0 — one-sided, no margin, falsified by any fall at any of the three — with the ≥ 0.50 rate demoted to an explicit degeneracy guard (without it, "did not fall" is trivially true where nothing was there), and the data-informed origin of 0.50 stated. Q2b now requires Δornament > band as well as the difference. §7's Q2 licences are rewritten to two flat descriptive sentences, with "does not establish that the device caused the source-belief" written into the licence itself.
2 MAJOR (restatement; v3's answer judged inadequate) restricting band to one-sided use does not make a realized maximum over six sham values a validated detection threshold, and sham loci sit in bodies that differ elsewhere, so their Δ absorbs arm-wide spillover taken as far as it is buyable; the residue goes to the limits verbatim. §4 now says band is a reference value, not a validated threshold, and every outcome sentence says "exceeds the sham reference" rather than "is detected". More sham loci cannot be bought without changing the stimulus and losing the identity with step 1. On the spillover half the finding is answered rather than conceded: for a per-locus contrast between two passages that differ at eleven other places, a reference that includes arm-wide spillover is the right one; a purely local estimate would be too small.
3 MAJOR (restatement; v3's answer judged inadequate) v3 said "every inferential word is removed" and then called the enumeration a conditional randomisation test of association under a sharp null; §7 still turned the count into a conclusion taken in full. All test, null and association language is gone. T and the 1,680-relabeling count are an algorithmic sensitivity summary, Q1 and Q4 carry no evidential weight, and §4 says in plain words that step 2 therefore cannot identify step 1's primary and that this is a loss, not a result. The arm's payload is Q2 and Q3, which do not touch the enumeration.
4 BLOCKING F3 allows an analysis with two live seats per arm without requiring the same seats in both arms; the seat-clustered Δ is undefined for a seat missing an arm taken in full. §4 defines a contributing seat as one with at least one usable body in both arms; Δ is computed over contributing seats only; F3 now requires two contributing seats; Q5's eligibility is defined separately and is UNAVAILABLE below two eligible seats.
5 MAJOR Q3's "6 of 6" and the algorithm's class-level unavailability contradict each other, letting a favourable three-locus class be reported after the other became unavailable taken in full, the critic's first option. Q3 is a single 6-of-6 prediction, UNAVAILABLE if any one of the six loci is excluded. No partial-class verdict exists.
6 MAJOR (restatement; v3's answer judged inadequate) d5 ≤ 0.111 is an untested acceptance decision on a point estimate, and a mean over 42 cells can hide large per-locus divergence; passing it should not authorise reproduction language taken in full. The margin is removed. Q5 is purely descriptive — mean and maximum, reported per locus and probe — and cross-run comparison is now UNCONDITIONALLY PROHIBITED, whatever it shows. Q2 and Q3 re-test §7.16's claims on this run's own data, which needs no cross-run comparison.
7 MINOR when F1 fires, Q2b is unavailable but composite Q2 is not explicitly set unavailable, leaving a route to narrate a Q2 failure that was really an unavailability taken in full. §5 and §6 step 7 now say composite Q2 is UNAVAILABLE whenever either component is, and the one-/two-/three-locus branches are usable only if both components were available and evaluated.