Repository path: workshop/experiments/E-20260809h-rule-execution/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260809h-rule-execution |
| status | frozen |
| created | 2026-08-09 |
| updated | 2026-08-09 |
| track | T3 |
| senses | accuracy, naturalness, style-correspondence, perceived-source-carriage |
| provisional | true |
| internal-judgment-only | true |
| links | wiki/arms/ARM-rule-execution.md, wiki/findings/results/RS-20260809b-programme-tax.md, wiki/findings/results/RS-20260808c-sense-tradeoff-de.md, workshop/experiments/E-20260809b-programme-tax/design.md, workshop/translations/kronenwaechter-abschied/R06-v1/translation.md, workshop/translations/kronenwaechter-abschied/R08-v1/translation.md, workshop/regimes/R06-lead-single-pass.md, workshop/regimes/R08-resistancy.md, config/models.md, wiki/method-notes.md |
E-20260809h — can a translation programme be executed by a hand that did not write it, and what does executing it cost?
Frozen before any generation or scoring call. The lead's two arms were frozen first, at commit
b3df80a, before this file existed. The retrieval probe (§2) ran before either arm was begun.
1. The question, and why it is not last session's
RS-20260809b (S141) put a three-regime ladder on a German source with no reachable published
English and found that the unruled arm's accuracy advantage is a property of the lead: the
lead's tax is +0.714 on one source and +0.354 on the other, and one unbriefed independent hand's
is +0.024 and +0.167, neither interval excluding zero. Method note (bkz) was written from it:
a finding about what a REGIME does needs at least one arm of that regime executed by a hand that is
not the lead.
That run had one such hand. Two readings predict its nulls, and they are opposite:
- HAND — following a declared programme costs this translator accuracy and costs translators in general nothing. The programme is innocent; the lead's rule-following is the cost.
- NON-EXECUTION — the unbriefed hand did not translate under the programme at all. Handed ten numbered rules and nothing else, it produced prose that is not measurably more source-oriented than its own unruled prose. A hand that did not execute the programme cannot have paid for it, and its null is not evidence about what programmes cost.
These are not distinguishable in a design that never measures whether the programme was
executed. S141 did not measure it. Neither did S129 or S134. RS-20260808c's §6 notes in passing
that the same rule set moved an unbriefed hand's blind naturalness by 0.810 and its
perceived-source-carriage by only 0.476, "an eighth of the lead's 3.76" — which is the
NON-EXECUTION reading showing through the floorboards of a run built for something else.
What this run teaches about translating literature. A translation programme — Venuti's
resistancy, or any framework's numbered recommendations, including the ones framework/v0.2
publishes — is a text handed to a translator who did not write it. Whether such a text can be
executed from itself, and what executing it costs in propositional accuracy, is the question the
framework's whole output form depends on. If a programme can only be executed by its author, a
framework of numbered rules is not a deliverable; it is a private notebook.
The unit's sentence (subject rule, continue-prompt.md §4.5): this unit measures whether a
written foreignizing programme changes seven translators' English at all, and what it costs their
accuracy when it does. That is a statement about translating, not about the project's apparatus.
2. Materials
Source. Arnim, «Die Kronenwächter», Erster Band, Erstes Buch, 7. Geschichte: Der Sturm (1817),
the closing three paragraphs, 751 German words in eight segments — the span immediately
after the one S141 used, with no overlap. Copy-text zeno.org, collated against the 1857 archive.org
Fraktur scan: 24 of 24 sentences present, character-ratio min 0.8913 / median 0.9569 / max
1.0000 (collate.py, materials/collation.json). The collation bounds omission and does not
certify orthography.
Why this source. It is the only German prose the project has established as having no
reachable published English translation at all (S141 §2), which removes the recall pathway that
RS-20260808c limit 4 left open, and it holds the language pair, the author's circle, the prose
type and both regimes constant against S141 and S134.
M1, the retrieval probe — run before either lead arm was translated. The eight German
segments, shuffled, to each of the seven hands: do you know of a published English translation of
this passage; if so quote its opening clause. Result, runs/probe-H*.json, 7 of 7 bodies clean:
| hand | slug | segments claimed known |
|---|---|---|
| H1 | x-ai/grok-4.5 |
0 of 8 |
| H2 | moonshotai/kimi-k3 |
0 of 8 |
| H3 | mistralai/mistral-medium-3-5 |
4 of 8, all attributed to "Roy Pascal" |
| H4 | qwen/qwen3.7-max |
0 of 8 |
| H5 | z-ai/glm-5.2 |
0 of 8 |
| H6 | nvidia/nemotron-3-ultra-550b-a55b |
0 of 8 |
| H7 | minimax/minimax-m3 |
0 of 8 |
The registered void condition of S141's probe — any plausible English quotation for a segment — is therefore triggered, by one hand, and it is reported rather than explained away. Three reasons it is read as confabulation, all checkable:
- The same slug answered NONE on 8 of 8 Arnim segments at S141, on the immediately preceding span of the same chapter, under the same prompt string. A memory that appears between two adjacent spans of one chapter is not a memory.
- No Roy Pascal translation of «Die Kronenwächter» is locatable. Pascal was a real Germanist; no catalogue reachable from this container attributes this work to him, and no English «Kronenwächter» is attributed to anyone.
- The quoted English misreads the German in the way an on-the-spot rendering would. Its S3 quotation — "Most of the people who had up to then treated the stranger as one of their own followed him" — reverses the relation of «welche die Fremde bis dahin als die Ihren behandelt hatte», in which the stranger is the subject and the people are her own. A published translation by a Germanist does not invert that clause; a model translating a fronted relative clause on the fly does.
Registered response, fixed before any generation call: H3 stays in. P1, P2 and P3 are
reported both over all seven hands and with H3 removed, and if the two differ in sign or in
whether a bar is cleared, the H3-removed version governs the licensed sentence.
Arms — eighteen, each of eight segments, 144 items.
| arm | who | what |
|---|---|---|
LEAD-R06, LEAD-R08 |
the lead | frozen at b3df80a with logs, before this design existed |
H1-R06 … H7-R08 |
seven independent hands | one unruled arm and one R08 arm each |
WRONG |
operator on LEAD-R06 |
eight single substitutions, each changing propositional content |
CLUNKY |
operator on LEAD-R06 |
eight single substitutions, each mangling English word order and changing no content |
The seven hands are seven labs, and none of them is a seat's lab — x-ai, moonshotai,
mistralai, qwen, z-ai, nvidia, minimax, against seats at openai, google and
deepseek. Declared fall-through if a hand returns malformed segments twice:
meta-llama/llama-4-maverick, once, then the hand is dropped and n is reduced before any P
value is computed (F4).
The prompts the hands receive are not new strings. The unruled preamble is imported from
E-20260809b's runner; the ruled preamble and the IND-R08 rule block are imported from
E-20260808c's; the Kleist→Arnim retarget is E-20260809b's .replace chain, copied verbatim. A
hand is given the German and the rules and nothing else — no mention of a comparison, a second arm,
or this question.
3. Instrument
SCORE_PROMPT, DEFS and SENSES are imported from E-20260808c's runner, unchanged: four
senses (accuracy, naturalness, style-correspondence, perceived-source-carriage), 1–7
integers, German shown, naturalness instructed to be judged on the English alone. Three seats,
J1 openai/gpt-5.6-terra, J2 google/gemini-3.6-flash, J3 deepseek/deepseek-v4-pro — the
same three as S134 and S141.
Blinding. Item ids are opaque and assigned after a deterministic shuffle; each seat gets its own
shuffle of all 144 items in eight blocks of eighteen; no item carries an arm, a hand, or a regime.
Nothing in an item names a programme. Seats are told that the same German may appear rendered by
different hands, which is E-20260808c's string and not an addition.
Known instrument limits, imported and not re-argued. perceived-source-carriage has never
been through Tier D and its own entry records a 0.40–0.60 false-positive rate on manufactured
oddity; Tier D is NOT PASSED; CLUNKY-style mangling leaks ≈0.55–0.58 into accuracy in both
prior runs. All three bear directly on this design and are named again in §7.
4. The statistics, fixed before any score exists
All quantities are per (arm, segment), averaged over the three seats.
- Δ_h, the tax: mean over the eight segments of
accuracy(h-R06) −accuracy(h-R08). - E_h, the execution: mean over the eight segments of
perceived-source-carriage(h-R08) −perceived-source-carriage(h-R06).
The executed bar is E_h ≥ +0.50, and it is fixed here, before dispatch. Its warrant is prior
data, not taste: across S134 and S141 this instrument's null separations on this scale sat at
0.04 to 0.15 (style-correspondence between the programme arms, twice) while its genuine
carriage separations sat at +2.33 to +2.81. A bar of +0.50 is above the observed null band and
far below every observed execution. It is not tuned to any number in this run, which does not exist.
| id | statement | test | bar |
|---|---|---|---|
P1 |
a written programme is executable by hands that did not write it | exact sign-flip over the seven E_h (2⁷ = 128), two-sided; and the count of hands with E_h ≥ +0.50 |
P ≤ 0.05 and ≥ 4 of 7 executing |
P2 |
following the programme costs accuracy, lead-free |
exact sign-flip over the seven Δ_h, two-sided |
P ≤ 0.05 |
P2e |
…among the hands that executed it | same test over the executing subset | reported; withheld if fewer than 4 execute (F1) |
P3 |
the cost is proportional to the execution | Spearman ρ(E_h, Δ_h), exact permutation over 7! orderings |
reported with its P; declared underpowered in advance |
P4 |
the lead is not an ordinary hand | rank of Δ_lead among the eight Δ values, exact one-sided P = rank/8 | reported; min attainable P = 0.125, so it can never reach 0.05 |
P2e is a selection on a measured variable and the design says so before running. Conditioning
on E_h after measuring it biases the subset's Δ upward if the two are correlated by noise;
declaring the rule in advance does not remove that bias, it only removes the freedom to choose
the rule afterwards. P2e is therefore evidence about what a tax looks like where a programme was
executed and is not an unbiased estimate of anything. P2 over all seven is the primary.
Failure criteria.
F1— fewer than 4 hands execute ⇒P2ewithheld, andP1is reported as a failure of executability, which is itself a result and is the one this design most expects to have to report honestly.F2— a seat × sense using fewer than 3 distinct integers over its 144 scores ⇒ that sense is withheld for that seat and the fact is reported before any mean.F3—C1fails ⇒ noaccuracyprimary is read at all, because a jury that cannot see eight planted content errors cannot be measuring a 0.3-point content difference.F4— a hand's arms malformed after one fall-through ⇒ hand dropped,nreduced, reported before any P.F5— fewer than 95% of the 1,728 expected (seat × item × sense) cells returned ⇒ the shortfall is reported per arm before any primary, and any arm missing more than one segment on any seat is dropped.
Gates.
| id | what it checks | bar |
|---|---|---|
C1 |
the jury sees planted content damage | Δaccuracy(LEAD-R06 − WRONG) ≥ +1.00, sign-flip P ≤ 0.05 over 8 segments |
C2 |
the jury sees mangled English | Δnaturalness(LEAD-R06 − CLUNKY) ≥ +1.00, sign-flip P ≤ 0.05 |
C3 |
cross-sense leak, measured not gated | Δnaturalness(R06 − WRONG) and Δaccuracy(R06 − CLUNKY) reported; S134 measured the second at 0.551 and S141 at 0.583 |
M1 |
recall-proofness | §2; fired on H3; registered response above |
There is no cross-run drift gate, deliberately: no seat has judged this span before, so every
comparison in this run is inside this run. The lead's Δ on S141's neighbouring span (+0.810 for
R06 − R08) is quoted in the result as a descriptive neighbour and is not a gate and not a
replication.
5. Predictions, written before any generation call
- ≥ 5 of 7 hands execute (
E_h ≥ +0.50). - The mean
E_his smaller than the lead'sE_lead— a hand executing a rule set it has read once moves less than the translator who assembled it. P2misses (P > 0.05 over seven hands): the lead-free tax is small, and seven is not many.P3is positive in sign (ρ > 0) and does not reach significance.C1andC2both pass,C1by more than +2.00.style-correspondenceseparates nothing — |Δ| < 0.30 between each hand's two arms, its third consecutive null on this instrument.
6. Procedure
snapshot open— done, 85.017990275.probe— done, §2, $0.033340.- Lead arms translated and frozen at
b3df80a— done, $0. - This design frozen; independent pre-run critic (
openai/gpt-5.6-terra, reasoning off, 16,000 cap) over the design, the rubric, the rule set, the operators and both frozen logs. Findings accepted or overruled in writing, before any generation call. translate— 14 calls, 7 hands × 2 arms, 4,000 cap, temperature 0.3.build_items.py— 144 items, opaque ids, per-seat shuffles.score— 3 seats × 8 blocks, 10,000 cap.analyse.py, thenverify.pyrecomputing every reported number by an independent path plus mutation tests.snapshot close; reconcile the key-usage delta against the per-responseusage.costsum.
Judgment is not parallelised. Generation calls may overlap; scoring calls are issued in sequence.
7. Threats, and which of them this design does not answer
perceived-source-carriageis not a calibrated instrument.P1rests on it entirely. Its own entry records a 0.40–0.60 false-positive rate on manufactured oddity, which means a hand that merely wrote odd English could clear the executed bar without carrying anything of the source. The design's partial answer isC2and theCLUNKYarm, which is manufactured oddity with no source-orientation at all:CLUNKY's own carriage score is reported beside the bar, and ifCLUNKYclears +0.50 overLEAD-R06the bar is disclosed as uninterpretable andP1is withheld. This is registered here asC4and is the sharpest check in the run.- Seven language models are not seven translators. Nothing here licenses a sentence about human translators. The population is hands that can be handed a rule set, and the finding transfers to human practice only as a hypothesis.
- The hands share a training distribution with the seats. S141's limit 5, unrepaired. The seven labs make this weaker than a single hand did, and do not remove it.
- The lead chose the source, wrote both lead arms knowing the hypothesis, and wrote the
operators. The repair is structural rather than rhetorical: the lead's arms enter no
primary.
P1,P2,P2eandP3are computed over the seven independent hands only.P4is the sole statistic the lead enters and it cannot reach significance by construction. accuracymay be partly definitional here.R08's R5 forbids supplying what the source withholds and R6 requires calquing figures; a rubric that reads unlicensed addition and misreading could in principle score a calque as a distortion, in which case the "tax" is the rubric disliking the programme rather than the programme losing content. The design does not resolve this. What it can do is reportstyle-correspondencebesideaccuracy: if theR08arms gain on form while losing on content, the trade is real; if they lose on both, the rubric explanation gains. This is named, not tested.- Eight segments, three seats, one passage, one programme.
R07is absent by choice — on this source S141 measured the fluency programme's cost at +0.125, ns, and the resistancy programme's at +0.583 — so this run puts its whole budget on the arm that carried the effect. Nothing here is evidence about domesticating programmes. - Tier D is NOT PASSED. Every sentence produced by this run is
provisionaland no jury verdict carries evidential weight (charter §2.4).
8. Budget
Worst case built from max_tokens and the config/models.md prices, not from expected output
length (note (abc)):
| stage | calls | cap | worst case |
|---|---|---|---|
| probe | 7 | 4,000 | $0.033340 actual |
| critic | 1 | 16,000 | $0.116 |
| translate | 14 | 4,000 | $0.362 |
| score | 24 | 10,000 | $1.290 |
| declared ceiling | $1.85 |
UTC day 2026-08-09 stands at $2.371826 of $5.00 before this run; $2.628174 headroom. The ceiling
fits with $0.778 to spare. The S022 routing caution applies — a slug can bill several times its
list price through provider routing — and the response provider is recorded on every body.
9. Verification
verify.py recomputes, by a path that imports nothing from analyse.py:
- every item's German and English against the frozen artifacts, byte for byte;
WRONGandCLUNKYagainstLEAD-R06as exactly the declared substitutions and no other character;- that the rubric strings are the imported ones, by identity against
E-20260808c's module; - that the hands' prompts are the imported ones and differ from each other only in the rule block;
- body counts with
runs/discarded/excluded; - a full deterministic re-run reproducing
results.jsonfield for field; - every reported mean recomputed outside the reporting path;
- the exact permutation P values by exhaustive enumeration;
- mutation tests, each of which must change a reported number: flip one
accuracyscore structurally; relabel one item's arm; relabel one item's segment; swap two hands' labels; drop one seat's block.
Per §9a of RS-20260809b: a mutation that cannot move the output is not a test, and each mutation
above is asserted to move a specific named figure.
Amendments — the pre-run critic, applied before any generation call
openai/gpt-5.6-terra, reasoning off, 16,000 cap, $0.06087125, stop, 37,996 characters.
VERDICT: NEEDS-REDESIGN — 8 BLOCKING, 16 SERIOUS, 4 MINOR. runs/critic.txt is the full text.
24 findings accepted, 4 overruled in writing. Nothing below was written after a score existed;
the generation stage was dispatched after this section, and the scoring stage after that.
A1 — the execution measure is rebuilt (findings 1, 2, 15, BLOCKING). The critic is right and the
error was a real one: the +0.50 bar's warrant was drawn from style-correspondence nulls, which is a
different construct from perceived-source-carriage, so nothing in it bounded this sense's null
behaviour. Three changes.
- The bar is now a within-run negative control, on the same sense and the same jury. A hand is
recorded as executing iff
E_h > E_CLUNKYandE_h ≥ +0.50, whereE_CLUNKY= carriage(CLUNKY) − carriage(LEAD-R06) — manufactured oddity with no source-orientation whatever, put through the identical instrument in the identical run. This is the discriminator the critic asked for at minimum. C4enters the gates table with a formula and a consequence.C4=E_CLUNKY, aggregated exactly asE_his (per segment, three seats averaged, mean over eight). IfE_CLUNKY ≥ median(E_h),P1is withheld entirely — the sense cannot then tell executing a programme from writing badly. IfE_CLUNKY ≥ +0.50, the +0.50 floor is disclosed as uninterpretable and only clause 1 governs.- A mechanical corroboration,
X_h, computed from the texts with no jury and at $0. R1 and R2 say to reproduce the source's discontinuity and to follow its period length. So per hand,X_h= mean over segments of |sentences(R06) − sentences(German)| − |sentences(R08) − sentences(German)|. A hand executing R1/R2 tracks the German's sentence count more closely underR08than under its ownR06, soX_h> 0. It is source-grounded, unblinded to nothing, and cannot be produced by generic oddity.
The +0.50 floor is declared provisional and unvalidated (finding 2 accepted in full): being
fixed before dispatch prevents threshold shopping and does not make a threshold calibrated, and no
distribution of E_h under non-execution has been estimated for this sense.
A2 — P2e is descriptive only (finding 3, BLOCKING). Accepted in full. It is reported as a
subgroup mean with no P value, no interval and no causal language. Selecting on a measured E_h
distorts the subgroup's Δ if the two share error structure, which the shared jury and shared items
make likely; declaring the rule in advance removes discretion and not bias.
A3 — P3 is exploratory (finding 5, BLOCKING). Accepted. ρ is reported with no P value,
alongside the critical |ρ| ≈ 0.79 it would have needed, and no directional conclusion is drawn from
it in any licensed sentence.
A4 — the population is seven named systems (findings 6, 26, 27, BLOCKING/SERIOUS). Accepted in full, and it changes the unit's own sentence. Every hypothesis, prediction and licensed conclusion now reads these seven named language-model systems, under these two prompts, on this passage, under this one programme. No sentence about human translators, about programmes in general, or about the causal effect of execution may be drawn from this run. §1's motivation stands as motivation; what the run can support is narrower than what motivated it, and the result page states the narrower thing.
A5 — the endpoint is renamed and its circularity conceded (findings 7, 8, BLOCKING/SERIOUS).
Accepted as unrepaired. Δacc is the accuracy-score difference, never "the cost" or "the tax"
without the caveat attached: R5 forbids supplying what the source withholds and R6 requires calquing,
and a rubric that reads unlicensed addition, omission, distortion may score faithful calques and
retained ellipsis as content failure. style-correspondence beside it does not decide between
the two readings. Repairing this needs proposition-level bilingual adjudication, which this run does
not have and does not pretend to.
A6 — C1/C2 are necessary and not sufficient (findings 9, 14, BLOCKING/SERIOUS). Accepted.
F3 keeps its blocking direction — a jury that misses eight planted reversals cannot measure
anything — but a pass licenses only "the jury detects gross planted content damage" and says
nothing about resolution at 0.3 points, which is the size of the difference of interest. C1 is
reported item by item as well as pooled (finding 12), because a pooled mean over heterogeneous
manipulations is not a sensitivity certificate.
A7 — the control operators are rewritten (findings 12, 13, BLOCKING). Accepted, and repaired
rather than caveated. Every CLUNKY substitution is now a strict word-multiset permutation of the
string it replaces, case-folded — no word added, removed or duplicated — and verify.py asserts
it mechanically, so changes no content is proved at the lexical level instead of asserted. The
critic's specific casualties (S2's stray it, S5's duplicated it, S6's added her, S3's and S8's
resumptive pronouns) are gone. WRONG S1 and S8 were replaced: held its post and cursed were
judged partly detectable from English alone, and are now ordered the carriage (German: «die
Pferde») and the celebration of the wedding (German: «des Einzugs»), each invisible without the
German.
A8 — H3 is excluded from every confirmatory analysis (finding 16, SERIOUS). Accepted in the
strong form, against this design's own registered response. The void condition fired; a rule that
keeps the hand when the results happen to agree is a conditional analysis rule. P1, P2, P2e
and P3 are computed on the six hands H1, H2, H4, H5, H6, H7. H3 appears only in a quarantined
sensitivity analysis. The cost is power: at n = 6 the minimum attainable two-sided sign-flip P is
0.03125, so P2 can only reach significance if all six point the same way. That is the price of
enforcing the condition as written and it is paid.
A9 — "recall-proof" becomes "limited public-retrieval screening" (finding 17, SERIOUS). Accepted. Unreachability from this container is not absence from a training corpus, and a self-report probe tests willingness to confabulate rather than memorisation — H3 is the demonstration. Possible memorisation is unresolved, and no sentence in the result may say otherwise.
A10 — arm dispatch order counterbalanced (finding 18, SERIOUS). Accepted and implemented: odd
hands R06 first, even hands R08 first. Every call is a fresh stateless HTTP request with no
system message and no conversation history, so no carryover mechanism exists; the counterbalance
costs nothing and removes the objection rather than arguing with it. Provider, finish reason and cost
are recorded on every body.
A11 — no substitute model (finding 19, SERIOUS). Accepted; the declared reserve is withdrawn.
A hand that cannot return well-formed segments is a registered outcome relevant to executability,
not an operational hiccup to be papered over with a different vendor under the same label. Such a
hand is dropped, n is reduced, and the fact is reported before any P value.
A12 — per-seat and leave-one-seat-out reporting (finding 20, SERIOUS). Accepted at $0: E_h,
Δ_h, C1 and C2 are reported per seat and recomputed with each seat removed in turn, and the
per-segment values are published. Averaging three seats does not make them independent, and a mean
can be stable because all three share a bias.
A13 — F2 is demoted (finding 21, MINOR). Accepted. Scale usage is reported as a response-style
statistic and withholds nothing; a seat using all seven integers can still be non-discriminating.
The leave-one-seat-out sensitivity of A12 replaces it as the check that does work.
A14 — the verifier's scope is named (finding 22, SERIOUS). Accepted. verify.py verifies code
and materials integrity, not scientific validity — it cannot certify the rubric, the blinding, or
the population claim, and §9's title now says so. Each mutation names the specific figure it must
move, and a negative mutation is added: permuting the ids inside a seat's block order must change
no arm mean, and the verifier asserts that it does not.
A15 — a hard abort rule (finding 23, SERIOUS). Accepted in substance, with one factual
correction: the worry named is a breach of the $5.00 UTC-day cap, and $2.371826 already spent plus
this run's full declared worst case of $1.85 is $4.22, so the day cap is not reachable even
if every call bills its ceiling. The routing multiplier is the real exposure, so: before the score
stage is dispatched, cumulative spend is checked against $0.60; if it is above, the score stage is
issued seat by seat and the run stops at whatever is affordable, reporting the shortfall as F5
rather than continuing.
A16 — condition recognition is conceded, not fixed (findings 10, 11, SERIOUS) — PARTLY
OVERRULED. Accepted: P1 measures perceived execution and the run cannot exclude that seats
recognise the foreignizing arm. That is what the sense's name says and it is now said in the
licensing section too. Overruled: the prescription that each judge see only one rendering per
German segment. Reasons, in order of weight. (i) Arithmetic: 144 items over 8 segments means every
block of 18 contains at least three items per segment; one-rendering-per-judge needs 18 blocks of 8
and 54 scoring calls, roughly $1.4 more than this run's whole score budget. (ii) Validity: the
exact sign-flip tests are paired within segment; a between-judge allocation confounds arm with
judge, which is a worse error than the one it repairs. (iii) Comparability: this block structure
and this instruction string are E-20260808c's and E-20260809b's, and moving them voids reading
this run's figures beside theirs. The per-seat independent shuffles mean any order-of-exposure
learning is spread across arms rather than confounded with them, which is stated as mitigation and
not as a fix.
A17 — predictions are separated from the report (finding 25, MINOR). Accepted: analyse.py
emits every pre-specified outcome into results.json mechanically, and the result page prints that
table before any interpretation.
A18 — R08's own under-operationalisation (finding 28, MINOR). Recorded as a limit and not
repaired: R08 is a frozen regime, and rewriting it to be decision-tree operational would break
every prior run that used it. The critic's point stands — two hands can both execute faithfully and
diverge — and it is one more reason P1 is a measure of perceived uptake rather than of compliance.
Overruled, with reasons: findings 10/11 in part (A16 above), finding 24, and the "delete P3" half
of finding 5.
- Finding 24 asks that the budget be reallocated from the seven-hand run to a pilot validating
the instruments. Overruled on the charter's own rule: a unit whose question is are this project's
instruments valid is method work, which
wiki/tracks.md's subject rule bars from being a session's principal unit. The endpoint limits are recorded in §7 and in A1/A5 instead, which is what that rule prescribes — a gate inside the unit, not a session spent on the apparatus. - The "delete
P3" half of finding 5 is overruled in favour of the critic's own fallback in the same paragraph: report it as exploratory with no directional conclusion (A3).
A19 — the generation stage was re-dispatched once, and the ceiling raised to $2.10 to pay for
it. Written after the generation stage and before any scoring call; no score existed. Five of
fourteen bodies returned finish_reason: length with ZERO visible content — moonshotai/kimi-k3
on both arms, and qwen/qwen3.7-max, z-ai/glm-5.2 and minimax/minimax-m3 on R08 — the whole
4,000-token cap spent on hidden reasoning, which is notes (bfb), (bga) and (bkw)'s shape and not a
property of the hands. The identical prompt was re-dispatched to the identical slug at a raised
cap of 8,000 with reasoning suppressed where the provider accepts it; the dead bodies are preserved
under runs/discarded/ and are counted as waste (note (bhd)). This is not amendment A11's
prohibited move — no substitute model is inserted under a hand label, and the prompt string is
byte-identical — but it is a repair made after seeing an outcome, so it is declared here rather than
absorbed. A hand that fails again is dropped, n reduced, and the fact reported before any P value.
The declared ceiling moves $1.85 → $2.10; the UTC day would then stand at $4.47 of $5.00,
so the cap is not reachable.