Repository path: workshop/experiments/E-20260809b-programme-tax/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260809b-programme-tax |
| status | frozen |
| created | 2026-08-09 |
| updated | 2026-08-09 |
| track | T3 |
| senses | accuracy, naturalness, style-correspondence, perceived-source-carriage |
| provisional | true |
| internal-judgment-only | true |
| links | wiki/arms/ARM-programme-tax.md, wiki/findings/results/RS-20260808c-sense-tradeoff-de.md, wiki/findings/results/RS-20260807e-sense-tradeoff.md, workshop/experiments/E-20260808c-sense-tradeoff-de/design.md, workshop/translations/kronenwaechter-sturm/R06-v1/translation.md, workshop/translations/kronenwaechter-sturm/R07-v1/translation.md, workshop/translations/kronenwaechter-sturm/R08-v1/translation.md, workshop/regimes/R06-lead-single-pass.md, workshop/regimes/R07-fluency.md, workshop/regimes/R08-resistancy.md, config/models.md |
E-20260809b — does a programme cost accuracy, or is unruled translation remembering?
Frozen before any scoring call. The three lead arms on the new source were frozen first, at
3c9ae09, before this file existed.
1. The question
RS-20260807e (Lu Xun, ZH→EN) and RS-20260808c (Kleist, DE→EN) both found the unruled arm
R06 scoring about +1.1 points of accuracy above either declared programme — the same size,
from opposite programmes, in two language pairs. Two readings predict that table:
- TAX — following any declared programme costs propositional accuracy.
- RECALL — the unruled arm drifts toward remembered published English, which is accurate, so it scores as accurate without translating better.
On the German pair the recall pathway was demonstrably open: RS-20260808c §2 measured lead
R06 at 110 shared 7-grams, 31 12-grams and a 21-token run against two published English hands,
five times what those two hands share with each other. The registered control C4 did not fire;
the continuous form returned ρ = +0.587, P = 0.173 on seven segments. RS-20260808c limit 4
records the pathway as unexcluded, and this experiment exists to exclude it or to concede it.
The design is an interaction, not a replication. A source whose published English does not exist is run against the same Kleist arms, re-judged by the same seats in the same blind run, so that the comparison does not have to trust a figure recollected across sessions.
2. Materials
Source A — recall-rich. Kleist «Michael Kohlhaas» (1810), the Lisbeth span, 7 segments, 728
German words. Unchanged from E-20260808c; its arms are the frozen S134 artifacts.
Source B — recall-proof. Ludwig Achim von Arnim, «Die Kronenwächter» I.7 «Der Sturm» (1817), 8 segments, 694 German words, frozen this session.
Why B is matched to A and not merely different. Same language pair; seven years apart
(1810 / 1817); the same Berlin Romantic circle — Arnim and Kleist were contemporaries and
associates; the same kind of prose (long-periodic narrative with embedded speech); the same
translator, the same regimes, the same arm order R06 → R08 → R07, the same jury, the same
session, and a rubric string that is byte-identical. What differs, by selection, is the one thing.
Why B is recall-proof. No English translation of «Die Kronenwächter» could be located in any
edition, anthology, catalogue or digitised text reachable from this container. The work is the
standing example of the untranslated Arnim; the Arnim titles that are in English — Isabella of
Egypt, The Mad Veteran of Fort Ratonneau, Melück Maria Blainville, The Marriage Blacksmith —
do not include it. This is bibliographic, not measured, and it is falsifiable: one located
English rendering of this chapter voids the manipulation. tools/dependence_check.py therefore
cannot be run on B's arms — there is no comparator — and that is declared on each artifact rather
than left as a silent contamination: none, per CLAUDE.md §Contamination. M1 below measures
the same thing a second way instead of assuming it.
Copy-text. B's German collated in full against an independent digitisation (projekt-gutenberg.org) at sequence ratios 1.0000 / 0.9995 / 1.0000; the two variants are punctuation and no word differs.
3. Arms
| source A (Kleist, 7 seg) | source B (Arnim, 8 seg) | |
|---|---|---|
| lead, no rule set | R06 (S134) |
R06 (S141) |
| lead, fluency | R07 (S134) |
R07 (S141) |
| lead, resistancy | R08 (S134) |
R08 (S141) |
| unbriefed hand, no rule set | IND-R06 new |
IND-R06 new |
| unbriefed hand, fluency | IND-R07 (S134) |
IND-R07 new |
| unbriefed hand, resistancy | IND-R08 (S134) |
IND-R08 new |
R06 + 1 content error/segment |
— | WRONG |
R06 + 1 syntactic mangling/segment |
— | CLUNKY |
42 + 64 = 106 items. WRONG/CLUNKY are carried on B only: they exist to gate this jury, and
one gated half is what the gates are for. The independent ladder gains an unruled arm, which
S134 did not have — without it there is no lead-free measurement of the tax at all, only of the
R08−R07 difference.
The unbriefed hand (mistralai/mistral-medium-3-5, non-panel) receives the German and, for the two
programme arms, the rule set verbatim and nothing else — no hypotheses, no other arms, no knowledge
that an experiment exists. The IND-R06 prompt is the identical preamble with the rule paragraph
and rule block deleted, so that the three independent arms differ in exactly what the three lead
arms differ in.
4. Procedure
Three non-Anthropic seats — J1 = P1 openai/gpt-5.6-terra, J2 = P2 google/gemini-3.6-flash,
J3 = P5 deepseek/deepseek-v4-pro — blind to arm and to source, item ids opaque, each seat's own
deterministic shuffle, six balanced blocks. Four senses, German present. A separate source-blind
naturalness pass on the English alone is dispatched first (S134 amendment A10), so that every
naturalness figure was taken before any seat saw a German word.
The scoring prompt string, the sense definitions and the item format are byte-identical to
E-20260808c. This is load-bearing: gate G1 compares this run's Kleist figures against S134's
published ones, and that comparison is meaningless if the string moved. For the same reason the
request body for the scoring and blind stages is left exactly as S134 sent it — note (bkw)'s
reasoning: {enabled: false} is applied to the critic and the retrieval probe only, where there
is nothing to be comparable to, and deliberately not to the seats.
Pre-run adversarial critic (nvidia/nemotron-3-ultra-550b-a55b, non-panel) over this frozen design
before any scoring call. Every raw body written to runs/ before anything is computed from it;
dead bodies to runs/discarded/, never overwritten (note (bhd)). Every dispatch in the background
(note (bid)).
5. Predictions, each with the outcome that falsifies it
Let Δ = accuracy(R06) − mean(accuracy(R07), accuracy(R08)), computed per segment as
the mean over the three seats, so Δ_A has 7 values and Δ_B has 8. All intervals are exact
permutation intervals; two-sample tests permute the 15 segment labels across the two sources.
P1 — the tax survives where recall cannot operate. Δ_B > 0, 90% CI excluding 0.
Fails if the CI includes 0.
P2 — recall does not explain the tax. The 90% CI on (Δ_A − Δ_B) lies inside ±0.75 — the
equivalence margin S134's critic imposed and S134's own lead pair failed by 0.027.
Fails if the interval leaves ±0.75 in either direction.
P3 — the lead-free ladder does the same. P1 and P2 recomputed on IND-R06/IND-R07/
IND-R08. The lead knew the hypothesis while translating B; the unbriefed hand did not, and at
S134 the independent ladder was the only arm that established anything.
How the four outcomes read, registered now so that none of them can be chosen afterwards:
P2 holds (no interaction) |
P2 fails (Δ_A > Δ_B) |
|
|---|---|---|
P1 holds |
TAX. The programme cost is real and is not recall. | BOTH. Real, and inflated on A by recall. |
P1 fails |
NULL, uninformative — neither source shows a tax; G1 must be checked before anything is said. |
RECALL. RS-20260807e §5 and RS-20260808c §5 are retracted and wiki/tracks.md's T3 line is corrected. |
6. Gates. Any failure withholds the primary and the withholding is reported
G1— drift. Δaccuracy(R06−R07) on source A must reproduce S134's +1.0952 within ±0.75, on byte-identical texts under a byte-identical prompt. This is method work and it is placed where method work belongs — as a gate inside the unit it blocks, not as a finding. If it fails, this jury is not the S134 jury and the interaction has no baseline;P2is withheld and only the within-runP1stands.C1— the jury sees content damage. Δaccuracy(R06−WRONG) on B ≥ +1.0, one-sided permutation P < 0.05.C2— the jury sees English damage. Δ blindnaturalness(R06−CLUNKY) on B ≥ +1.0, P < 0.05.C3— cross-sense specificity.WRONGleak ratio (Δnaturalness/ Δaccuracy) < 0.50. S134 recordedCLUNKYfailing this at 0.551; that failure is expected to recur and is a standing limit, not a new finding.F1— scale usage. Each seat uses ≥ 4 distinct integers on each sense.
7. M1 — the manipulation, measured rather than assumed
The three seats and the unbriefed hand are each shown the 15 German segments, interleaved and
unlabelled, and asked: do you know a published English translation of this passage, and if so quote
its opening clause; otherwise answer NONE.
Registered prediction: the hit rate on A exceeds the hit rate on B, and the hit rate on B is 0.
M1 does not gate the primary, and the reason is worth stating in advance: a null on A would
show only that retrieval on demand is not the pathway, not that overlap in generation is absent
— and overlap in generation was measured directly at S134 (110 shared 7-grams). M1 can support
the manipulation; it cannot by itself defeat it.
8. Threats, stated before the numbers exist
- A and B are different texts. Every match in §2 narrows this and none removes it. A finding of no interaction is therefore stronger than a finding of interaction: an interaction is what a text difference would also produce.
- The lead knew the hypothesis while translating B — worse than S134's "knew the result",
because here the lead knows which arm the hypothesis is about.
P3is the answer to this and is the reason the independent ladder gained an unruled arm. - Recall-proofness is bibliographic. §2 and
M1. - 8 segments against 7, and the two-sample test is on 15 values. Small.
- Both halves share the arm order
R06→R08→R07. Deliberate — an order that differed between halves would confound the interaction — but it means any order effect is common to both. - Tier D is NOT PASSED. No verdict here carries evidential weight (charter §2.4).
9. Budget
Pre-flight worst case built from max_tokens, not from expected output (note (abc)):
| stage | bodies | worst case |
|---|---|---|
| score, 3 seats × 6 blocks, 10,000 cap | 18 | $0.99 |
| blind, 3 seats × 2 blocks, 12,000 cap | 6 | $0.40 |
| critic, 16,000 cap | 1 | $0.06 |
| independent hands, 4,000 cap | 4 | $0.03 |
| retrieval probe, 3,000 cap | 4 | $0.03 |
| total | 33 | $1.51 |
Declared worst case $1.60 against a UTC-day headroom of $4.313164 ($0.686836 already spent
by S140). Expected actual, scaling S134's $0.348 by item load, ≈ $0.55. Key-usage snapshot open
and close; per-response usage.cost primary, snapshot delta as the cross-check.
10. Verification
verify.py recomputes every number the result page reports from the raw bodies, checks each item's
English against the frozen artifact byte for byte, checks that WRONG/CLUNKY differ from R06
by exactly the declared substitutions and in no other character, and runs 5 seeded mutations that
must each be caught.
11. Amendments — adopted from the pre-run critic, before any scoring call
nvidia/nemotron-3-ultra-550b-a55b, runs/critic.txt, VERDICT: NEEDS-REDESIGN, 6 BLOCKING,
2 ADVISORY. Seven findings accepted, two overruled with reasons. No score had been dispatched
when these were written; the four independent-hand arms and this critic were the only bodies.
A1 (critic (b), BLOCKING — accepted, and it changes the primary). The critic's arithmetic is
right: a 90% two-sample permutation interval on 7 against 8 segment-level values has a width of
order 2.5–3.0 scale points, so an equivalence claim at ±0.75 cannot pass, and S134's own
paired n = 7 missed the same margin by 0.027. P2 as written made the headline outcome
unreachable. P2 is therefore no longer an equivalence test and is no longer a primary. It
becomes P2′: a two-sided exact permutation test for a difference in the tax between sources,
reported with its interval, secondary and descriptive, in a segment-level form (7 vs 8) and a
(seat × segment) cell-level form (21 vs 24) reported alongside as S129/S134 reported theirs.
The primary is now P1 alone, and the inference does not need the interaction. If the tax
appears on a source where the recall pathway is closed, recall is not what produces it. That
argument is within-run, is paired over 8 segments (sign-flip minimum P = 0.0078), and does not
borrow anything from source A.
A2 (critic (a), BLOCKING — accepted). Recall-proof is withdrawn as a description of source
B and replaced by bibliographically untranslated. Absence from every catalogue reachable in
2026 is not absence from a training corpus, and M1 probes retrieval-on-demand, which is not the
only pathway — S134's 110 shared 7-grams were produced without any model claiming to know it was
reproducing anything. Registered void condition: if any M1 respondent returns a plausible
English quotation for any source-B segment, the manipulation is void and P1 is reported as
uninterpretable. What survives the amendment is a difference in probability of recall between
two texts, one of which was measured with an open pathway and one of which has no located English
at all — and the result page will say that and no more.
A3 (critic (c), BLOCKING — accepted). §8 threat 1's claim that a null interaction is
stronger than an interaction is withdrawn. Registered instead: any interaction found here is
uninterpretable as recall versus text difficulty, because two texts by two authors differ in many
ways that could modulate a programme's cost. P2′ is reported as a description of two numbers, and
no causal reading of it will be offered.
A4 (critic (d), BLOCKING — accepted in part, overruled in part). Accepted: G1 also reports
the Spearman correlation between this run's and S134's seven segment-level Δ values and the
per-seat Δ, and G1 is declared a drift check only — re-judging the same texts with the same
model slugs is a re-test, not an independent replication, and it is not claimed to be one.
Overruled: the critic asks that a G1 failure abort the run. It does not, because P1 is a
within-run paired comparison on source B that borrows nothing from S134; a G1 failure withholds
P2′ and G1's own comparison and leaves P1 standing. Aborting would discard the one measurement
that does not depend on the gate.
A5 (critic (g), BLOCKING — accepted). The critic is right that WRONG S6
(constellations → townspeople) changed a figure rather than a proposition, and would have
been scored against style-correspondence rather than accuracy. Replaced, before any scoring
call, with a propositional reversal: who no longer believed themselves safe → who still
believed themselves safe. The other seven substitutions are accepted as content errors; the
critic's own table agrees on all seven.
A6 (critic (g), second half — overruled with reason). The critic asks that C3 be moved to
test CLUNKY's leak into accuracy. C3 exists to show that the content probe is
content-specific, which is what licenses reading C1 as a content gate, and moving it would delete
that check. C3 stands on WRONG at a bar of 0.50. CLUNKY's leak into accuracy is
computed and reported, is expected to fail (S134 measured 0.551), and is carried as a standing limit
on every accuracy figure in this project rather than as a gate — which is where RS-20260808c §7
already put it.
A7 (critic (e), BLOCKING — accepted as far as it can be). A rule-compliance audit is added
(S134's A4): an independent non-panel agent, given only the two frozen rule sets, the German, and
the two programme renderings of source B, judges each of the translator's logged divergence sites
REQUIRED / LICENSED / NOT-SUPPORTED. This bounds how much of the R07/R08 contrast the lead's
hand could have chosen. It does not repair the fact that the lead knew the hypothesis and which
arm it concerns while translating, and no claim here will say it does. The critic's further point
— that the unbriefed hand shares a training distribution with the seats, so P3 is a correlated
rather than an independent replication — is accepted in full and recorded as a limit; P3 is
reported as lead-free, which is what it is, and never as independent.
A8 (critic (f), ADVISORY — accepted). "Blind to arm and to source" is restated: the seats are
blind to labels and to the structure of the experiment. The German reveals which work an item
comes from, and the English of a programme arm is recognisable as such. Blinding here is on
identity and grouping, not on inference, and §4's wording is corrected to say so.
A9 (critic (h) — overruled with reason). The critic asserts that OpenRouter does not return
cost in usage and that cost must be computed from token counts and list prices. This is false for
this endpoint: with "usage": {"include": true} in the body, usage.cost is returned per response,
and this project has reconciled per-response sums against key-usage deltas to within 3 × 10⁻⁹ across
many sessions, including S140 yesterday. The stated method stands. Accepted from the same
finding: the five verification mutations are named explicitly in verify.py rather than promised.