Repository path: workshop/experiments/E-20260801e-lead-carryover/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260801e-lead-carryover |
| status | frozen |
| created | 2026-08-01 |
| updated | 2026-08-01 |
| senses | naturalness, style-correspondence |
| internal-judgment-only | true |
| links | wiki/arms/ARM-carryover.md, workshop/experiments/E-20260731h-carryover/design.md, wiki/findings/results/RS-20260731h-carryover.md, workshop/regimes/R10-anchored-register.md, workshop/regimes/R07-fluency.md, config/models.md, wiki/method-notes.md, tools/dependence_check.py |
E-20260801e-lead-carryover — the archive supplies the independent arm
FROZEN before any translating and before any dispatch. Amendments made after the pre-run critic pass are appended to §12 with the finding that produced them, and nothing above §12 is edited in place.
1. Question
ARM-carryover step 1 (E-20260731h, S076) measured carryover in three panel seats and
wrote, in its own §2, that the lead's could not be measured: "The lead cannot render a unit in
ignorance of its own prior rendering of that unit. No procedure available to it removes the
first arm from memory."
That sentence is true about one rendering and false about the pair. The thing the arm needs is not a lead that has forgotten; it is a first arm the second arm's author never saw. The archive holds five matched pairs written months of sessions ago, and this session's lead has no memory of any of them. So:
Is a second rendering closer to the first arm it was written beside than to an independently written first arm by the same agent under the same brief?
If yes, the within-pair comparisons this project has published on those pairs are confounded in the direction the arm suspected, and the size of the confound is now a number.
2. The construction
For an archived matched pair, write X for the arm written first and Y for the arm written second, both by the lead in one session, in one context.
This session, blind to both, the lead renders the same source unit under X's own brief, producing X′. Then:
| quantity | definition | what varies |
|---|---|---|
| Seq | ov(Y, X) | Y against the first arm it had in context |
| Ind | ov(Y, X′) | Y against a first arm by the same agent under the same brief, written in a different context |
| Carry | Seq − Ind | Y is held byte-identical. Only whether its author had seen the comparator differs |
This is the mirror of E-20260731h's Seq(C)/Ind(C), which held the comparator fixed and varied
the measured text. Here the measured text is fixed and the comparator varies, which is the
stronger form: Y is an artifact frozen in a previous session and nothing this session does can
touch it.
Two further quantities, both free:
| quantity | definition | why |
|---|---|---|
| Self | ov(X, X′) | how much the lead repeats itself across contexts under one brief. Load-bearing: see the guard in §8.3 |
| Floorj | ov(Y, Sj) | the same comparison with an outside agent's independently written first arm, seat j |
| Floor′j | ov(X, Sj) | ditto against X |
The ladder the design is built to price. Floor is source forcing plus generic English. Ind adds the same agent's idiolect across contexts. Seq adds contact. Each rung is an increment and each increment is measured, not argued.
3. Materials — three archived pairs, all written before this question existed
| pair | work, source language | source unit | X (first) | Y (second) | session |
|---|---|---|---|---|---|
| K | Kielland, «Karen» (1882), ¶1–15 — Norwegian | workshop/translations/karen/source-unit.txt, 799 words |
T-karen-R10c-v1 (R10, A-mchugh-presence, C01–C19) |
T-karen-R10p-v1 |
S074 |
| C | d'Annunzio, «La fine di Candia» (1886), ¶1–24 — Italian | workshop/experiments/E-20260731g-provenance/material/source-unit.txt, 800 words |
T-candia-R10p-v1 (R10, A-mansfield-garden-party, M01–M11) |
T-candia-R10c-v1 |
S075 |
| O | Tarchetti, «Un osso di morto» (1869), opening — Italian | workshop/translations/osso-di-morto/source-it.txt, 357 words |
T-osso-di-morto-R07-v1 (R07, F1–F10) |
T-osso-di-morto-R08-v1 |
S048 |
Arm order is not inferred from git — the squash-merge collapses intra-session order. K and C
come from wiki/arms/ARM-carryover.md §"There is a specific two-point observation to explain"
(S074 wrote the centre arm first; S075 wrote the centre arm second); O from
E-20260728e-venuti-specifiability/design.md §Procedure steps 2–3 (R07 then R08).
The three new renderings, each under X's own regime at X's own version, filed as first-class artifacts before any measurement exists:
Excluded, with the reason. The S076 uj-pap pairs (u1, u2) are not in the primary
set. Their translator had constituted ARM-carryover earlier in the same session and knew the
hypothesis while writing both arms — the one condition that would suppress the effect. Mixing a
treated condition into a three-item primary is not a robustness check. The two Kleist/Sōseki
pairs the arm also names are excluded for a different reason: they are draft-and-revision
pairs, where the second text is an edit of the first and dependence is definitional, not
inherited.
4. Blindness protocol, and what enforces it
- This design is committed before the sources are read for translation. The commit is the freeze.
- The lead reads only the source unit, the regime page, and the frozen proposition file. It
does not open
X,Y, their translator's logs, the result pages that quote them, or any published English of these works. material/extract.py— which slices the archived English out of the sixtranslation.mdfiles — is written and first executed only after all three new renderings are committed. Nothing displays an archived rendering to the lead before that commit.- The freeze is a git commit per rendering, before the next is begun.
- What this does not enforce, stated rather than hedged: the lead knows the hypothesis. It cannot deliberately diverge from a text it has not seen, but it could write atypically. §8.3 registers the guard that would catch that, and §11 states what remains.
5. Subjects for the Floor rung
Three non-Anthropic seats from config/models.md, used as contrast subjects — the same
three E-20260731h used, so the two steps' floor and seat cells are commensurable:
| role | slug |
|---|---|
| S1 | openai/gpt-5.6-terra (P1) |
| S2 | x-ai/grok-4.5 (P3) |
| S3 | deepseek/deepseek-v4-pro (P5) |
Each seat receives, in a fresh context, the source unit and X's own brief in the wording the lead worked from, and returns one rendering. Nine logical seats (3 passages × 3 seats). Seats are not told what is measured, that another rendering exists, or that a second brief exists.
No model judges anything in this design. There is no jury, no rating, no quality claim. Every reported number is a mechanical token overlap. Tier D's failure therefore does not bear on any number here.
Role collision (the S053 fix). Critic seat qwen/qwen3.7-max (probed-but-not-selected),
declared reserve google/gemini-3.6-flash (P2, not a subject here) per note (bfc). Neither
is among S1–S3.
6. The statistic
tools/dependence_check.py, used unmodified (ARM-carryover constraint 5). Per pair of
texts it returns longest_run (run), n7, n12, n15.
Only the raw columns are used. The name-excluded columns are not read anywhere in this
design, on RS-20260801d-name-ground-truth: the exclusion rule's precision is 0.679
token-weighted, the error is one-sided over-exclusion, and 11 of 19 measured false positives
come from a punctuation habit that differs between translators. A column biased by whose comma
it is cannot compare two renderings.
Length is a confound and is handled before the data, not after. n7 and n12 are counts,
so they rise with text length, and X, X′ and Sj are three different texts of three
different lengths. The primary statistic is therefore normalised:
n7n= 1000 × n7 / min(tokens(a), tokens(b))
run, raw n7, n12 are reported beside it and are secondary. Registered guard L: within a
pair, if tokens(X′) differs from tokens(X) by more than 25%, the raw comparison of Seq with
Ind is declared length-confounded for that pair and only n7n is read; the flag is reported
whether or not it fires.
7. Controls, registered before the data
- CTRL-POS — the instrument responds to dependence it should find. ov(
T-bettelweib-locarno-R06-v1,T-bettelweib-locarno-R04m-v1) — a matched-length self-revision of that same draft, dependent by construction. Passes ifrun≥ 20 andn12> 0. If it fails, the instrument is not measuring contact on this project's own artifacts and the whole experiment reports NO RESULT. - CTRL-NEG — the statistic is not saturated by agent identity alone. The three cross-pair
cells ov(YK, YC), ov(YK, YO), ov(YC, YO)
— same agent, same session-era, different works, different source languages, no shared
content. Passes if
n12= 0 in all three cells. If it fails,n12is measuring the lead's English and not the pair, andn12is withdrawn from the report. - The Floor rung is itself a control on Ind (§8.2): if an outside model's independent rendering scores as high against Y as the lead's own does, then Ind carries no agent-specific signal and the middle rung of the ladder is empty. That is a reportable outcome, not a failure.
8. Predictions and guards, registered before any translating and any dispatch
8.1 Primary
| # | prediction | fails if |
|---|---|---|
| Q1 | Carry > 0 on n7n at 3 of 3 pairs |
≤ 2 pairs |
| Q2 | Carry > 0 on run at 3 of 3 pairs |
≤ 2 pairs |
Three of three, not a majority: with n = 3 a two-of-three rule is one coin flip from noise, and
E-20260731h's pre-run critic made the same correction to that design's seat rule.
8.2 Secondary
| # | prediction | fails if |
|---|---|---|
| Q3 | Ind > meanj(Floorj) on n7n at 3 of 3 pairs — the lead's independent first arm resembles Y more than an outside agent's does |
≤ 2 pairs |
| Q4 | reported, not predicted: Self, Floor′, raw n7, n12, every token count, and the per-seat spread of Floor |
8.3 Guard T — is X′ a typical lead rendering at all?
The design's own weak point is that X′ is written by an agent that knows what is being measured. It cannot avoid a text it has not read, but it could write unlike itself, and that would inflate Carry.
Guard T fires for a pair when Self ≤ maxj(Floor′j) on
n7n— i.e. X′ is no more like X than an outside model's rendering of the same source under the same brief is. For a pair where guard T fires, Ind is not usable as the lead's own baseline, Carry is reported as UNINTERPRETABLE for that pair, and it counts as a failure against Q1 and Q2, not as an exclusion.
The guard can fire on all three pairs, in which case the experiment's answer is the construction does not work, and that is the result.
9. Void rules and failure criteria, written before the data
- A rendering (lead or seat) is VOID if its English word count is below 0.7× or above 2.0× the source unit's word count, or if it contains a run of 5+ consecutive source-language words, or (seats only) if the delimiters are missing.
- Attempts are capped at 2 per logical seat (the S079 overrun: the estimate assumed one dispatch per seat and the caller can dispatch three attempts across two slugs). A void seat is reported void. There is no re-briefing (note (bda)).
- If fewer than 2 seats survive on a passage, Q3 reports NO RESULT for that passage and the Floor rung for it is reported as the surviving cells only, labelled.
- If a lead rendering is void, the experiment reports NO RESULT for that pair; there is no re-translation, because a second attempt would have the first in context and would be exactly the thing under study.
- No threshold is set on the size of Carry. The project has no prior for what a normalised 7-gram point means here, and inventing one after the data is note (o)'s failure.
- Q1–Q3 are reported separately and a prediction failing is the result, not a reason to re-cut the statistic (the S075 lesson: the discarded statistic may be the one that moves).
- Carry < 0 is a finding, not noise. A second rendering that resembles the first less than an independent one does is deliberate contrast, and it is as interesting as inheritance.
10. Verification
analysis/verify.py recomputes every reported number from the stored bytes — the archived
translation.md files, the new translation.md files, and the seats' .raw bodies — imports
nothing from analyse.py, and must fail when a body is damaged: for every .raw it asserts
finish_reason == "stop" and the exact stored content length, and its mutation tests damage a
real body and require a non-zero exit and are asserted to change the bytes they claim to
change (note (bgu)).
11. What this experiment cannot show
- It cannot separate carryover from anything else that differs between one session's context and another's. Ind is measured across a session boundary; Seq within one. The lead's own model build is not recorded per session anywhere in this repo, so if the Routine configured a different Anthropic model at S048, S074, S075 and S081, Ind is partly a cross-model figure. This is unremovable with what the repo stores and it is not hedged: it is the single largest threat to the reading, and Guard T is only a partial check on it.
- It measures three pairs, two of them Italian, one brief-pair used twice. It is not a rate.
- The lead knew the hypothesis while writing X′ (§4.5, §8.3).
- It says nothing about whether any of these renderings is any good. No evaluative claim is made or elicited anywhere.
- A confirmed Carry does not retract any published figure. It prices a confound in comparisons between the two arms of a pair. Figures comparing one arm to a published third-party translation are a different quantity and are untouched by anything here.
12. Pre-flight budget
Worst case built from max_tokens, not from an assumed output length (note (abc)), with P3
priced at twice list and P5 at four times list for routing (S022 caution), and attempts at the
registered cap of 2.
| line | dispatches | max_tokens |
worst case |
|---|---|---|---|
| pre-run critic (+1 reserve) | 1 | 12,000 | $0.16 |
| seats, K and C | 2 × 3 × 2 | 3,000 | $0.32 |
| seats, O | 3 × 2 | 1,600 | $0.09 |
| total | $0.57 |
Today's headroom at session open: $2.075053177 of the $5.00 UTC-day cap, four sessions
already run (config/budget.md). The run fits with margin.
13. Amendments after the pre-run critic pass
(appended below; nothing above this line is edited in place)
Critic: qwen/qwen3.7-max, provider Alibaba, stop, 172.7 s, $0.049782725. Declared
reserve google/gemini-3.6-flash not used. Verdict NEEDS-AMENDMENT, five findings, one
BLOCKING. All five accepted in substance; two prescribed remedies declined in writing
(A4, A5). Raw at runs/critic.raw, text at critic.txt.
A1 — the normalisation denominator is wrong, and the fix makes the design stricter (finding 1, BLOCKING — accepted, with a stronger remedy than prescribed)
§6 sets n7n = 1000 × n7 / min(tokens(a), tokens(b)). Because Y is byte-identical in both Seq
and Ind, a min denominator moves between the two comparisons whenever X and X′ straddle Y in
length, and Carry then moves with no change in raw overlap at all. The critic is right and the
defect would have run.
Amended. Every family of cells is normalised by the fixed text's 7-gram position count:
- Seq, Ind, Floorj — all compare against Y:
n7n = 1000 × n7 / (tokens(Y) − 6). - Self, Floor′j — all compare against X:
n7n = 1000 × n7 / (tokens(X) − 6).
Within a family every cell now shares one denominator exactly, so all between-condition comparisons — which are the only comparisons any prediction is registered on — are exact.
And the remedy goes further than the critic's, because normalising by Y does not remove the comparator's length. A longer X′ has more 7-grams and therefore more chances to hit Y's set, whatever the denominator. Guard L is therefore restated, not replaced, and given a direction:
Guard L (amended). Report tokens(X), tokens(X′), tokens(Sj) for every pair. If tokens(X′) < tokens(X), the length bias runs in the direction of Q1 and Q2, and that pair's Carry is reported as an upper bound. If tokens(X′) > tokens(X), the bias is conservative and Carry is a lower bound. The ±25% flag is retained and reported whether or not it fires.
A pre-registered statistic that can be biased by the length of the text the lead happens to write is worth less than one whose bias direction is declared before the text exists.
A2 — Carry is contact plus everything else that differs between two sessions (finding 2, SUBSTANTIVE — accepted literally)
§11's first bullet already names the model-build threat. The critic's point is wider and is adopted in its own words: Carry measures contact + shared session context + generation order, and a positive result does not strictly isolate lexical dependence from session-level variance. X and X′ are matched on being the first arm written under their brief, which removes the order component; nothing removes the session component. Every reported sentence uses "Carry", never "carryover", except where the distinction is being discussed.
A3 — Guard T catches extreme atypicality only (finding 3, SUBSTANTIVE — accepted literally)
Stated rather than repaired, because the prescribed alternative (a historical confidence interval for Self) needs a history of cross-context lead self-similarity that does not exist — this experiment produces the project's first three such measurements. Moderate atypicality in X′ will inflate Carry and Guard T will not catch it. Self is reported beside Floor′ for every pair so the margin is visible rather than asserted.
A4 — Q1/Q2 test direction only; Q3 is not a replication (finding 4, ADVISORY — half accepted, half DECLINED with a reason)
Accepted: Q1 and Q2 test directionality, no minimum effect size is set, and §9.5 says why (the project has no prior for what a normalised 7-gram point means, and inventing one after the data is note (o)'s failure). The magnitude of Carry, not its sign, is the finding, and the report leads with the magnitudes.
Declined: the critic states that Q3 "is already known from E-20260731h, which established
this exact ladder". It did not. E-20260731h compared a seat's rendering against the same
seat's other-context rendering; it has no outside-agent rung at all. Q3 asks whether the
lead's independent arm resembles Y more than an outside model's does — a comparison that
experiment could not make. Q3 stays a registered prediction.
A5 — senses: on a page that disclaims evaluation (finding 5, ADVISORY — accepted in substance, remedy DECLINED)
The critic is right that senses: [naturalness, style-correspondence] sits oddly beside §11's
"It says nothing about whether any of these renderings is any good."
The prescribed remedy is declined: lexical-dependence and carryover are not ids in
wiki/goodness-senses.md, and minting sense ids to fix a front-matter smell would breach
CLAUDE.md rule 3, which requires the ids to come from that page. The field is retained because
it records which senses the arm's question lives under, and it is not edited in place mid-freeze.
What the finding actually exposes is a schema gap, and it is filed as one: nothing in the front-matter vocabulary distinguishes a page that judges on these senses from a page whose question is about these senses and which judges nothing. A reader has to reach §11 to tell. New method note (bha).
A6 — a BLINDNESS BREACH on pair K, declared before the rendering is committed (not a critic finding — the lead's own defect, found mid-translation)
What happened. During the archive survey that chose the pairs — before this design was
written — the lead printed the summary block of workshop/translations/karen/gate/dependence.json
in order to establish whether pair K's Seq had ever been computed. That block carries a
max_run_text field. Three strings of archived English were therefore displayed to the lead
before it translated pair K's source, and §4's blindness protocol is breached for that pair and
only that pair. It was noticed while drafting the paragraph containing the third string.
Exactly what was seen — 26 tokens in all, named here so the breach is bounded rather than described:
| string | cell it came from | which archived text it is in |
|---|---|---|
down with a bang and |
karen-R10c-vs-collier1907 |
X (R10c) and Collier 1907 |
to him at first it had been a friendly invitation to |
karen-R10p-vs-collier1907 |
Y (R10p) and Collier 1907 |
swore and the hens shrieked and in the kitchen they were |
karen-R10c-vs-R10p |
both X and Y |
Nothing else of pair K's English was seen, and nothing at all of pairs C and O was: the
candia and osso-di-morto directories hold no computed gate, and no archived rendering of
either was opened or printed.
The direction of the bias, which is the reason this is declared and continued rather than the pair being dropped. Two of the three strings are in Y. If X′ reproduces them, Ind rises, and since Carry = Seq − Ind, Carry falls. The breach can only make Q1 and Q2 harder to satisfy on pair K. The third string is in X only, so reproducing it inflates Self, which makes Guard T easier to pass — that one is not conservative and is flagged as such.
Registered now, before the rendering exists:
analysis/verify.pyreports, for each of the three strings, whether it appears in X′, verbatim after the tool's own tokenisation. This is reported whether or not any appears.- The seat renderings are the test of whether the leak matters. S1–S3 receive pair K's source and brief in fresh contexts and have seen none of this. If a leaked string appears in a seat's rendering, the string is forced by the source and its presence in X′ is not evidence of the breach. Reported per string per seat.
- Carry for pair K is reported as a lower bound for this reason in addition to any length reason under Guard L, and the result page states the breach wherever pair K's number appears.
- Guard T's verdict on pair K is reported with the breach attached, because the breach runs in the permissive direction for that guard alone.
What is not done. The rendering is not adjusted to avoid the strings. Writing a translation so as to miss a string one has seen is fitting the artifact to the hypothesis, which is worse than the breach.