Repository path: workshop/experiments/E-20260731h-carryover/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260731h-carryover |
| status | frozen |
| created | 2026-07-31 |
| updated | 2026-07-31 |
| senses | naturalness, style-correspondence |
| internal-judgment-only | true |
| links | wiki/arms/ARM-carryover.md, workshop/regimes/R10-anchored-register.md, workshop/translations/uj-pap/opportunity.md, workshop/translations/uj-pap/u1-R10p-v1/translation.md, workshop/translations/uj-pap/u1-R10c-v1/translation.md, workshop/translations/uj-pap/u2-R10c-v1/translation.md, workshop/translations/uj-pap/u2-R10p-v1/translation.md, config/models.md, wiki/method-notes.md, wiki/backlog.md |
E-20260731h-carryover — does the second rendering inherit from the first?
FROZEN before any dispatch. Amendments made after the pre-run critic pass are appended to §9 with the finding that produced them, and nothing above §9 is edited in place.
1. Question
A translator writes two renderings of the same source unit under two different register targets, one after the other. Does the second inherit wording from the first, over and above what the source forces?
The question is not academic here. Every matched pair this project has produced — S042 Sōseki, S046 Kleist, S048 Tarchetti, S074 Kielland, S075 d'Annunzio, and the four renderings frozen this session — was written by one continuous agent in one session. If the second arm inherits from the first, every within-pair comparison the project has run is confounded in a direction nobody has measured (wiki/backlog.md, row opened S075).
And there is a specific two-point observation to test. S074 wrote the centre arm first and measured its longest shared run with a published period translation at 5 tokens / 0 shared 7-grams; S075 wrote the centre arm second and measured 11 / 13. The period arm read 11 in both positions. That is the signature carryover would produce — but the two sessions used two different source works, so source and position are confounded, and it is two points.
2. Why this cannot be answered on the lead, and what is done instead
The lead cannot render a unit in ignorance of its own prior rendering of that unit. No procedure available to it removes the first arm from memory, and the lead also knows what is being measured. The four lead renderings frozen this session (8d065a0 → 6a47c8b) therefore cross the arm order within one work for the first time and are reported as a description with n = 1 per cell, in which unit and position are perfectly confounded (U1 is narration, U2 is more than half dialogue). They are not this experiment's test and no prediction below is registered on them.
Panel seats can be given the same brief with and without the first rendering in context, because each conversation is a fresh context. That difference is constructible, and it is the measurement.
The inference from panel seats to the lead is an argument, not a measurement, and is labelled as one wherever it is made. Charter §4: AI-only convergence is weak evidence.
3. Subjects, and what they are not
Three non-Anthropic seats from config/models.md, used as contrast subjects:
| role | slug |
|---|---|
| S1 | openai/gpt-5.6-terra (P1) |
| S2 | x-ai/grok-4.5 (P3) |
| S3 | deepseek/deepseek-v4-pro (P5) |
No model judges anything in this design. There is no jury, no rating, no quality claim. Every reported number is a mechanical token-overlap computed by tools/dependence_check.py, used unmodified. Tier D's failure therefore does not bear on this experiment's numbers, and nothing here gains evidential weight from the panel's agreement.
Role collision (the S053 fix). The pre-run critic must not be a subject. Critic seat: qwen/qwen3.7-max (probed-but-not-selected, config/models.md §reserves), declared reserve google/gemini-3.6-flash (P2, a different lab, not a subject here) per note (bfc). Neither is among S1–S3.
4. Materials
- Source unit:
workshop/translations/uj-pap/source-u1.txt— Mikszáth Kálmán, Szent Péter esernyője I.iii ¶1–9, 285 Hungarian words, PG #68911, public domain. The project's first Hungarian and its fifteenth source language. - Brief P: propositions M01–M11 of
E-20260731f-catalogue-reach/claims.jsonat109b103(A-mansfield-garden-party). - Brief C: propositions C01–C19 of the same file (
A-mchugh-presence). - Briefs are given to the seats in the same wording the lead worked from. Seats are not told that a second rendering will be requested, are not told what is being measured, and are not told that another model or the lead has translated the same passage.
- Comparator:
workshop/translations/uj-pap/comparator-worswick-ch3.txt— B. W. Worswick, St. Peter's Umbrella, 1900, chapter III entire (2,980 words), PG #31945. Extracted by program; its contents were never displayed to the lead, and the lead's own renderings were written without opening it. - Why the whole chapter rather than an aligned span. Aligning the span would have required reading the English. The whole chapter is a superset of the aligned text, so every longest-run figure is an upper bound on the aligned figure — but it is the identical comparator for every cell, so all between-condition comparisons, which are the only comparisons any prediction below is registered on, are exact. Absolute run lengths here are not comparable to S074's and S075's aligned-span figures and must not be quoted beside them.
5. Procedure
Per seat, two conversations, two turns each, temperature 0:
| conversation | turn 1 | turn 2 (turn 1's reply in context) | renderings |
|---|---|---|---|
| A | brief P | brief C | A.P (period, FIRST) · A.C (centre, SECOND) |
| B | brief C | brief P | B.C (centre, FIRST) · B.P (period, SECOND) |
Twelve dispatches in all (3 seats × 2 conversations × 2 turns). max_tokens 2200 per dispatch. Each raw response body is written to runs/<tag>.raw before anything is parsed (note (bdt)).
Renderings are delimited by <<<BEGIN>>> / <<<END>>> and extracted mechanically.
6. Measurements
All by tools/dependence_check.py, unmodified. Two statistics per pair: run = longest common contiguous token run; g7 = shared 7-gram count. S075 found that run failed to replicate across two sources where g7 did, so both are carried and neither is privileged.
Per seat:
| quantity | definition | what it is |
|---|---|---|
| Seq(C) | overlap(A.C, A.P) |
a centre rendering against a period rendering it had in context |
| Ind(C) | overlap(B.C, A.P) |
a centre rendering by the same seat against the same period text, written without it in context |
| Carry(C) | Seq(C) − Ind(C) | carryover into the centre arm. Same seat, same source, same comparison text; only context differs |
| Seq(P) | overlap(B.P, B.C) |
mirror image |
| Ind(P) | overlap(A.P, B.C) |
mirror image |
| Carry(P) | Seq(P) − Ind(P) | carryover into the period arm |
| W(x) | overlap(x, Worswick ch. III) | for each of the four renderings per seat — the statistic this project actually quotes |
7. Predictions, registered before any dispatch
| # | prediction | fails if |
|---|---|---|
| Q1 | Carry(C) > 0 on run in ≥ 2 of 3 seats |
≤ 1 seat, or the majority is negative |
| Q2 | Carry(P) > 0 on run in ≥ 2 of 3 seats |
as above |
| Q3 | Carry(C) > Carry(P) on run in ≥ 2 of 3 seats — the S074/S075 asymmetry, in which the centre arm moved and the period arm did not |
≤ 1 seat |
| Q4 | W(A.C) > W(B.C) on run in ≥ 2 of 3 seats — writing the centre arm second lifts its measured overlap with a published period translation, which is the S074→S075 observation |
≤ 1 seat |
| Q5 | reported, not predicted: W(A.P) against W(B.P) |
— |
| Q6 | reported, not predicted: the g7 counterpart of Q1–Q4 |
— |
The lead's own four renderings get the identical measurements and are reported beside these, labelled description.
8. Failure criteria and void rules, written before the data
- A cell is VOID if its rendering is not a translation of the unit: fewer than 150 or more than 700 English words, or containing Hungarian source strings of 5+ words, or missing the delimiters. A void cell is reported void. There is no re-briefing (note (bda) — re-specifying after seeing an inconvenient result is fitting materials to the hypothesis).
- If all three seats are void on any one condition, every prediction needing that condition reports NO RESULT.
- Interpretability floor. If Ind(C) is at or above Seq(C) and Ind(P) at or above Seq(P) for a seat, that seat is reported as showing no carryover, not as noise to be excluded.
- Q1–Q4 are reported separately and a majority failing is the result, not a reason to re-cut the statistic. The S075 lesson stands: the discarded statistic (
g7) may be the one that moves. - No threshold is set on the size of Carry. The project has no prior for what a token or two means here, and inventing one after the data is note (o)'s failure.
9. What this experiment cannot show
- It measures carryover in three specific models on one 285-word Hungarian passage. It does not measure the lead's.
- It cannot separate carryover from deliberate contrast: a seat that has just written a period rendering may write a more sharply un-period one because it remembers, which is carryover with a negative sign and would show as Carry < 0. Both directions are findings and both are pre-registered as such.
- It says nothing about whether either rendering is any good. Tier D has not passed; no evaluative claim is made or elicited anywhere in this design.
10. Verification
analysis/verify.py recomputes every reported number from the stored .raw bytes, imports nothing from analyse.py, and must fail when a body is damaged — the S075 defect, note (bgg): for every .raw it opens it asserts finish_reason == "stop" and the exact stored content length, and its mutation tests damage a real body and require a non-zero exit.
11. Pre-flight budget
| line | max_tokens |
worst case |
|---|---|---|
| pre-run critic, 1 dispatch (+1 reserve) | 12,000 | $0.25 |
| 12 subject dispatches | 2,200 each | $0.20 |
| declared worst case | $0.45 |
Worst case built from max_tokens, not from an assumed output length (note (abc)). Day headroom at session start: $3.099624934 of the $5.00 UTC cap. Lead translation is free and is not ledgered.
12. Amendments after the pre-run critic pass
Critic: qwen/qwen3.7-max, provider Alibaba, finish_reason: stop, 10,775 completion tokens, 195.6 s, $0.05315015. Declared reserve google/gemini-3.6-flash was not used; note (b) did not fire. Verdict NEEDS-REDESIGN — the second this project has taken since S021 (after S062's). Seven findings: three BLOCKING, three MANDATORY, one ADVISORY. All seven accepted. Three are answered with a stronger remedy than the one proposed; one is accepted in substance and its literal remedy declined in writing, with the reason. Raw: runs/critic.raw. Nothing above this section was edited; every change is stated here.
A1 — BLOCKING, §6. The Seq/Ind contrast did not isolate context. ACCEPTED, STRONGER REMEDY.
Finding. A.C is a turn-2 rendering and B.C a turn-1 rendering, so the contrast confounded "the prior translation is in context" with turn number, context length, and having seen the other brief.
Applied. Two further conversations per seat supply a matched independent baseline. The prior translation is present, of the same length, under the other brief, in turn 1 — but of a different passage (U2):
| conv | turn 1 | turn 2 | yields |
|---|---|---|---|
| A | brief P, unit U1 | brief C, unit U1 | A.P (P first) · A.C (C second, same-passage prior) |
| B | brief C, unit U1 | brief P, unit U1 | B.C (C first) · B.P (P second, same-passage prior) |
| X | brief P, unit U2 | brief C, unit U1 | X.C (C second, different-passage prior) |
| Y | brief C, unit U2 | brief P, unit U1 | Y.P (P second, different-passage prior) |
Ind(C) is now overlap(X.C, A.P) and Ind(P) is overlap(Y.P, B.C). Turn position, context length, both-briefs-seen and prior-translation-present are matched; the only thing that differs is whether the prior translation is of the passage being rendered. That is exactly the quantity §1 names. Eight dispatches per seat.
A2 — BLOCKING, §5. Temperature 0. ACCEPTED IN SUBSTANCE; THE LITERAL REMEDY DECLINED, WITH THE REASON, AND ANSWERED WITH A CONTROL.
Finding. At temperature 0 the brief may dominate, A.C and X.C come out identical, and Carry is trivially 0.
Declined, and why. Raising temperature to 0.7 would make every Seq − Ind difference a mixture of a context effect and sampling noise with n = 1 per cell and no repeats to separate them, which is strictly worse than the stated problem. The project runs at temperature 0 throughout, and S074 measured a 14.7% flip rate on byte-identical repeats at temperature 0 (RS-20260731f), so temperature 0 is not in fact deterministic on this stack.
Accepted, and answered. (i) Identical renderings are reported as Carry = 0, a null, and never excluded; byte-identity between any pair of a seat's renderings is reported as a fact. (ii) A byte-identical repeat of X.C is dispatched for every seat — same prompt, same seat, same session — giving a same-prompt dispatch-noise floor measured on this material rather than assumed. Nine dispatches per seat. What the repeat bounds is dispatch noise, not context-sensitivity noise, and that limit is stated wherever the floor is used.
A3 — BLOCKING, §4. The whole-chapter comparator. ACCEPTED.
Applied. The comparator is now comparator-worswick-u1.txt, 435 words, built by a stated mechanical rule and still never displayed to the lead: the English chapter's paragraphs 0–8, i.e. from the chapter's first paragraph through the last paragraph containing the string Bjela — the Vistula / Bjela-Voda aside, which is U1's closing sentence. Located by printing paragraph indices only. Worswick merges the U1/U2 boundary inside that paragraph, so the span over-runs U1 by the opening clause of U2; the over-run is identical for every cell, so between-condition comparisons are unaffected. The whole-chapter figures are still computed and reported beside the span figures, so a reader can see what the critic's objection was worth. comparator-worswick-u2.txt (paragraphs 9–18, 247 words) is built the same way for the lead's description.
A4 — MANDATORY, §§7–8. "2 of 3 seats" has a null probability of 0.5. ACCEPTED.
Applied. A prediction is CONFIRMED only at 3 of 3 seats (null probability 0.125 under a fair-coin null). 2 of 3 is reported as SPLIT and licenses no directional sentence anywhere on the result page. 1 of 3 or 0 of 3 is FAILED. This replaces the "≥ 2 of 3" wording in §7 for every one of Q1–Q4 and is registered before any subject dispatch.
A5 — MANDATORY, §9. Negative carryover is not distinguishable from noise. ACCEPTED.
Applied. A negative Carry is reported as deliberate contrast only at 3 of 3 seats, on the same rule as A4, and otherwise as a null. In addition, the A2 repeat gives a per-seat dispatch-noise magnitude; any |Carry| at or below that seat's repeat difference is reported as within dispatch noise, whatever its sign. Where the repeat comes back byte-identical, the floor is 0 and the page says so rather than treating 0 as an estimate.
A6 — MANDATORY, §7. The lead's own renderings could be woven in to rescue a failed panel prediction. ACCEPTED.
Applied. analyse.py computes and writes analysis/panel.json first and closes it, then computes analysis/lead.json. The result page's panel sections cite panel.json only, and the lead's four renderings are reported in a separate section that no panel sentence may reference. analysis/verify.py asserts that panel.json contains no key drawn from the lead cells.
A7 — ADVISORY, §§1, 6. The source's contribution is asserted to cancel and is never measured. ACCEPTED, STRONGER REMEDY.
Applied. The subtraction removes the source's expected contribution, since Seq and Ind compare two C renderings to the same A.P; what it does not bound is the residual variance. That is now measured, for free and with no extra dispatch: Cross(C) = overlap(B.C of seat i, B.C of seat j) over the three seat pairs, and Cross(P) = the same over A.P. Three independent seats rendering one source under one brief share only what the source and their common priors force, which is the baseline the finding asks for. It is reported beside Carry. The phrase "over and above what the source forces" is retained only where Cross is quoted alongside it.
13. Revised dispatch count and budget
27 subject dispatches (3 seats × 9) + the critic already spent. max_tokens unchanged at 2,200.
| line | worst case |
|---|---|
| pre-run critic | actual $0.05315015 |
27 subject dispatches, from max_tokens (note (abc)) |
$0.70 |
| revised declared worst case | $0.75 |
Day headroom after the critic: $3.046474784.
14. Amendment A8 — a run-time defect, written before the re-dispatch
Two defects, one of them the lead's own, found by running the design.
- Note (b), twenty-first firing, and the first on a translation SUBJECT rather than a critic.
deepseek/deepseek-v4-proreturnedfinish_reason: lengthwith 2,200 completion tokens and ZERO characters of content on all three of its dispatched cells. The cap was the design's declared 2,200. - The declared reserve for that seat was
qwen/qwen3.7-max, which is this design's own pre-run critic. That is the role collision the S053 fix forbids: the critic read the frozen design, so a rendering it produces is a subject that knows what is being measured. The collision was written intorun.py'sRESERVEtable by the lead and was not caught by the critic, who was not shown the runner. Three qwen renderings were produced before the run was stopped.
Applied, before any figure was computed and before the analysis was run:
- Every cell answered by
qwen/qwen3.7-maxis VOID. Run 1 of seat S3 is preserved entire underruns/void-run1-qwen/so a reader can reject this disposition. Cost of the void run: $0.1501180769. - The reserve for S3 becomes
google/gemini-3.6-flash(P2), which is neither the critic nor a subject. - S3's
max_tokensis raised from 2,200 to 8,000, for that seat only. Caps are already not uniform in effect — S1's largest completion was 1,717 tokens and S2 returned 4,231 completion tokens against a 2,200 cap, which OpenRouter did not enforce for that slug — and a cap that is not reached cannot affect a rendering. - The panel composition is NOT changed. Substituting a different lab would change the subject population after seeing that one seat failed, which is note (bda)'s shape. Raising a cap on the declared seat is not.
- If S3 fails again, the design runs on two seats and §12/A4's rule is restated on the result page in its weakened form: CONFIRMED would then require 2 of 2, at a null probability of 0.25 rather than 0.125, and that weakening is stated wherever a verdict is quoted.
Note (abc) took a hit and it is recorded rather than smoothed. The worst case was built from max_tokens, as the note requires, and max_tokens was not an upper bound on billed completion for x-ai/grok-4.5. The note's arithmetic is sound and its input was not.
Revised declared worst case: $0.95 — critic $0.05315015 (actual) + S1 $0.0712635 (actual) + S2 $0.1660948 (actual) + void run $0.1501180769 (actual) + up to $0.35 for the S3 re-run at max_tokens 8,000 on the cheapest panel seat, priced at the note (x) four-times routing caution. Day headroom before the re-dispatch: $2.659.