Repository path: workshop/experiments/E-20260728b-forced-run-ru/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260728b-forced-run-ru |
| status | frozen |
| created | 2026-07-28 |
| updated | 2026-07-28 |
| senses | — |
| internal-judgment-only | true |
| links | workshop/experiments/E-20260726c-forced-or-borrowed-ru/design.md, wiki/findings/results/RS-20260726c-forced-or-borrowed-ru.md, wiki/findings/results/RS-20260728-forced-run.md, workshop/translations/senilia/R04-v1/translation.md, workshop/translations/jeli-il-pastore/R05-v1/translation.md, wiki/arms/ARM-forced-defence.md, CLAUDE.md |
The forcing defence, retested on the run the standing rule is built on
Design frozen before dispatch. S045. ARM-forced-defence step 1.
0. What is being tested and why it is not a repeat
RS-20260728-forced-run (S044) tested the objection "the source forces it" on a lead
translation for the first time and refused it, on a 12-token Chinese run. Its §5 says what it left
owed:
Every published contamination figure in this project is retestable for under two cents each, and none has been tested. The five prior lead translations with measured runs — above all the 21-token Turgenev run that the standing rule in
CLAUDE.mdis built on — were each declared on the run alone.
This session runs the Turgenev cell. It is the cell that matters, because CLAUDE.md's standing
rule — the lead cannot serve as the independent third translator for any work whose standard English
translation is canonical — cites exactly one measurement, and it is here.
And it adds the control S044 could not run. That page's own limit 3 reads: "Three model outputs are not three translators, and all three share whatever the Yangs contributed to their training." On Garnett this is far sharper than on the Yangs: Garnett's Turgenev is public domain, sits on Project Gutenberg, and is among the most reproduced English translations of the nineteenth century. If the panel models simply hold Garnett's sentences, a FORCED verdict on this material means nothing at all — and nobody has ever measured whether they do. §4 measures it.
1. Materials, and the reproduction check that comes with them
The subject is T-senilia-R04-v1 — eight Turgenev prose poems, translated by the lead from the
Russian alone, contamination: high, frozen at 5f78bab with the answer key committed one commit
earlier at 4d59b31 and never opened.
RS-20260726c-forced-or-borrowed-ru reports the eight lead-vs-Garnett longest runs as lengths
only — 21, 18, 17, 16, 16, 15, 13, 11 — and prints the text of two of them. loci.py recomputes
all eight from the lead's frozen English and Garnett 1897 (Gutenberg #8935, fetched fresh) using
tools/dependence_check.longest_common_run, unmodified.
All eight lengths reproduce exactly. That is an independent reproduction of a published figure by a script written two sessions later against a freshly fetched comparator, and it is recorded as such rather than assumed.
| # | unit | poem | run | tok |
|---|---|---|---|---|
| L1 | 4 | Соперник | i was not frightened i was not even surprised but raising myself a little and propping myself on my elbow i | 21 |
| L2 | 39 | Монах | it he would stand so long on the cold floor of the church that his legs below the | 18 |
| L3 | 11 | Восточная легенда | part what need have you of the white apple you are wiser than solomon as it is | 17 |
| L4 | 3 | Собака | storm is howling the dog sits in front of me and looks me straight in the | 16 |
| L5 | 33 | Голуби | of snow it was a white dove flying from the direction of the village it flew | 16 |
| L6 | 27 | Пир у Верховного Существа | the supreme being took it into his head to give a great banquet in his | 15 |
| L7 | 17 | Последнее свидание | my heart sank i sat down on a chair beside him and involuntarily | 13 |
| L8 | 20 | Насекомое | among us were women children old men we were all talking | 11 |
Displaying these runs primes nothing. Every one of them is text the lead itself wrote and froze in S026; what is new is only that it also appears in Garnett. Method note (bcp) — the caution that measuring between a draft and its revision tells the translator which of its own strings match a published version — does not bite here, because there is no further Turgenev to translate in this arm.
A ninth locus, and it is prospective. T-jeli-il-pastore-R05-v1 span 4 was translated and frozen
in this same session at c6f19d7, before any comparator was touched, and then measured against
Dole 1896:
| tokens | 7-grams | 7-grams, names excl. | 12-grams | 12-grams, names excl. | longest run | |
|---|---|---|---|---|---|---|
| span 4 vs Dole | 2,579 / 12,970 | 103 | 10 | 7 | 0 | 14 |
| span 4 vs null control | 2,579 / 15,342 | 0 | 0 | 0 | 0 | 6 |
| # | unit | work | run | tok |
|---|---|---|---|---|
| L9 | — | Verga, «Jeli il pastore» ¶114 | mara no i haven't forgotten you but i was so tired after the fireworks | 14 |
L9 is the other end of the variable the standing rule turns on. Garnett is canonical for Turgenev in English; Dole 1896 is an obscure Boston selection and D. H. Lawrence's better-known Vita dei campi has never been reachable here. A run against a canonical comparator and a run against an obscure one, put through the same instrument in the same session, is the contrast this design exists to draw.
2. Panel and conditions
Three models, one lab each, temperature 0, max_tokens 3000, one call per model per condition —
twelve calls. Slugs resolved from config/models.md at run time: P1 openai/gpt-5.6-terra,
P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5. The same three as S044, deliberately, so
the two results are comparable.
| condition | prompt | what it measures |
|---|---|---|
| A — blind RU | the eight Russian passages, "render each into English prose". No author, no title, no attribution, no English. | would an independent translator produce the run? |
| B — recall RU | the same eight passages, "write out Constance Garnett's published English rendering as you recall it; if you do not recall it, write UNKNOWN. Do not translate." | does the model hold Garnett's text? |
| C — blind IT | the Italian passage at L9, same instruction as A | the same question at the non-canonical end |
| D — recall IT | the Italian passage at L9, naming Nathan Haskell Dole 1896 | is the obscure comparator held? |
Naming the author in A would destroy the measurement, so A and C name nothing. B and D name the translator on purpose and are separate calls; the API is stateless, so no ordering effect is possible, and A/C are nonetheless dispatched first.
3. Scoring, frozen, and the rule is deliberately the permissive one
For every (locus, model, condition), the datum is the longest contiguous stretch of that locus's
run-tokens appearing anywhere in that model's output for that locus, computed by score.py under
tools/dependence_check.tokenise. That integer is rule-free and is reported in full.
Two verdict rules, both registered here:
- PRIMARY — S044's rule, unchanged. A locus is FORCED if ≥2 of 3 outputs contain the run verbatim or ≥7 of its tokens contiguously. Otherwise NOT FORCED.
- SECONDARY — proportional. FORCED if ≥2 of 3 outputs contain ≥⅔ of the run's tokens contiguously (14 of 21, 12 of 18, 12 of 17, 11 of 16, 10 of 15, 9 of 14, 9 of 13, 8 of 11).
Why the absolute rule is primary, stated before the numbers exist. It is permissive on long runs — a fixed floor of 7 asks a 21-token run to clear a third of itself and an 11-token run to clear two thirds — and being permissive means it is biased toward FORCED, which is the verdict that would overturn this project's own standing rule. A rule change that made the project's existing position easier to keep would be a degree of freedom taken in the convenient direction; keeping S044's rule is not. The proportional rule is reported because the absolute one plainly under-describes a 21-token run, and where the two disagree that disagreement is the result, not something to be resolved in favour of either.
Confound flag. A locus where condition B returns ≥7 contiguous tokens from ≥2 of 3 models is marked CONFOUNDED: at such a locus the models demonstrably hold Garnett, so a FORCED verdict in condition A cannot be attributed to source constraint and is not counted as one.
Is the recall condition inert? A model may answer B by translating and calling it recall. So
score.py also reports, per (locus, model), the token-level similarity between that model's A output
and its B output. If A ≈ B everywhere, condition B measured nothing and must be reported as
uninformative rather than as evidence of no recall.
4. Pre-registered predictions
- P-a. L1 is FORCED under the primary rule and NOT FORCED under the secondary. «Я не
испугался — даже не удивился» has a near-forced English of about eight tokens; «опершись на локоть»
→ "propping myself on my elbow" is elective, and
RS-20260726c§the-worst-case already argued so in writing. - P-b. At least 4 of the 8 Russian loci are FORCED under the primary rule. The runs are long and contain long forced stretches.
- P-c. No more than 2 of the 8 are FORCED under the secondary rule.
- P-d. The Garnett-recall confound is real: at least one model at at least one locus returns ≥7 contiguous tokens in condition B, and condition B's outputs are measurably unlike condition A's for at least one (model, locus).
- P-e. L9 is FORCED under both rules, and condition D returns no locus with ≥7 contiguous. The obscure comparator is not held in memory and the run against it is source-constrained.
5. Failure criteria, and they bind pages outside this one
FB1 — the branch that costs this project its standing rule. If L1 is FORCED under the
secondary rule as well — ≥2 of 3 models independently producing ≥14 contiguous tokens of the
21-token run — and L1 is not flagged CONFOUNDED, then the 21-token run is source-constrained,
CLAUDE.md's parenthetical "canonical and unavoidable" is literally true in a sense the project
never established, the run stops being evidence of recall, and the standing rule loses its
headline support. CLAUDE.md, ARM-overlap-dependence §consequences and
CL-20260726-lead-centrality §range must then be restated on whatever the remaining seven loci
support, or the rule withdrawn. This is written before dispatch and it binds.
FB2 — the branch that costs the instrument. If ≥4 of the 8 loci are flagged CONFOUNDED, the
forcing probe is confounded on canonical material as a class, and RS-20260728-forced-run's licensed
sentence must be narrowed to: a shared run is not explained by source constraint unless the
constraint has been demonstrated on somebody else who does not already hold the comparator**. That
would make the S044 method weaker than S044 claimed, and it would be this session's headline.
FB3. If P-e fails and L9 is NOT FORCED under the primary rule, then span 4 of the Verga
carries a 14-token run that is not source-constrained against an obscure comparator, and
T-jeli-il-pastore-R05-v1's contamination: suspected must be reconsidered upward on the page.
A control that can only confirm the convenient answer is not a control — S044's words, and the reason all three branches are written here rather than after the numbers.
6. What this cannot settle, listed before it can be forgotten
- Three model outputs are not three translators, and condition B bounds only what a model will emit on request. A model may hold Garnett and fail to produce it when asked; a low recall score weakens the confound, it does not eliminate it. The control is one-sided in the same way the probe is.
- Mechanism is untouched. Recall, a shared frequency prior, and free convergence all predict a match. The probe refuses a defence; it identifies nothing.
- A FORCED verdict never clears a run. It removes one reading of it.
- n = 8 loci on one translator, one language pair, one comparator, plus one Italian locus. Nothing here says how often the defence fails in general.
- Two of the eight runs are known to be near-all function words (
RS-20260726c§what-the-shared- runs-actually-are: L4's poem and L7's poem supplied the two runs with three and two distinct content tokens). They are kept in, because dropping them after seeing that would be selecting the data, and they are expected to behave differently.
6b. Amendments from the independent pre-run critic pass (2026-07-28, before dispatch)
Critic: qwen/qwen3.7-max, the declared first reserve — P1, P2 and P3 are subjects in this
design and may not critique it, P4 is off the call list (S044), and P5 deepseek/deepseek-v4-pro
was tried first and returned finish_reason: length with content: null at max_tokens 6000,
$0.019293024 for nothing. Standing note (b), sixth session bitten; dropped rather than retried, on
S044's own lesson. The runner overwrote that raw body on the substitute's first attempt — a
defect in this session's own script, recorded rather than hidden; the provider (Novita), finish
reason and cost were read off it before it was lost and are ledgered.
Nine findings accepted, one rejected. §§0–6 above are not edited; these amendments govern.
- A1 — accepted, and it is the sharpest. "the run stops being evidence of recall" over-reaches: a NOT FORCED verdict shows the panel does not produce the run. If the lead and Garnett share a prior the panel lacks — nineteenth-century literalism, or simply a translationese register — then the run is convergence on a shared prior and neither recall nor source constraint. The probe separates "source-forced" from "not source-forced" and nothing finer, and every sentence in the result page must be written to that limit.
- A2 — accepted. The CONFOUNDED flag is one-sided in both directions: a model may produce a
Garnett-flavoured pastiche and be flagged CONFOUNDED without holding the text, and a model may hold
the text and decline to recite it. Amendment: an explicit refusal is scored
REFUSED, kept separate from a low score and fromUNKNOWN; a locus where ≥2 of 3 models refuse is flaggedRECALL-UNTESTEDand may not be read as NOT CONFOUNDED. - A3 — accepted as a real gap in the design, resolved here. FB1 and FB2 can both fire (L1 clean and forced while L2–L8 are confounded). They are not in conflict and neither takes precedence: FB2 narrows what the method licenses in general, FB1 acts on L1 in particular, and if both fire the result page states both. Written now so it is not decided after the numbers.
- A4 — rejected, with the reason. The critic says a refusal-driven NOT CONFOUNDED would "conveniently allow the FORCED verdict to stand and protect the standing rule". The direction is backwards. A clean FORCED verdict at L1 fires FB1, which costs this project its standing rule. The underlying observation (refusals produce spurious NOT CONFOUNDED) is correct and is handled at A2; the claim about whose convenience it serves is not.
- B1 — accepted. A 7-token floor on a 21-token run is 33% and is reachable by common idiom. The
design already reports both rules; added:
score.pynow reports, for every matched stretch, the number of distinct content tokens in it (stoplist-free definition: tokens not in the run's own function-word set), so a match made of of the direction of the is visible as one. - B2 — partly rejected on the implementation.
tools/dependence_check.tokenisestrips punctuation and works on whitespace words, so a sentence boundary does not break contiguity and subword splitting does not arise. And the metric is slice-based, not prefix-based — it takes the longest contiguous slice of the run occurring in the output — so L2 beginning mid-phrase with "it" costs at most that one token. Stated because the critic could not see the implementation. - C1 — accepted. "one clause of surrounding context" is not a constant: the passages run from 15 to 47 source words. The design does not claim uniformity and now says so; the per-locus context length is reported with the results.
- C2 — accepted. L3's run contains solomon and L9's contains mara. Amendment: the run length with name tokens excluded is reported for every locus, and no verdict on L3 or L9 is stated without it.
- C3 — accepted, and it belongs in §6. These are the longest run per poem, i.e. an extreme-value selection, so they are systematically likelier than a random span to be boilerplate, formula, or function words. Any FORCED verdict here is therefore easier to obtain than at a random locus, and the NOT FORCED verdicts are correspondingly the stronger ones.
- E1 — accepted, and the design withdraws an over-claim. §1's "L9 is the other end of the variable the standing rule turns on" claims more than one locus can carry. Restated: L9 is a single prospective case whose value is that its prediction was registered before the run existed and before the comparator was fetched, not that it characterises obscure comparators as a class. P-e is a single-case prediction and is reported as one.
- G1 — accepted, and it is the hole the design did not see. The inference is about the lead's contamination, and the panel is a proxy for an independent translator, not for the lead. Nothing here can establish whether the lead holds Garnett; it can only establish whether the wording was available to translators who do not. Added to §6 as limit 6.
6c. Condition E — added POST HOC, after conditions A–D were scored, and labelled as such
Why. Condition B returned UNKNOWN at 24 of 24 (locus × model) cells, and condition D at 3 of
3. Every recall score is therefore 0, and that is not a measurement: it is the floor the critic
predicted at finding D — "UNKNOWN is not usable; it yields zero tokens, acting as a false
negative for the CONFOUNDED flag (the model might hold the text but output UNKNOWN to avoid
hallucinating)." The memory control did not run. The confound RS-20260728-forced-run §4 limit 3
named is still unmeasured, and no locus may be described as not confounded on the strength of
condition B.
What is added, and its weight. A cheap partial proxy: identification. Three calls, the same eight Russian passages, one instruction — name the author and the work. Identification is a task models attempt rather than decline, so it has no UNKNOWN floor.
It bounds the confound only in one direction and weakly. A model that cannot identify Turgenev from three lines of his prose is unlikely to be holding Garnett's English of it; a model that can identify him may still hold no English at all. Recognition and translation-memory are different capacities, and condition E cannot tell them apart.
Pre-registered before dispatch, after A–D were scored and before E was called:
- P-f. All three models name Turgenev for ≥6 of 8 passages, and at least two models name the individual prose poem for ≥4 of 8.
- Reading, fixed now. If P-f holds, the models demonstrably hold this corpus and the Garnett-recall confound stays live and unrefuted — every FORCED verdict in §4 keeps an alternative explanation. If P-f fails, the confound is weakened and not removed, for the reason in the paragraph above.
This is post-hoc and is reported as post-hoc. It was added because a pre-registered control failed, its prediction was written before it ran, and it does not touch any verdict in conditions A–D.
7. Provenance
- Span 4 of the Verga, and its whole log, frozen at
c6f19d7before any comparator was fetched. loci.pyandspan4_contamination.pydisplay no comparator prose except the shared runs themselves, which are by construction text the lead already wrote.- §§0–6 of this page, including all five predictions and all three failure branches, frozen in git before dispatch; the freezing commit is the one that creates this file.
- Raw responses, costs and provider fields under
runs/. Comparator texts live in the session scratchpad and are not stored in the repository.