Repository path: workshop/experiments/E-20260730f-recall-floor/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260730f-recall-floor |
| status | frozen |
| created | 2026-07-30 |
| updated | 2026-07-30 |
| senses | — |
| internal-judgment-only | true |
| links | wiki/arms/ARM-forced-defence.md, wiki/findings/results/RS-20260728b-forced-run-ru.md, wiki/findings/results/RS-20260728-forced-run.md, workshop/translations/senilia-openings/R04-v1/translation.md, workshop/translations/korolenko-spans/R04-v1/translation.md, config/models.md, CLAUDE.md |
The recall control with a floor at chance, which ARM-forced-defence may not proceed without
Design frozen before dispatch. S065. ARM-forced-defence step 2.
0. What is owed, in the words of the page that owes it
ARM-forced-defence, absorbing a backlog row that hit the review-or-retire age at S056:
Step 2 may not report a FORCED verdict on canonical material until this control exists, and what it needs is forced choice against a distractor, where a model that knows nothing scores 50% and cannot decline.
The reason is RS-20260728b-forced-run-ru §3. Asked to write out Constance Garnett's published
English as they recalled it, all three panel models returned UNKNOWN at 24 of 24 cells. Every
recall score was 0, which is a refusal floor and not a measurement, and the same three models,
asked instead to identify the passages, named Turgenev for 8, 7 and 8 of 8 and the individual
prose poem for 6, 4 and 5 of 8 from three lines each. So the Garnett-recall confound is live and
unrefuted, and the six FORCED verdicts that session recorded — including the one supporting
CLAUDE.md's standing rule that the lead cannot be the independent third translator for a work
with a canonical English translation — each keep an alternative explanation nobody has removed.
This design removes the floor. Two options, no third; a model that knows nothing is at 50% and cannot retreat to UNKNOWN.
1. The three conditions, and why the third exists
Every item is the same shape: a Russian passage, and two English renderings of it, labelled 1
and 2. One is the published translation. The other was written by the lead from the Russian
alone, frozen at ff64960 before any comparator was opened. The question is always which is
the published one.
| condition | published side | canonicity | |
|---|---|---|---|
| G | 10 items | Constance Garnett, Dream Tales and Prose Poems, 1897 | the default English Turgenev; on Gutenberg; named by CLAUDE.md's standing rule |
| N | 10 items | Marian Fell, Makar's Dream, and Other Stories, 1916 | the only freely reachable English of any of these four Korolenko works; nobody's default |
| P | 6 items | the King James Version | the most reproduced English translation there is |
N is the empirical floor. A model with no memory of the comparator is not necessarily at 50% — it may prefer one option on register, on fluency, or because the rival was written by an author it resembles. That bias, whatever it is, is present in N exactly as it is in G: same task, same instruction, same distractor author, same language pair, same period, same regime, same session. What N supplies is the score this instrument returns when the comparator is not held, measured rather than assumed. The primary result is the difference G − N, not G against 0.5.
P is the positive control, and it is the part of this design S064 makes mandatory. If G and N both come back at chance, there are two readings — these models do not hold either comparator, and this task cannot be done by anybody. P separates them. KJV wording is held verbatim by every model that has read English text at scale, so a model that cannot pick the KJV out of a matched pastiche cannot do this task at all, and a null in G and N would be uninterpretable.
P's shape differs and the difference is stated now rather than defended later. P shows a chapter-and-verse reference instead of a source passage, because the KJV's source is Hebrew and Greek and showing it would change what the task measures. P therefore certifies the response format and the forced choice; it does not certify the source-shown shape. It is also the extreme of the canonicity range and says nothing about sensitivity in the middle of it.
2. Materials
Spans. Selected by materials/select_spans.py, committed at 9682fdb before any span was
seen. Turgenev: poems in published order, skipping every poem the lead has already translated and
every poem under 120 Russian words, opening paragraphs until the span reaches 35 words. Korolenko:
fixed offsets in four stories, first paragraph at or after the offset with ≥35 words. Two source
repairs, both declared on the translation pages (finalise_spans.py).
Rivals. T-senilia-openings-R04-v1 and T-korolenko-spans-R04-v1, 939 Russian words, frozen
with their translator's logs at ff64960. The rule the translator worked under is on those pages
and is repeated here because it is a design commitment: translate each span exactly as R04 would
have it translated if there were no experiment. A distractor made worse is a rigged control; a
distractor made Garnett-like is an invented agreement.
P's rivals are pastiches written for this design in matched register, in materials/kjv.json,
against KJV text extracted verbatim from Gutenberg #10.
Turgenev spans are openings on purpose. An opening is the most reproduced part of a short work. Korolenko's spans sit deeper in their works and are therefore, if anything, less recallable. That biases G − N upward, i.e. toward confirming the confound and leaving the arm blocked — which is the conservative direction, and it is chosen deliberately.
2b. Amendments made after §§0–7 were frozen and before any dispatch, labelled as such
A1 — two N spans are dropped, and the reason is in the comparator, not in the result. Fell 1916
is a freer translator than Garnett, and at two of the ten Korolenko spans she restructures the
passage so completely that no stretch of her English corresponds to the Russian: the Khapun exchange
(N09) and the matchmaking dialogue (N10). An item cannot be built where there is no counterpart.
Condition N therefore has 8 items and 24 primary units, not 10 and 30. Both spans remain in
T-korolenko-spans-R04-v1, which is a translation and not an item list.
A1 has a consequence the design must own: requiring alignment selects, inside Fell, for the passages she rendered closely. G is Garnett throughout, who is close throughout; N is Fell where Fell is close. If free rendering were itself a cue — the option that departs from the Russian is the published one — the selection removes it from N and there was never any to remove from G. The direction is toward comparability; the selection is still a selection and is reported.
A2 — two published spans carry one added closing quotation mark each (N05, N08), because
Fell's quotation continues past the span boundary and an unbalanced quotation mark is a typographic
cue with nothing to do with memory. Declared in materials/build_items.py; the verbatim assertion
strips the added mark before comparing, so it is still an assertion against the real text.
A3 — the FB3 covariates were computed before dispatch, because they are properties of the
materials and not of the outcome (analysis/covariates.py, runs/covariates.json).
| G (10) | N (8) | Welch t | p | |
|---|---|---|---|---|
| length ratio published ÷ rival | 0.934 | 0.988 | −1.021 | 0.323 |
| token Jaccard, published ~ rival | 0.485 | 0.400 | 2.011 | 0.066 |
| longest common run, tokens | 7.00 | 6.13 | 0.905 | 0.382 |
FB3 does not fire. But the Jaccard row is close to firing and it points the way that matters: the lead's Turgenev rivals are more similar to Garnett than its Korolenko rivals are to Fell, so G's items are the harder discrimination of the two. That is a bias against finding G above N — against confirming the confound, and therefore toward FB1. It is a limit on the null reading and it is stated here rather than after the numbers. Added to the analysis: the Spearman correlation between item Jaccard and item accuracy, pooled over G and N, so that the size of this cue is reported rather than argued about.
A5 — the design would have measured punctuation, and it was caught by reading the built prompt
rather than the design. Gutenberg's Garnett prints ellipses as three ASCII dots and speech in
single curly quotes; Fell prints double curly quotes; the lead's own prose uses the ellipsis
character and, in one poem, guillemets. A model could have picked the published rendering off the
punctuation alone with no memory whatever, and the first built prompt shows it doing so would have
been trivial — Windlessness, warmth ... air like new milk! against warm weather… the air is.
typographic_neutralise() now rewrites every option in every condition: ellipsis character to three
ASCII dots at a fixed spacing, every quotation mark to the straight double mark, every dash form to
a spaced em dash, whitespace collapsed. It touches no word. assert_no_typographic_cue()
asserts afterwards that no banned mark survives anywhere, and the assertion runs again in
analysis/verify.py.
A6 — the same problem one level up: dialogue marking. R04's standing typographic ruling marks
spoken dialogue with a leading dash, the Russian convention; both comparators use quotation marks.
That is a regime decision rather than a wording decision, and left standing it would let a model
pick the published option off the dialogue marking at four of the eight N items — inflating N,
shrinking G − N, and pushing the result toward FB1, the branch that unblocks the arm. The four
rivals with spoken dialogue are therefore re-marked with quotation marks for the item only; the
translation artifacts keep the dash exactly as frozen. assert_punctuation_only() asserts that each
override differs from the frozen rival in punctuation alone — letters stripped, the two strings are
identical — so no wording can hide inside the normalisation.
A5 and A6 are the two amendments this session made without a critic telling it to, and both were found by reading the dispatch text. Recorded that way rather than folded into §2 silently.
A4 — pre-flight estimate, note (abc), built from max_tokens and not from an assumed output.
Worst case per choice call at list out-price: P1 $0.035, P2 $0.036, P3 $0.032 on a 4,000-token cap;
18 calls ≈ $0.65. Critic at max_tokens 16,000 on P4 ≈ $0.26. Declared worst case
$0.96, against $3.27 of headroom on the UTC day. Routing can multiply a per-call figure by ~4
(S022), so the declared figure is a list-price worst case and not a bound.
2c. Amendments from the independent pre-run critic pass (before dispatch)
Critic: deepseek/deepseek-v4-pro (P5), the declared reserve. P1, P2 and P3 are subjects and
may not critique the design; P4 moonshotai/kimi-k3 was called first and returned
finish_reason: length with content: null at max_tokens 16,000, provider Fireworks,
$0.4149495 for nothing — standing note (b), thirteenth session, and the single most expensive
firing of it in this ledger. The rejected body is preserved at critic.attempt0.raw and was not
overwritten; the reserve's body is critic.attempt1.raw. Verdict NEEDS-AMENDMENT, seven
findings — two BLOCKING, two MANDATORY, three ADVISORY. All seven accepted. Note (rr),
twenty-third consecutive session in which a critic changed a design.
- C1 — BLOCKING, accepted, and it was live in the dispatch text. At G10 the published rendering
was a single block and the rival carried line breaks between speaker turns, because
typographic_neutralise()collapsed spaces and tabs and not newlines. A model could have taken that item off the paragraphing. All whitespace is now collapsed. A second instance the critic did not see was found while fixing this one: the rival extractor was picking up the artifact's---rule, leaving G10's rival ending in a stray dash. - C2 — BLOCKING, accepted, and it is the finding that would have decided the session. Fell's
transliterations are not the lead's: Aksana/Oksana, Raman/Roman, Lavrovski/Lavrovsky,
Tiburtsi/Tyburtsy, Tartar/Tatar. A model that knows the published spellings could have scored in
N without recalling one word of the passage. That cue exists in N and not in G, so it inflates
N, shrinks G − N and pushes the result toward FB1 — the branch that unblocks the arm. The
rival's spellings are conformed to the published ones at four items, five substitutions, declared
in
build_items.pyand asserted there and in the verifier. - C3 — MANDATORY, accepted. The permutation test shuffled condition labels across all units without stratifying, so a model main effect could enter the null. It now shuffles within each model.
- C4 — MANDATORY, accepted, and it narrows this design's central claim. Fell 1916 is on Gutenberg too. N is not a zero-memory floor; it is a low-canonicity comparator, and the ID condition measures recognition of the Russian author, not memory of Fell's English. G − N is therefore a lower bound on Garnett-specific memory, not an estimate of it, and §1's phrase "the score this instrument returns when the comparator is not held" is withdrawn and replaced by that sentence.
- C5 — ADVISORY, accepted. ID runs after the choice conditions, so exposure to the passages may inflate recognition. ID scores are an upper bound on prior recognition and are reported as one.
- C6 — ADVISORY, accepted, and it weakens FB1. P shows a verse reference where G and N show a source passage. A null in G and N with P passing therefore leaves open that the source-shown shape is itself the obstacle, rather than that memory is absent. §1's claim that P separates the two readings is true of the response format and the forced choice, and not of the shape. FB1's conclusion inherits this and says so.
- C7 — ADVISORY, accepted. FB3 covers length and Jaccard only. C1 and C2 are the two cues it
would have missed; a manual audit of the built prompts was run for others (orthography, dialogue
marking, paragraphing, trailing punctuation) and its results are the assertions now in
assert_no_typographic_cue(). No claim is made that the audit is exhaustive.
3. Panel, calls, and the counterbalance
Three models, one lab each, temperature 0, max_tokens 4000. Slugs resolved from config/models.md
at run time: P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5
— the same three as S044 and S045, deliberately, so the three results are comparable.
| calls | what |
|---|---|
| 6 | G, three models × two orders |
| 6 | N, three models × two orders |
| 3 | P, three models × one order |
| 3 | ID, three models, one call each: the 20 Russian spans, unlabelled, name the author and the work |
| 18 | total dispatches, plus one independent pre-run critic call |
Order is counterbalanced because this project has measured slot preference at 0.50–0.85 and has never found it absent. Order a puts the published rendering in slot 1 for odd-numbered items and slot 2 for even; order b is the exact swap. Every item is therefore seen by each model twice, once in each arrangement, and the pair doubles as a within-day repeat control — which note (bev) made non-optional yesterday, when a byte-identical same-day prompt moved 24.3% of one model's codes on a classification task.
ID exists so that a null in N is readable. If N comes back at chance, the first objection is that the models do not recognise Korolenko at all, and that objection is answerable only by measurement. It also replicates S045's condition E on Turgenev — 8/8, 7/8, 8/8 — which nobody has repeated. ID is dispatched after G, N and P are complete, so that naming the authors cannot cue the recall task; the API is stateless and no ordering effect is possible, but the ordering is kept anyway.
No model is asked to judge its own output, and nothing here is a quality judgment: the question is a matter of textual fact, which is the register in which S015 found this panel strong.
4. Scoring, frozen
The datum for each (item, model, order) is the chosen slot, mapped to correct / incorrect.
Refusals and unparseable answers are scored NOANSWER and reported separately; the instruction
forbids them and there is no UNKNOWN option.
The primary unit is the (item, model) pair, scored across its two orders:
- 1.0 if the model picked the published rendering in both orders
- 0.0 if it picked the rival in both
- 0.5 if it flipped
So G and N each have n = 30 primary units and a chance value of 0.5. P, run in one order only, has 18 (item, model) judgments and is scored as a plain proportion.
Reported alongside, all pre-registered:
- slot-1 rate per model per condition — the position bias the counterbalance is neutralising
- flip rate per model per condition — the same-day, same-content inconsistency of note (bev), measured here for the first time on a two-alternative task
- discriminability covariate: for each item, the token-level Jaccard similarity between the published and rival renderings, and the longest common run between them. Secondary pool = items whose two options share fewer than 40% of tokens; the primary pool is all items, and no item is dropped.
- length cue: mean and SD of the word-count ratio (published ÷ rival), per condition. If these differ materially between G and N, the G − N comparison is confounded by a cue that has nothing to do with memory, and the result page says so.
Statistical rule, fixed now. G − N is tested by a two-sided exact permutation test over the 60 primary units (10,000 permutations, seed 20260730, labels shuffled between conditions). G against 0.5 and N against 0.5 are each tested by a two-sided exact binomial on the counts of 1.0 and 0.0 units, with 0.5-units split evenly and rounded down in the direction that makes the test less significant. α = 0.05, and there are three tests, so the Holm correction is applied and the adjusted values are what the verdicts are read off.
5. Pre-registered predictions
Written before any comparator was opened, before any item was built, and before any call.
- P-a. P ≥ 15 of 18. The KJV is held; a pastiche does not survive next to it.
- P-b. G is above chance — G ≥ 0.65 on the primary unit — because S045's condition E showed these three models identify these prose poems from three lines, and identification at that level makes wording memory likely rather than merely possible.
- P-c. N is at chance, within [0.40, 0.60].
- P-d. G − N ≥ 0.15 and the permutation test rejects. This is the prediction the arm's status turns on: it says the confound is real, Garnett-specific, and measurable.
- P-e. The flip rate is not zero and is at least 10% in at least one model × condition cell. Note (bev) was measured on a classification task; two-alternative forced choice is an easier task and should be steadier, but not perfectly steady.
- P-f. ID: all three models name Turgenev for ≥6 of the 10 Turgenev spans, replicating S045's condition E, and no model names Korolenko for more than 5 of the 10 Korolenko spans.
6. Failure criteria, and they bind pages outside this one
FB1 — the branch that unblocks the arm and reopens six verdicts. If P ≥ 15/18 (the
instrument works) and G is not above chance after Holm correction and G − N does not reject,
then these models do not discriminate Garnett's wording from a matched rival at these ten loci.
The Garnett-recall confound is then refuted for this corpus, RS-20260728b-forced-run-ru §3's
"live and unrefuted" must be restated, §5's FB2 stops being Untestable and becomes did not
fire, and the six FORCED verdicts of that session become readable as source constraint. CLAUDE.md's
standing contamination paragraph and ARM-forced-defence are amended in the same session, not a
later one.
FB2 — the branch that costs the instrument. If P < 12/18, the forced-choice format does not work on these models and nothing in G or N may be read at all — not as evidence of memory, not as evidence of its absence. The session reports a void and the arm stays blocked. A positive control that is allowed to fail quietly is not a positive control, which is S064's lesson at the cost of a whole session's headline.
FB3 — the branch that costs the comparison rather than the instrument. If the length-cue or discriminability distributions differ between G and N at a two-sided Welch p < 0.05, then N is not a matched floor for G and G − N is reported as confounded, whatever it comes to. The individual condition results survive; the contrast does not.
FB4 — the branch that says the rivals were badly made. If G > 0.85 and N > 0.85, the models are picking the published rendering everywhere, which on this design most likely means the lead's rivals are identifiable as non-published prose rather than that Fell is held in memory. That is a defect in the materials, it is this session's finding, and no memory claim follows from it.
These four are written before dispatch and they bind. FB1 and FB4 point in opposite directions and cannot both fire; FB2 voids FB1, FB3 and FB4 alike.
7. What this cannot settle, listed before it can be forgotten
- Nothing about the lead's own memory. The panel is a proxy for an independent reader who does
not hold the comparator, not for the lead. This project cannot measure what its own lead agent
holds.
RS-20260728b-forced-run-rulimit 4, inherited unchanged. - Recognition is not recall and neither is discrimination. A model may hold Garnett's cadence without holding her words, and pick correctly for that reason. A positive G is consistent with several mechanisms and identifies none.
- The rivals are the lead's, and the lead has read a great deal of nineteenth-century English translation. If the lead's prose is systematically Garnett-flavoured, G is depressed — the conservative direction for FB1 is the opposite one, so this is a limit on the positive reading, not on the null.
- n = 10 items per condition, three models, one language pair, two comparators. Nothing here says how the forcing probe behaves on Chinese, Italian or French material, where this project also has measured figures.
- A refuted confound does not clear a run. It removes one alternative explanation of the S045 FORCED verdicts. Those verdicts still only separate source-forced from not source-forced, and the panel's convergence on a shared prior remains untouched (A1, S045).
- P is the extreme of the canonicity range. The natural next measurement — Garnett's own most read book, against these same rivals' method — is not in this design and is named as step 3.
8. Provenance
- Selection rule frozen at
9682fdb, before any span was seen. - Both translations and both translator's logs frozen at
ff64960, before any comparator was opened. - §§0–7 of this page, all six predictions and all four failure branches frozen in git before the comparators were aligned and before any dispatch; the freezing commit is the one that creates this file.
- Raw responses, costs and
providerfields underruns/. Comparator texts live in the session scratchpad and are not stored in the repository, exceptfell1916-makar.txt, which S026 already stored and this session did not open.