Repository path: wiki/findings/results/RS-20260730f-recall-floor.md · rendered 2026-09-09
Page metadata (front matter)
| type | result |
|---|---|
| id | RS-20260730f-recall-floor |
| status | active |
| created | 2026-07-30 |
| updated | 2026-07-30 |
| senses | — |
| internal-judgment-only | true |
| links | workshop/experiments/E-20260730f-recall-floor/design.md, wiki/arms/ARM-forced-defence.md, wiki/findings/results/RS-20260728b-forced-run-ru.md, wiki/findings/results/RS-20260728-forced-run.md, workshop/translations/senilia-openings/R04-v1/translation.md, workshop/translations/korolenko-spans/R04-v1/translation.md, wiki/method-notes.md, CLAUDE.md |
The recall control has a floor at chance now, and the floor condition inverted: three models name the work at 18 of 18 and pick the lead's translation over the published one at 46 of 48
S065. E-20260730f-recall-floor, ARM-forced-defence step 2. 18 dispatches, all accepted; plus
two critic calls, one wasted. $1.0005113722. Selection rule frozen at 9682fdb before any span was
seen; both translations and both translator's logs frozen at ff64960 before any comparator was
opened; design, six predictions and four failure branches frozen at 272c9c3 before the comparators
were aligned and before any dispatch. Verification: 452 checks, 0 failures, by a verifier that
imports nothing from the scorer or from tools/.
1. What was owed, and what was built
ARM-forced-defence could not proceed. Its step 2 carries an absorbed precondition:
Step 2 may not report a FORCED verdict on canonical material until this control exists, and what it needs is forced choice against a distractor, where a model that knows nothing scores 50% and cannot decline.
The reason is RS-20260728b-forced-run-ru §3: asked to write out Garnett's published English, three
models returned UNKNOWN at 24 of 24 cells. A refusal floor, not a measurement.
This session built the control. A Russian passage, two English renderings, one published and one written by the lead from the source alone; pick the published one; guessing required, refusal scored as an error. Ten items on Turgenev against Constance Garnett 1897 (condition G), eight on Korolenko against Marian Fell 1916 (condition N), six on the King James Version (condition P, the positive control). Three models, and G and N in both slot orders, so every (item, model) pair is seen twice and scores 1.0 / 0.5 / 0.0 for consistently right, flipped, consistently wrong.
2. Result
| items | units | mean | 1.0 | 0.5 | 0.0 | flip rate | vs chance (Holm) | |
|---|---|---|---|---|---|---|---|---|
| P — KJV | 6 | 18 | 1.000 | 18 | 0 | 0 | 0.0% | — |
| G — Garnett | 10 | 30 | 0.467 | 9 | 10 | 11 | 33.3% | p = 0.856 |
| N — Fell | 8 | 24 | 0.042 | 0 | 2 | 22 | 8.3% | p = 8.9e-06 |
G − N = +0.425, stratified permutation p = 1.0e-04 (Holm 2.0e-04).
Per model, and the pattern is the same in all three: G 0.50 / 0.50 / 0.40 (P1 / P2 / P3), N 0.063 / 0.000 / 0.063.
Three things happened, and only one of them was predicted.
- P passed at ceiling. 18 of 18, no flips, no refusals. The response format works, the forced choice works, and these models can express a hit. FB2 does not fire, so nothing here is void on instrument grounds. P-a held.
- G is at chance. 0.467 against 0.5, p = 0.856, and a third of the (item, model) pairs flipped when the two options changed slots — the signature of guessing rather than of a weak signal. P-b is falsified: the prediction was G ≥ 0.65.
- N did not return a floor. It inverted. Zero of twenty-four units correct in both orders; 46 of 48 individual judgments picked the lead's rendering as the published one. P-c is falsified, and not in the direction any branch anticipated.
P-d "held" and the way it held makes it worthless. The prediction was G − N ≥ 0.15 with the test rejecting, offered as evidence that the confound is Garnett-specific. The difference is +0.425 and the test rejects at 1e-04 — but it is produced entirely by N collapsing, not by G rising. A prediction that is satisfied by the wrong term of its own difference is not confirmed; it is uninformative, and it is recorded that way.
3. Why N inverted, and why the answer is not ignorance
The obvious explanation for N — the models do not know Korolenko — is refuted by this session's own identification condition.
| answered | named the author | named the individual work | |
|---|---|---|---|
P1 gpt-5.6-terra |
18/18 | 10/10 Turgenev, 8/8 Korolenko | 7/10, 8/8 |
P3 x-ai/grok-4.5 |
18/18 | 10/10 Turgenev, 8/8 Korolenko | 7/10, 8/8 |
P2 gemini-3.6-flash |
0/18 | — | — (truncated at the token cap; a truncated identification call can only lose answers, so nothing is inferred from it) |
Both models that answered named Korolenko at every Korolenko span and «Лес шумит», «Сон Макара» and «В дурном обществе» correctly at all eight — from forty words each. They know exactly what they are looking at, and they still take the lead's rendering over Fell's, unanimously.
So the property they are using is not memory of the published wording. The most economical reading is that Fell 1916 renders freely — she compresses, reorders, and at N05 reverses a clause the Russian negates — and the lead's R04 rendering is close. Asked which of these two was published, the models appear to answer which of these two is the better or more faithful translation, and the close one wins. That is a systematic non-memory cue, and it is the same class of thing this design spent four amendments removing from the surface of the items (punctuation, dialogue marking, paragraphing, transliteration). It survived because it is in the prose, not on it.
This is the session's finding for the method, and it is a general one. S045 established that free recall does not measure memory on canonical material — it returns a refusal floor at zero. S065 establishes that forced choice does not either, at least not where the comparator is unusual: it returns a confident answer to a different question. Two elicitations, two failures, and the failures are of opposite shapes. Method note (bez).
4. What follows for the arm, stated against the branches as written
FB1 does not fire on its literal wording, and the literal wording was the wrong test. FB1 required P ≥ 15/18 (held, 18/18), G not above chance after Holm (held, p = 0.856), and G − N not rejecting (failed). The third clause was written as a proxy for no Garnett-specific signal, on the tacit assumption that N would sit near 0.5 and could only be exceeded. It cannot express N ≪ chance, so it fires against a null in G for a reason that has nothing to do with G. Recorded as a defect in the branch, not as a result.
FB4 does not fire; its mirror image did, and was not registered. FB4 anticipated G > 0.85 and N > 0.85 — rivals identifiable as unpublished prose. What happened is N < 0.15: the published option identifiable as the odd one out. An unregistered outcome, named as unregistered.
FB3 does not fire. Computed before dispatch: length ratio G 0.934 / N 0.988 (Welch p = 0.323), Jaccard 0.491 / 0.413 (p = 0.125), longest common run 7.50 / 6.38 (p = 0.354). The Jaccard row still says the lead's Turgenev rivals are closer to Garnett than its Korolenko rivals are to Fell, so G is the harder discrimination of the two — a bias against a positive G, stated before the numbers existed. Within condition, Spearman(Jaccard, item accuracy) = +0.379 in G and +0.247 in N: similarity does not depress accuracy here, so the null in G is not an artefact of the items being too alike.
What the arm may now say, and it is less than it hoped.
- Licensed. At ten openings of Turgenev prose poems, three models that name the author at 10 of 10 cannot pick Garnett's wording out of a matched rival better than chance, on a task they perform at 18 of 18 on the King James Version. Measured, counterbalanced, verified.
- Licensed. A forced-choice "which is published" instrument returns a confident, unanimous, wrong answer when the published comparator translates freely. 46 of 48.
- Not licensed: that the Garnett-recall confound is refuted. Three reasons, and each is sufficient on its own. (i) The matched floor failed, so there is no measured value for what this instrument returns when the comparator is genuinely not held. (ii) Critic finding C6: P shows a verse reference where G shows a source passage, so a null in G with P passing leaves open that the source-shown shape is the obstacle rather than absent memory. (iii) Recognition and discrimination are different capacities from recall, and a model may hold Garnett and still fail to pick her out of a close rival at forty-word spans.
- The confound is therefore WEAKENED, not refuted, and
RS-20260728b-forced-run-ru§3's "live and unrefuted" becomes "live, and now measured against at one canonical pair, where the panel does not discriminate the comparator's wording." The six FORCED verdicts of S045 are not reopened and not cleared.
5. The translation limb, and two prospective contamination figures
939 Russian words, twenty spans, translated blind and frozen before any comparator was opened:
T-senilia-openings-R04-v1 (ten prose-poem openings Turgenev's lead translator had never touched)
and T-korolenko-spans-R04-v1 (ten spans from four Korolenko stories). Both were written as the
instrument's distractors and both are first-class artifacts with frozen logs.
Because the prose was frozen before the comparators were opened, the contamination measurements are prospective — the number did not exist when the words were chosen, which is the property S045 identified as the strongest kind of cell this arm can build.
| lead tokens | comparator | whole-set longest common run | the run | |
|---|---|---|---|---|
T-senilia-openings-R04-v1 |
621 | Garnett 1897, 66,047 tok | 10 | i looked round and saw a little bent old woman |
T-korolenko-spans-R04-v1 |
662 | Fell 1916, 73,371 tok | 11 | did not like it when he talked like that she would |
The non-canonical comparator gives the longer run. Against this project's measured range — Ovid 0, Turgenev/Garnett 21 — a canonical pair returns 10 and an obscure one returns 11, on prose written the same afternoon by the same translator under the same regime. Per span the Garnett runs are 4, 5, 6, 7, 7, 7, 7, 8, 9, 10 and the Fell runs 4, 5, 5, 5, 6, 6, 6, 7, 11.
This is a third independent line against reading run length as a contamination signal, after
S045's "6 of 8 forced on the frozen rule, 4 of 8 on a two-thirds rule" and its correction of the
project's "11–21 tokens at eight of eight loci" framing. Both new runs are almost entirely function
words. Neither artifact's contamination: suspected is revised — the front matter records the
declaration made before measuring, and the two-step is the record.
6. Instrument notes
- Position bias is real and the counterbalance earned its cost. Slot-1 rates in G: P2 0.90 and 0.70 across its two calls, P3 0.70 and 0.30, P1 0.50 and 0.50. A single-order design would have been reading P2's preference for slot 1.
- The within-day repeat is the other thing the two orders bought. G's flip rate is 33.3%, N's 8.3%, P's 0%. Note (bev) — S064's finding that a byte-identical same-day prompt moved 24.3% of one model's codes — now has a companion measurement on a two-alternative task, and the flip rate tracks how hard the item is rather than being a constant of the model: at ceiling it is zero, at chance it is a third.
- Note (b) fired, and this is its most expensive firing.
moonshotai/kimi-k3returnedfinish_reason: lengthwithcontent: nullatmax_tokens16,000, providerFireworks, $0.4149495 for nothing, thirteenth session. The rejected body was preserved rather than overwritten (the S045 defect, not repeated) and the declared reservedeepseek/deepseek-v4-prowas accepted first call at $0.0233. - P2's identification call truncated at the token cap and returned no parseable answers. Its figures are floors and nothing is inferred from them. The verifier asserts that exactly one ID call truncated and that no choice call did.
- The verifier itself found two defects in its own first pass — a cost-row check that counted derived files as response bodies, and a truncation check that would have failed the deliberately exempt ID call. Both were repaired in place before the 452-check pass.
7. What this cannot settle
- Nothing about the lead's own memory. The panel is a proxy for an independent reader, not for the lead. Inherited unchanged from S045 limit 4.
- n = 10 items, 3 models, one translator, one comparator in the arm's headline condition.
- Mechanism is untouched in every direction: a null in G is consistent with no memory, with memory that discrimination cannot reach, and with memory cancelled by whatever inverted N.
- P is the extreme of the canonicity range and a different item shape. It certifies that the format can express a hit. It does not calibrate sensitivity anywhere between the KJV and here.
- The two contamination figures are single measurements on single sets, and the comparison between 10 and 11 is one pair against one pair.
- Four amendments removed surface cues and one cue survived in the prose. No claim is made that the audit was exhaustive; critic finding C7 said so before the run and it is still true.
8. Verification
452 checks, 0 failures (runs/verification.json). The verifier re-reads every raw body,
reimplements the answer parser, the slot-assignment rule, the primary unit, the exact binomial, the
stratified permutation and the item assertions from the design text, imports nothing from
score.py, build_items.py or tools/, asserts every published span verbatim against the
comparator file, asserts that every rival is the frozen translation with punctuation and the five
declared renames set aside, asserts one cost row per stored body with no duplicates, and carries
three mutation tests.