Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260730f-recall-floor/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260730f-recall-floor
statusfrozen
created2026-07-30
updated2026-07-30
senses—
internal-judgment-onlytrue
linkswiki/arms/ARM-forced-defence.md, wiki/findings/results/RS-20260728b-forced-run-ru.md, wiki/findings/results/RS-20260728-forced-run.md, workshop/translations/senilia-openings/R04-v1/translation.md, workshop/translations/korolenko-spans/R04-v1/translation.md, config/models.md, CLAUDE.md

The recall control with a floor at chance, which ARM-forced-defence may not proceed without

Design frozen before dispatch. S065. ARM-forced-defence step 2.

0. What is owed, in the words of the page that owes it

ARM-forced-defence, absorbing a backlog row that hit the review-or-retire age at S056:

Step 2 may not report a FORCED verdict on canonical material until this control exists, and what it needs is forced choice against a distractor, where a model that knows nothing scores 50% and cannot decline.

The reason is RS-20260728b-forced-run-ru §3. Asked to write out Constance Garnett's published English as they recalled it, all three panel models returned UNKNOWN at 24 of 24 cells. Every recall score was 0, which is a refusal floor and not a measurement, and the same three models, asked instead to identify the passages, named Turgenev for 8, 7 and 8 of 8 and the individual prose poem for 6, 4 and 5 of 8 from three lines each. So the Garnett-recall confound is live and unrefuted, and the six FORCED verdicts that session recorded — including the one supporting CLAUDE.md's standing rule that the lead cannot be the independent third translator for a work with a canonical English translation — each keep an alternative explanation nobody has removed.

This design removes the floor. Two options, no third; a model that knows nothing is at 50% and cannot retreat to UNKNOWN.

1. The three conditions, and why the third exists

Every item is the same shape: a Russian passage, and two English renderings of it, labelled 1 and 2. One is the published translation. The other was written by the lead from the Russian alone, frozen at ff64960 before any comparator was opened. The question is always which is the published one.

condition published side canonicity
G 10 items Constance Garnett, Dream Tales and Prose Poems, 1897 the default English Turgenev; on Gutenberg; named by CLAUDE.md's standing rule
N 10 items Marian Fell, Makar's Dream, and Other Stories, 1916 the only freely reachable English of any of these four Korolenko works; nobody's default
P 6 items the King James Version the most reproduced English translation there is

N is the empirical floor. A model with no memory of the comparator is not necessarily at 50% — it may prefer one option on register, on fluency, or because the rival was written by an author it resembles. That bias, whatever it is, is present in N exactly as it is in G: same task, same instruction, same distractor author, same language pair, same period, same regime, same session. What N supplies is the score this instrument returns when the comparator is not held, measured rather than assumed. The primary result is the difference G − N, not G against 0.5.

P is the positive control, and it is the part of this design S064 makes mandatory. If G and N both come back at chance, there are two readings — these models do not hold either comparator, and this task cannot be done by anybody. P separates them. KJV wording is held verbatim by every model that has read English text at scale, so a model that cannot pick the KJV out of a matched pastiche cannot do this task at all, and a null in G and N would be uninterpretable.

P's shape differs and the difference is stated now rather than defended later. P shows a chapter-and-verse reference instead of a source passage, because the KJV's source is Hebrew and Greek and showing it would change what the task measures. P therefore certifies the response format and the forced choice; it does not certify the source-shown shape. It is also the extreme of the canonicity range and says nothing about sensitivity in the middle of it.

2. Materials

Spans. Selected by materials/select_spans.py, committed at 9682fdb before any span was seen. Turgenev: poems in published order, skipping every poem the lead has already translated and every poem under 120 Russian words, opening paragraphs until the span reaches 35 words. Korolenko: fixed offsets in four stories, first paragraph at or after the offset with ≥35 words. Two source repairs, both declared on the translation pages (finalise_spans.py).

Rivals. T-senilia-openings-R04-v1 and T-korolenko-spans-R04-v1, 939 Russian words, frozen with their translator's logs at ff64960. The rule the translator worked under is on those pages and is repeated here because it is a design commitment: translate each span exactly as R04 would have it translated if there were no experiment. A distractor made worse is a rigged control; a distractor made Garnett-like is an invented agreement.

P's rivals are pastiches written for this design in matched register, in materials/kjv.json, against KJV text extracted verbatim from Gutenberg #10.

Turgenev spans are openings on purpose. An opening is the most reproduced part of a short work. Korolenko's spans sit deeper in their works and are therefore, if anything, less recallable. That biases G − N upward, i.e. toward confirming the confound and leaving the arm blocked — which is the conservative direction, and it is chosen deliberately.

2b. Amendments made after §§0–7 were frozen and before any dispatch, labelled as such

A1 — two N spans are dropped, and the reason is in the comparator, not in the result. Fell 1916 is a freer translator than Garnett, and at two of the ten Korolenko spans she restructures the passage so completely that no stretch of her English corresponds to the Russian: the Khapun exchange (N09) and the matchmaking dialogue (N10). An item cannot be built where there is no counterpart. Condition N therefore has 8 items and 24 primary units, not 10 and 30. Both spans remain in T-korolenko-spans-R04-v1, which is a translation and not an item list.

A1 has a consequence the design must own: requiring alignment selects, inside Fell, for the passages she rendered closely. G is Garnett throughout, who is close throughout; N is Fell where Fell is close. If free rendering were itself a cue — the option that departs from the Russian is the published one — the selection removes it from N and there was never any to remove from G. The direction is toward comparability; the selection is still a selection and is reported.

A2 — two published spans carry one added closing quotation mark each (N05, N08), because Fell's quotation continues past the span boundary and an unbalanced quotation mark is a typographic cue with nothing to do with memory. Declared in materials/build_items.py; the verbatim assertion strips the added mark before comparing, so it is still an assertion against the real text.

A3 — the FB3 covariates were computed before dispatch, because they are properties of the materials and not of the outcome (analysis/covariates.py, runs/covariates.json).

G (10) N (8) Welch t p
length ratio published ÷ rival 0.934 0.988 −1.021 0.323
token Jaccard, published ~ rival 0.485 0.400 2.011 0.066
longest common run, tokens 7.00 6.13 0.905 0.382

FB3 does not fire. But the Jaccard row is close to firing and it points the way that matters: the lead's Turgenev rivals are more similar to Garnett than its Korolenko rivals are to Fell, so G's items are the harder discrimination of the two. That is a bias against finding G above N — against confirming the confound, and therefore toward FB1. It is a limit on the null reading and it is stated here rather than after the numbers. Added to the analysis: the Spearman correlation between item Jaccard and item accuracy, pooled over G and N, so that the size of this cue is reported rather than argued about.

A5 — the design would have measured punctuation, and it was caught by reading the built prompt rather than the design. Gutenberg's Garnett prints ellipses as three ASCII dots and speech in single curly quotes; Fell prints double curly quotes; the lead's own prose uses the ellipsis character and, in one poem, guillemets. A model could have picked the published rendering off the punctuation alone with no memory whatever, and the first built prompt shows it doing so would have been trivial — Windlessness, warmth ... air like new milk! against warm weather… the air is. typographic_neutralise() now rewrites every option in every condition: ellipsis character to three ASCII dots at a fixed spacing, every quotation mark to the straight double mark, every dash form to a spaced em dash, whitespace collapsed. It touches no word. assert_no_typographic_cue() asserts afterwards that no banned mark survives anywhere, and the assertion runs again in analysis/verify.py.

A6 — the same problem one level up: dialogue marking. R04's standing typographic ruling marks spoken dialogue with a leading dash, the Russian convention; both comparators use quotation marks. That is a regime decision rather than a wording decision, and left standing it would let a model pick the published option off the dialogue marking at four of the eight N items — inflating N, shrinking G − N, and pushing the result toward FB1, the branch that unblocks the arm. The four rivals with spoken dialogue are therefore re-marked with quotation marks for the item only; the translation artifacts keep the dash exactly as frozen. assert_punctuation_only() asserts that each override differs from the frozen rival in punctuation alone — letters stripped, the two strings are identical — so no wording can hide inside the normalisation.

A5 and A6 are the two amendments this session made without a critic telling it to, and both were found by reading the dispatch text. Recorded that way rather than folded into §2 silently.

A4 — pre-flight estimate, note (abc), built from max_tokens and not from an assumed output. Worst case per choice call at list out-price: P1 $0.035, P2 $0.036, P3 $0.032 on a 4,000-token cap; 18 calls ≈ $0.65. Critic at max_tokens 16,000 on P4 ≈ $0.26. Declared worst case $0.96, against $3.27 of headroom on the UTC day. Routing can multiply a per-call figure by ~4 (S022), so the declared figure is a list-price worst case and not a bound.

2c. Amendments from the independent pre-run critic pass (before dispatch)

Critic: deepseek/deepseek-v4-pro (P5), the declared reserve. P1, P2 and P3 are subjects and may not critique the design; P4 moonshotai/kimi-k3 was called first and returned finish_reason: length with content: null at max_tokens 16,000, provider Fireworks, $0.4149495 for nothing — standing note (b), thirteenth session, and the single most expensive firing of it in this ledger. The rejected body is preserved at critic.attempt0.raw and was not overwritten; the reserve's body is critic.attempt1.raw. Verdict NEEDS-AMENDMENT, seven findings — two BLOCKING, two MANDATORY, three ADVISORY. All seven accepted. Note (rr), twenty-third consecutive session in which a critic changed a design.

3. Panel, calls, and the counterbalance

Three models, one lab each, temperature 0, max_tokens 4000. Slugs resolved from config/models.md at run time: P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5 — the same three as S044 and S045, deliberately, so the three results are comparable.

calls what
6 G, three models × two orders
6 N, three models × two orders
3 P, three models × one order
3 ID, three models, one call each: the 20 Russian spans, unlabelled, name the author and the work
18 total dispatches, plus one independent pre-run critic call

Order is counterbalanced because this project has measured slot preference at 0.50–0.85 and has never found it absent. Order a puts the published rendering in slot 1 for odd-numbered items and slot 2 for even; order b is the exact swap. Every item is therefore seen by each model twice, once in each arrangement, and the pair doubles as a within-day repeat control — which note (bev) made non-optional yesterday, when a byte-identical same-day prompt moved 24.3% of one model's codes on a classification task.

ID exists so that a null in N is readable. If N comes back at chance, the first objection is that the models do not recognise Korolenko at all, and that objection is answerable only by measurement. It also replicates S045's condition E on Turgenev — 8/8, 7/8, 8/8 — which nobody has repeated. ID is dispatched after G, N and P are complete, so that naming the authors cannot cue the recall task; the API is stateless and no ordering effect is possible, but the ordering is kept anyway.

No model is asked to judge its own output, and nothing here is a quality judgment: the question is a matter of textual fact, which is the register in which S015 found this panel strong.

4. Scoring, frozen

The datum for each (item, model, order) is the chosen slot, mapped to correct / incorrect. Refusals and unparseable answers are scored NOANSWER and reported separately; the instruction forbids them and there is no UNKNOWN option.

The primary unit is the (item, model) pair, scored across its two orders:

So G and N each have n = 30 primary units and a chance value of 0.5. P, run in one order only, has 18 (item, model) judgments and is scored as a plain proportion.

Reported alongside, all pre-registered:

Statistical rule, fixed now. G − N is tested by a two-sided exact permutation test over the 60 primary units (10,000 permutations, seed 20260730, labels shuffled between conditions). G against 0.5 and N against 0.5 are each tested by a two-sided exact binomial on the counts of 1.0 and 0.0 units, with 0.5-units split evenly and rounded down in the direction that makes the test less significant. α = 0.05, and there are three tests, so the Holm correction is applied and the adjusted values are what the verdicts are read off.

5. Pre-registered predictions

Written before any comparator was opened, before any item was built, and before any call.

6. Failure criteria, and they bind pages outside this one

FB1 — the branch that unblocks the arm and reopens six verdicts. If P ≥ 15/18 (the instrument works) and G is not above chance after Holm correction and G − N does not reject, then these models do not discriminate Garnett's wording from a matched rival at these ten loci. The Garnett-recall confound is then refuted for this corpus, RS-20260728b-forced-run-ru §3's "live and unrefuted" must be restated, §5's FB2 stops being Untestable and becomes did not fire, and the six FORCED verdicts of that session become readable as source constraint. CLAUDE.md's standing contamination paragraph and ARM-forced-defence are amended in the same session, not a later one.

FB2 — the branch that costs the instrument. If P < 12/18, the forced-choice format does not work on these models and nothing in G or N may be read at all — not as evidence of memory, not as evidence of its absence. The session reports a void and the arm stays blocked. A positive control that is allowed to fail quietly is not a positive control, which is S064's lesson at the cost of a whole session's headline.

FB3 — the branch that costs the comparison rather than the instrument. If the length-cue or discriminability distributions differ between G and N at a two-sided Welch p < 0.05, then N is not a matched floor for G and G − N is reported as confounded, whatever it comes to. The individual condition results survive; the contrast does not.

FB4 — the branch that says the rivals were badly made. If G > 0.85 and N > 0.85, the models are picking the published rendering everywhere, which on this design most likely means the lead's rivals are identifiable as non-published prose rather than that Fell is held in memory. That is a defect in the materials, it is this session's finding, and no memory claim follows from it.

These four are written before dispatch and they bind. FB1 and FB4 point in opposite directions and cannot both fire; FB2 voids FB1, FB3 and FB4 alike.

7. What this cannot settle, listed before it can be forgotten

  1. Nothing about the lead's own memory. The panel is a proxy for an independent reader who does not hold the comparator, not for the lead. This project cannot measure what its own lead agent holds. RS-20260728b-forced-run-ru limit 4, inherited unchanged.
  2. Recognition is not recall and neither is discrimination. A model may hold Garnett's cadence without holding her words, and pick correctly for that reason. A positive G is consistent with several mechanisms and identifies none.
  3. The rivals are the lead's, and the lead has read a great deal of nineteenth-century English translation. If the lead's prose is systematically Garnett-flavoured, G is depressed — the conservative direction for FB1 is the opposite one, so this is a limit on the positive reading, not on the null.
  4. n = 10 items per condition, three models, one language pair, two comparators. Nothing here says how the forcing probe behaves on Chinese, Italian or French material, where this project also has measured figures.
  5. A refuted confound does not clear a run. It removes one alternative explanation of the S045 FORCED verdicts. Those verdicts still only separate source-forced from not source-forced, and the panel's convergence on a shared prior remains untouched (A1, S045).
  6. P is the extreme of the canonicity range. The natural next measurement — Garnett's own most read book, against these same rivals' method — is not in this design and is named as step 3.

8. Provenance