Repository path: workshop/experiments/E-20260804i-name-or-prose/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260804i-name-or-prose |
| status | frozen |
| created | 2026-08-04 |
| updated | 2026-08-04 |
| senses | accuracy, naturalness |
| provisional | true |
| internal-judgment-only | true |
| links | wiki/arms/ARM-tierP.md, wiki/findings/results/RS-20260804c-peer-record.md, wiki/decisions/resolved/D-20260725-07-athenaeum-1906-condition-ii.md, workshop/experiments/E-20260804c-peer-record/design.md, workshop/translations/pevtsy/R15-crib-v1/translation.md, workshop/translations/pevtsy/R15-english-v1/translation.md, config/models.md, config/budget.md, wiki/goodness-senses.md |
E-20260804i — the name or the prose
ARM-tierP step 2, T3. Frozen before any of this session's translation was written and before
any dispatch. Nothing below may be revised after the first API call; amendments are appended with a
letter and a reason, as E-20260804c and E-20260804h did.
0. The subject-rule sentence, written before the unit was designed
What does this unit teach about translating literature or evaluating translations? — It tests whether the dimensional split a critic who read the Russian drew in 1904 between two English Turgenevs — Hapgood the more accurate, Garnett the better English — is a property of the prose readable at the places where the Russian actually forces that trade-off, or a property of the two translators' names. The object of study is a human critical record, the two translations it is about, and a third pair built this session to carry the same two properties under no name at all.
1. What RS-20260804c left, and what this run is for
S103 asked whether three non-Anthropic seats reproduce the ratified 1904/1906 record on six loci of
«Певцы». Result: naturalness came back at 30–6 to Garnett, accuracy at 19–17 with one seat
reversing — and all three seats named both translators correctly, so failure criterion F4 fired
and any reproduction is reported confounded. ARM-tierP step 2 therefore asks for materials on
which the seats do not identify the translators, or a statement of what Tier P can mean when the
reader always knows the author.
This run does both halves at once, on one set of sites:
- the historical pair — Garnett and Hapgood, which the seats demonstrably recognise — under two attribution conditions, one silent and one false; and
- a constructed pair the seats cannot recognise because it did not exist before this session: the same six loci rendered twice by the lead under two frozen purposes, one for a reader who will check the English against the Russian and one for a reader who will only ever read the English.
And it adds the variable S103 had no way to measure: where in the text the Russian forces the trade-off at all, classified by seats shown the Russian and no English whatever.
2. Materials, and what is inherited unchanged
| item | provenance |
|---|---|
| Russian «Певцы», six loci, 2,437 words | E-20260804c/materials/loci.json, unchanged |
Garnett 1895 (G) and Hapgood 1903 (H), same six loci |
same file, two-scan verified at S103, unchanged |
C — lead, purpose CRIB |
workshop/translations/pevtsy/R15-crib-v1/, this session |
E — lead, purpose ENGLISH |
workshop/translations/pevtsy/R15-english-v1/, this session |
| 12 sites | materials/sites-ru.json, extracted by the frozen positional rule in §3 |
Contamination, measured before this design existed and not as a diagnostic inside it. The gate
is E-20260804c/materials/gate-result.json, run at S103 on a held-out paragraph of this same story:
the lead reproduces Garnett at 12 contiguous tokens with one shared 12-gram and Hapgood at 9
with none. The lead is therefore Garnett-ward on this text, which is why no claim in this design
turns on the lead being independent of any published rendering (CLAUDE.md, standing rule). C
and E are compared with each other, never against G or H as an independent third opinion.
The measurement is repeated on the two new arms as control FC3.
The lead has read both published renderings (S103, and they are stored in this repo). That is why the lead's own fork map is a secondary here and the primary classifier is two seats that see the Russian and nothing else (§4).
3. The site rule, frozen before translating and blind to content
Sites were extracted mechanically, by materials/sites-ru.py, before a word of this session's
translation was written and before any site was read for what it contains:
Split each locus into sentences on
[.!?…]+ whitespace. Take the sentence at index 1 and the sentence at index ⌊n/2⌋. Two sites per locus, twelve in all.
No site was selected, dropped, lengthened or shortened for what it says. The rule can produce a
trivial site (L6-s15 is eight words) and a dense one (L1-s10 is the "Russian, truthful, ardent
soul" sentence); both stay. This is the answer to the objection an adversarial reviewer raised
against E-20260804h two hours ago in a different context — that sites chosen by the party with a
hypothesis are sites commissioned rather than found.
Each site is presented with its preceding and following sentence as context, in every language and arm, with the graded span marked. Uniform across sites; this is what makes an eight-word site gradeable.
4. FORK classification — the independent variable, and the lead does not set it
Primary classifier: two seats shown the Russian span in context and NO ENGLISH OF ANY KIND, and
asked the open form of the question the project has learned to ask (RS-20260804h §8; the closed
form failed twice):
Rendering this span into English, does a translator face a real trade-off between staying close to what the Russian says and how it says it, and writing English that reads well on its own? Describe what the trade-off is, or say that there is none here. Then answer FORK or NO-FORK.
Seats: A1 = P5 deepseek/deepseek-v4-pro, A2 = qwen/qwen3.7-max (reserve; the same two seats
that authored E-20260804h's source-only yardsticks, where they agreed 12 of 12). Neither sees any
English, so neither can be classifying from Garnett or Hapgood.
- A site is FORK if both seats answer FORK, NO-FORK if both answer NO-FORK.
- Sites where they disagree are excluded from the primaries and reported by name.
Secondary, and it is the wire to the translation limb: the lead's own map, which is mechanical
— a site is a lead-FORK where arms C and E make different choices at the graded span, and
lead-NO-FORK where their renderings differ in nothing a reader would call a decision. The lead map is
frozen in materials/lead-fork-map.md before the seats are dispatched and is never used to
classify; it is only compared. Prediction P4 is about that comparison.
5. Stages
Stage 0 — fork classification. 2 calls (A1, A2), max_tokens 6,000, 12 answer lines each.
Stage 1 — U4, the unattributed four-arm task. 3 seats (P1, P2, P3), one call each,
max_tokens 8,000. At each of the 12 sites the seat sees the Russian span in context and four
English arms under per-site shuffled keys — sha256(design id | seat | site) fixes the
permutation, written to runs/keymaps.json before dispatch. No attribution anywhere. For each
site the seat returns:
ACC:a ranking of the four keys onaccuracy— the propositional content, imagery and stated detail of the Russian carried over without unlicensed addition, omission or distortion;NAT:a ranking of the four keys onnaturalness, scored on the target text alone (the revised wording,D-20260802-13), against the period-idiomatic register anchor.
Stage 2 — M2, the misattributed two-arm task. 3 seats (P1, P2, P3), one call each,
max_tokens 6,000. The same 12 sites, only the two published arms, and each is given an
attribution that is false: the Garnett text is presented as "Isabel F. Hapgood, 1903" and the
Hapgood text as "Constance Garnett, 1895". Same two senses, same output format. No probe asks
the seat who wrote what: S103 already measured that these three seats identify both translators on
this text at 3 of 3, and importing that measurement keeps stage 1 blind and the two stages
symmetrical in everything but the attribution line.
Stages are dispatched in order. Judgment is never parallelized (charter §8.5).
6. The primaries
PR1 — localization, the primary. In stage 1, restricted to the published pair G vs H, the
record's dissociation — H above G on ACC and G above H on NAT at the same site —
occurs at FORK sites and not at NO-FORK sites.
Bar: (dissociation rate at FORK) − (dissociation rate at NO-FORK) ≥ 0.25, the margin
E-20260804hregistered and cleared at 0.500. Rates are over (site × seat) cells.
PR2 — the same split on a pair that has no name. In stage 1, restricted to C vs E: C above
E on ACC and E above C on NAT, at ≥ 8 of 12 sites pooled over seats, and the same
localization margin ≥ 0.25.
PR3 — the label effect. For each sense, the pooled direction of G vs H in stage 2
(misattributed) is the same as in stage 1 (unattributed).
PR3holds if both senses keep their direction. It fails if either sense reverses.
7. What each outcome licenses, registered before dispatch
| PR1 | PR2 | PR3 | what the arm may say |
|---|---|---|---|
| holds | holds | holds | The 1904 dissociation is located where the Russian forces the trade-off, appears on a pair carrying no name, and does not follow stated attribution. Tier P on an identifiable pair is confounded but not empty: what it can certify is a dimensional claim at choice sites, and the norm-familiarity confound of RS-20260804c §10.2 is still not excluded |
| holds | fails | holds | The dissociation is real and located, but is not reproducible by construction — two purposes do not manufacture the property two translators had. The record is about these translators, not about the dimensions |
| fails | — | — | The dissociation, such as S103 found of it, is not located at the sites where the source forces a choice; PR2 and PR3 are reported descriptively and step 2 closes on the negative |
| — | — | fails | Verdict follows the name. Tier P is not testable on any pair these seats can identify, and the arm closes saying so. This is the outcome that would end Tier P as a certification within reach of this project |
No motion will be opened from this run. D-20260804-16's ratifying condition 4 requires a
result→option map registered before dispatch; this design registers instead that its output is an
arm-closure statement and an entry in config/models.md's Tier P record, not an amendment to any
page's wording.
8. Failure criteria and controls, all registered here
- FC1 — classification power. Fewer than 4 FORK or fewer than 4 NO-FORK sites under unanimous seat classification → PR1 and PR2 are withheld, and the run reports only PR3.
- FC2 — seat degeneracy. A seat returning the identical ranking at ≥ 11 of 12 sites within a
sense is reported separately and excluded from that sense's pooled rate. (
RS-20260804c§8 found P2 collapsing senses post hoc; this registers the check the last run could not.) - FC3 — contamination of the constructed pair.
tools/dependence_check.pyonCandEagainst both published arms. IfE's longest common run with Garnett exceeds the 12 tokens S103 measured for the lead on this story, orC's with Hapgood exceeds it, PR2 is reported confounded — the constructed pair would then be partly a copy of the historical one. - FC4 — body integrity. A body with fewer than 12 answer lines per sense, or
finish_reason: length, is a seat failure, retried once to the same slug then to the declared reserve. - FC5 — the no-information benchmark (note (bht)). Spearman between site length in Russian words and the per-site dissociation count. If |ρ| ≥ 0.6, PR1 is reported confounded by length, because a longer span offers more chances for the two to diverge on either sense.
- FC6 — the global-quality artefact (the check
E-20260804h's critic forced in, kept here by the same reasoning). If in stage 1 armEoutranks armCon both senses pooled, orCoutranksEon both, the pair differs in overall quality rather than in dimension, and PR2's dissociation claim is not attributable to the two purposes. - FC7 — order. Per-site key shuffling is the order control; a seat whose per-key first-place distribution over the 12 sites departs from uniform by more than one site's worth in a single key is reported.
9. Predictions, written before dispatch
- PR1 holds. The record's split is a claim about how the two translators handled difficulty, so it should live where difficulty is.
- PR2 holds. Two purposes pointed at the two readers should reproduce the dimensional trade.
- PR3 holds — the labels will not move the verdicts, because
RS-20260804c§4 showed these seats' recognition and their memorisation come apart, and a false label conflicts with a recognition they already have. - P4, the risky one. The lead's mechanical fork map agrees with the seats' at ≥ 9 of 12
sites.
RS-20260804c§5's registered prediction from a translator's log failed and a no-information benchmark would have made it anyway; this is the same class of claim, stated again and stated sharper, and it may fail the same way.
10. Pre-flight cost, built from max_tokens with a ×2 routing margin
Note (abc) bounds the token count; note (bhq) says the price per token needs its own margin. Nine dispatches.
| stage | calls | max_tokens |
worst case |
|---|---|---|---|
| critic | 1 | 20,000 | $0.20 |
| 0 — fork | 2 | 6,000 | $0.07 |
| 1 — U4 | 3 | 8,000 | $0.37 |
| 2 — M2 | 3 | 6,000 | $0.22 |
| declared worst case | 9 | $0.86 |
Headroom at design freeze: $1.953751992 of the $5.00 UTC day, after this session's $0.327597600
ratification gate. The run fits; no stage is deferred. Reserves declared before dispatch (note
(bfc)): qwen/qwen3.7-max for any stage-1 or stage-2 seat, z-ai/glm-5.2 for the critic.
Amendment A1 — written after FC3 was computed and before any API call of this run
FC3 was run as §8 requires, on the two new arms against both published ones, over all six loci
(materials/fc3-result.json). It fires, on both of its registered legs, and one thing it shows
was not anticipated by the design at all. Both are recorded here, before dispatch, so that no seat's
verdict can influence how either is stated.
(a) The registered leg fires; PR2 is confounded, as pre-registered. R15-ENGLISH's longest
common run with Garnett is 16 tokens, and R15-CRIB's with Hapgood is 19, both above the 12
that S103 measured for the lead on this story. Per §8 FC3, PR2 is reported confounded and
that is not renegotiable after the fact. The constructed pair is not clean of the historical one.
(b) The unanticipated part, and it is a double dissociation in the overlap itself. The two arms do not merely overlap the published pair — each overlaps a different member of it, and in the direction its brief points:
| vs Garnett | vs Hapgood | Δ | |
|---|---|---|---|
R15-CRIB (accuracy-forward) |
run 16, 132 shared 7-grams | run 19, 176 | Hapgood-ward |
R15-ENGLISH (English-forward) |
run 16, 73 | run 15, 56 | Garnett-ward |
| (the published pair with each other) | run 14, 102 |
And it is a within-translator manipulation with a measured effect: S103 measured this same lead, unbriefed, on this same story, as Garnett-ward (run 12 against 9). The CRIB brief moves it to Hapgood-ward. Having read both cannot explain a brief-dependent direction; it predicts a constant bias, which is what S103 measured and what this is not.
How this is registered. It is post hoc with respect to FC3's data and is labelled so
wherever it appears. It is not a fourth primary and it does not become one whatever the seats
return. It is stated here, before dispatch, at the strength it can carry: on this text, under two
briefs, a translator's verbatim overlap with each of two published translators moves in the direction
of the dimension that translator's reception record credits. Every published run-length figure in
this project sits under the standing forced-run finding (CLAUDE.md; RS-20260728b, notes (bez),
(bgf), (bgi)), which is imported, not re-derived here: what is claimed is the contrast between
two arms of the same translator on the same source, where whatever the source forces is forced
equally on both.
One thing (b) does not do: it does not rescue PR2. PR2 asks whether seats read the
dimensional split off the constructed pair. FC3(b) is a string measurement with no reader in it.
Amendments A2–A6 — from the pre-run critic, all before any grading call
z-ai/glm-5.2, one pass, NEEDS-REDESIGN, six findings, three BLOCKING, all six accepted
(critic.md, $0.0619817352). Dispatched before every other call of this run.
A2 — from F1. The critic found that P4 holding in the most direct way (the seats reproducing
the lead's 9/3 map) would fire FC1 and withhold both primaries, and that §7 has no row for it.
FC1's threshold is NOT lowered — moving a bar after seeing which way the lead map leans is
exactly the move RS-20260802-tierD-verdict refused to make. §7 gains the missing row:
If
FC1fires,PR1andPR2are withheld, the run reportsPR3(which needs no classification), the pooled, unlocalizedGvHandCvEdissociation rates as descriptive,FC3(b), andP4.ARM-tierPstep 2 then closes on the recognition question alone, and records that a content-blind positional site rule did not produce a usable contrast of difficulty — which is a finding about the instrument the next design needs, and is stated as one.
A3 — from F2 and F5, the expensive fix and the one that can convict the design. R15-CRIB
leaves Russian common nouns in its English (rjadchik, grosh, the Wild Barin), which a seat can
read as fidelity and as bad English at once, without reading anything else. The briefs are frozen and
are not rewritten. Instead:
- The restricted set is the sites whose presented text — span or context, in any arm —
contains no untranslated Russian common noun. Computed mechanically before dispatch: 7 sites,
L1-s10 L2-s1 L2-s9 L3-s1 L3-s11 L4-s1 L5-s1. Excluded:L1-s1 L4-s8 L5-s5 L6-s1 L6-s15. PR2is judged on the restricted set, at ≥ 5 of 7 rather than ≥ 8 of 12.- Registered in advance: if
PR2holds on the full twelve and fails on the restricted seven, it is reported as an artefact of lexical foreignness and not as evidence that two purposes reproduce a dimensional trade. FC6's power is declared compromised by the same mechanism, and the restricted comparison replaces it as the real control onPR2.- The lead map has only 2 NO-FORK sites inside the restricted seven, so the restricted
PR2may have no localization arm at all; if the seats' map leaves fewer than 2 of either class there, restrictedPR2reports the pooled rate only, and says so. PR1also gains the restricted recomputation, reported beside the full figure, as F5 asks.
A4 — from F3. §7's first row is replaced with the critic's own statement of it: the dissociation
is located at FORK sites and survives false labels, but the constructed pair is contaminated
(FC3) and the seats recognise the published pair, so neither "no name" nor "no attribution" is
established. The phrase "a pair carrying no name" is withdrawn from this design.
A5 — from F4, accepted in full rather than merely registered. A four-arm presentation lets a seat
that recognises Garnett and Hapgood notice that C resembles Hapgood and E resembles Garnett, and
rank by analogy. The four-arm stage is abolished. Stage 1 becomes two two-arm blocks in separate
calls:
- Stage 1a —
GvH, 12 sites, per-site shuffled keys, unattributed. 3 seats. - Stage 1b —
CvE, 12 sites, per-site shuffled keys, unattributed,GandHabsent from the payload entirely. 3 seats.
Three extra dispatches. Every figure in PR1 now comes from a block in which the constructed pair is
not present, and every figure in PR2 from one in which the published pair is not.
A6 — from F6. Stage 2 gains a final line: independently of the attributions given above, name
the translator of each arm if you believe you can, otherwise UNKNOWN. This measures site-level
recognition, which S103 established only at locus level, and measures whether recognition or the
false label wins. It cannot break a blind: stage 2 is the condition in which attributions are stated,
and it is a separate call from stages 1a and 1b.
Revised pre-flight. Twelve dispatches, not nine.
| stage | calls | max_tokens |
worst case |
|---|---|---|---|
| critic | 1 | 20,000 | spent, $0.0619817352 |
| 0 — fork | 2 | 6,000 | $0.07 |
1a — G v H |
3 | 8,000 | $0.37 |
1b — C v E |
3 | 8,000 | $0.37 |
2 — M2 |
3 | 6,000 | $0.24 |
| declared worst case, remaining | 11 | $1.05 |
Against $1.891770257 of headroom after the critic. The run fits.