Repository path: workshop/experiments/E-20260803d-class-uniformity/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260803d-class-uniformity |
| status | frozen |
| created | 2026-08-03 |
| updated | 2026-08-03 |
| senses | consistency, cultural-mediation, style-correspondence |
| purpose | Evidence for the derivation of the goodness sense `consistency`; no translation is being evaluated for quality and no reader-facing edition is at stake. |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-typology-derivation.md, wiki/goodness-senses.md, wiki/findings/sense-dossier.md, workshop/translations/bezhin-lug/R04-v2/translation.md, config/models.md, config/budget.md |
E-20260803d — is the drift caused? A class of fourteen items, three translators, one text
Frozen before dispatch. ARM-typology-derivation step 5 (T2).
1. The question, and why it is not the one the project has been asking
wiki/goodness-senses.md defines consistency as: "Internal coherence across the whole text: names,
terminology, motifs, register do not drift without cause."
Against it the project holds a census: seven translators, five language pairs, five eras
(1857–2026), and not one of them uniform on a forked class — Shaw four ways over Buddhist toponyms,
Garnett two ways over two dogs in a sentence, Baudelaire three ways over three italicised foreignisms,
four translators four ways over lind-plega. wiki/findings/sense-dossier.md records the standing
reading: "the alternative reading — that consistency as written is a prescription the observed norm
contradicts — has kept getting more expensive to deny."
The census cannot settle that, because it has never looked at the two words that do the work. Every instance in it measures outcome — that a class came out non-uniform. The definition does not forbid non-uniformity. It forbids non-uniformity without cause. Nobody has asked whether the drift the census records is caused.
This run asks it, on a class large enough to answer with, in a text where three renderings exist.
Sentence test (wiki/tracks.md §subject rule). What does this unit teach about translating
literature or evaluating translations? — Whether holding one handling across a class of culture-bound
items is a norm translators observe, or a prescription this project wrote into its definition of "good"
and the practice contradicts. That is a question about what "good" means in translation, which is the
charter's central object; it is not a question about the project's instruments.
2. Materials
| source | Turgenev, «Бежин луг» (1851), the boys' night talk entire — 143 paragraphs, 3,825 Russian words, SHA-256 9033b289… (workshop/translations/bezhin-lug/R04-v2/source-ru.txt) |
| T1 | Constance Garnett, 1895, A Sportsman's Sketches (Project Gutenberg 8597), the matching section — 4,707 words |
| T2 | Isabel Hapgood, 1903, Memoirs of a Sportsman (archive.org memoirsofsportsm00turg, OCR), the matching section — 5,950 words |
| T3 | the lead, 2026-08-03, T-bezhin-lug-R04-v2, R04, log frozen and committed before this design was written |
| class | 14 items, 43 tokens, fixed from the Russian alone before translating — materials/items.json |
All three sources are public domain (1851 / 1895 / 1903).
3. The three declarations that constrain what may be concluded
(a) The published pair is measurably DEPENDENT on this stretch, and this is registered before the
run, not discovered after it. tools/dependence_check.py over the matched sections:
| shared 7-grams | 53 |
| shared 12-grams | 5 |
| shared 15-grams | 0 |
| longest common run | 14 tokens |
| verdict | DEPENDENT? |
The name-excluded counts the tool also prints are void on this pair and are not reported: OCR
running heads in the Hapgood scan make name_tokens() ban a, the, that, he, which is the
instrument artifact RS-20260725c-contamination-sweep recorded on this same scan. The five shared
12-grams reduce to two distinct runs, both in the narrator's voice.
Registered consequence. Convergence between Garnett and Hapgood is inflated by whatever dependence this is. Therefore: divergence between them is evidential and convergence is not. No conclusion in this run may rest on the two published translators agreeing.
(b) The lead is primed on this class and is excluded from any convergence test. Before translating,
a string-frequency probe printed counts (not sentences) for domovoy, nymph, sprite, water-,
wood-, rusalk, Trishka and eleven others over both comparators, and the dependence gate printed
the shared 14-token run verbatim. Both are declared in full in T-bezhin-lug-R04-v2 §Priming. The
lead's renderings are used here for one purpose only: as the sites at which its frozen log recorded a
reason. Reasons are not made independent or dependent by knowing another translator's lexical choices.
(c) Nothing here is a quality judgment. No rendering is scored better or worse than another. Tier D
is NOT PASSED; every evaluative sentence anywhere downstream of this run carries provisional: true
and internal-judgment-only.
4. The lead's own coding, frozen here before any seat is dispatched
Strategy vocabulary — the project's own set (wiki/goodness-senses.md, cultural-mediation), plus
descriptive paraphrase for the case none of the nine names:
copy-opaque · transparent calque · substitute · gloss · omit · scaffold ·
measure conversion · re-foreignisation · added realia · descriptive paraphrase
| item | Garnett 1895 | Hapgood 1903 | lead 2026 |
|---|---|---|---|
| I01 домовой | copy-opaque | copy-opaque | transparent calque |
| I02 русалка | copy-opaque | substitute (two different substitutes) | copy-opaque |
| I03 лесная нечисть | descriptive paraphrase | descriptive paraphrase | descriptive paraphrase |
| I04 Тришка | copy-opaque + gloss | copy-opaque | copy-opaque |
| I05 леший | transparent calque | transparent calque | transparent calque |
| I06 погань | substitute | substitute | substitute |
| I07 водяной | transparent calque | substitute | transparent calque |
| I08 крестная сила | substitute | transparent calque | transparent calque |
| I09 нечистое место | substitute | substitute | transparent calque |
| I10 разрыв-трава | descriptive paraphrase | substitute + gloss | transparent calque |
| I11 родительская суббота | substitute | transparent calque + gloss | transparent calque |
| I12 предвиденье небесное | transparent calque + gloss | transparent calque + gloss | transparent calque |
| I13 лесное зелье | descriptive paraphrase | substitute | descriptive paraphrase |
| I14 бучило | substitute | substitute | substitute |
This coding is one coder's and the project has a measurement of what that is worth — the same arm's
RS-20260802-voice-warrant §6 found an independent seat agreeing with the lead's sense-home assignment
on 9 of 24. That is why §5 dispatches three seats and why the primary figures are the seats',
not these.
5. Procedure
Stage 1 — strategy coding (3 blind seats, one call each)
Each seat receives all 14 items. For each item: the Russian word, a one-line gloss of the thing, the
Russian context sentence, and the three renderings labelled A / B / C in an order permuted per item
by sha256(item_id) and frozen in materials/order.json — no translator named, no century given, no
mention of Garnett, Hapgood, a lead agent, or this project. The seat assigns exactly one strategy from
the vocabulary above to each of A, B, C, and may append +gloss where a footnote is described.
The seat is not told what the run is about, that uniformity is at issue, or that any item is expected to differ from any other.
Stage 2 — the cause property (3 blind seats, one call each, no English shown at all)
The registered site-property, chosen before any coding because it is the one the class's dominant strategy turns on:
BLOCKS-CALQUE. Is a transparent English calque unavailable for this item? — i.e. can an English word or short compound built out of the Russian word's own literal sense denote this thing accurately enough that an English reader takes the right sense, without a footnote?
Each seat sees only Russian: the 14 class items plus 8 filler items — five ordinary nouns from
the same passage and three culture-bound positive calibrators, each with its expectation justified
individually in materials/items.json (amendments A3/A10) — shuffled
into one frozen order, each with gloss and context, and answers YES (a calque is unavailable /
blocked) or NO (a calque is available) with a one-clause reason. No English rendering of any item is
shown. The seats cannot see what any translator did, so this coding cannot be contaminated by the
outcome it is used to predict, and it cannot be inflated by the Garnett–Hapgood dependence.
Stage 3 — analysis (local, no API)
analysis/score.py, written and committed before dispatch, computes from the stored raw bodies:
- Uniformity. Per translator, the number of distinct strategies over the 14 items, and the modal strategy's share.
- Divergence. The number of items on which the two published translators' seat-majority strategies differ.
- Cause. Per translator:
P(non-modal strategy | BLOCKS-CALQUE majority = YES)minusP(non-modal strategy | NO). - The log check. For the lead only: agreement between the seats' BLOCKS-CALQUE majority and the site-local-reason column frozen in the translator's log before this design existed.
6. Predictions, registered
| prediction | threshold | what it would mean | |
|---|---|---|---|
| P1 | Neither published translator is uniform | ≥ 3 distinct strategies each over 14 items | positive control on the census; a failure here would say this class is unlike the seven the census rests on |
| P2 | Drift is caused: calque-availability predicts whether a calque is used — restated by amendment A12; see §10 | pooled published gap ≥ 0.40 | consistency's "without cause" clause survives — the census measured licensed drift |
| P3 | The two published translators diverge | ≥ 3 of 14 items differ | dependence-safe direction; a translator-borne component exists |
| P4 | The lead's frozen log agrees with the seats on where cause lies | ≥ 10 of 14 items | contemporaneous process evidence is recoverable by outside readers from the source alone |
P2 is the run's primary. If P2 fails for both published translators, the honest reading is that
non-uniformity in this class is not explained by the availability of a calque, and consistency's
escape clause is doing no work here.
7. Failure criteria, registered
- F1. If a seat-majority (2 of 3) does not exist on ≥ 12 of 14 class items in stage 2, the property is not a property: P2 and P4 return NO VERDICT.
- F2 — the sham and calibrator control. If two or more of the 8 fillers take a
YESmajority, the instrument fires where nothing blocks a calque, and P2 and P4 return NO VERDICT whatever the class items show. (Threshold raised from one to two by amendment A3.) - F5 — conditioning-cell minimum (amendment A6). Fewer than 4
YES-majority class items, or fewer than 4NO-majority class items → P2 and P4 return NO VERDICT. A gap computed off one or two items is one miskey wide. - F3. If stage 1 yields no seat-majority strategy on ≥ 36 of 42 cells, stage 1 is void; P1 and P3 fall back to the lead's frozen §4 coding and are reported explicitly as one coder's.
- F4. Runner guards, unchanged from
E-20260803c/call.py:finish_reason == "length"is a seat failure; a body with fewer answer lines than the payload has items is a seat failure; every attempt uniquely labelled (note (bgz)); raw bytes to disk before any parse (note (bdt)).
No threshold in §6 or §7 may be revised after any body is read.
8. The decision this feeds, with its default fixed now
The run's output goes to a decision page on consistency's wording, opened this session and
ratifiable only by a later one (charter §8). Options, with the provisional default fixed here,
before the numbers:
- A — no change (provisional default). The clause survives; the census measured licensed drift.
- B — amend the wording so that what is scored is drift the text does not motivate, with the measured base rate of within-class non-uniformity stated on the entry.
- C — narrow the sense to referential and terminological consistency (names, terms), excluding handling-strategy uniformity over culture-bound classes.
- D — retire
consistencyfor want of derivation. - E — split: within-class handling uniformity and referential consistency as two things.
9. Panel and pre-flight
Seats are non-Anthropic (config/models.md); the critic takes no part in either coding stage.
| stage | seats | max_tokens |
worst case |
|---|---|---|---|
| pre-run critic | P4 moonshotai/kimi-k3 |
16,000 | $0.26 |
| stage 1 strategy | P1, P3, P5 | 4,000 | $0.07 |
| stage 2 property | P1, P3, P5 | 3,000 | $0.05 |
second critic pass, if the first returns NEEDS-REDESIGN |
P4 | 16,000 | $0.26 |
| retry reserve (note (bfc)) | — | — | $0.20 |
| declared worst case | $0.84 |
Worst case is built from max_tokens and the published output price, not from an expected answer
length — note (abc). Today's headroom at session open: $3.407322239 of the $5.00 UTC-day cap, three
sessions already run. P2 google/gemini-3.6-flash is not seated in either coding stage: it billed
ten times P1 or P5 for the same payloads at S096 and returned an unusable body on 4 of 5 calls (notes
(bhq), (bhf)).
10. Amendments after freezing
A1 — the critic dispatch, after a seat failure that note (b) had already prescribed the fix for.
The first critic dispatch (critic_pass1, P4 moonshotai/kimi-k3, provider Fireworks, max_tokens
16,000, no reasoning cap) returned finish_reason: length, 16,547 reasoning tokens, zero
content characters, $0.260301 for nothing, in 223.2 s. Guard (b) in call.py rejected it correctly
and the raw body is preserved at runs/critic_pass1_kimi-k3_try1.raw.
This is note (b)'s twenty-fourth firing and the fourth on this same seat, and the note's own
standing amendment — "on a seat with a recorded history of this failure, the reasoning cap is not
advisory; effort: low is set on the first dispatch, not the retry" — was not applied. S094
applied it to this seat and got a clean verdict; this session inherited call.py from S096, which had
no reasoning parameter at all, and did not check. The cost is the operator's, not the seat's.
The amendment, made before any coding seat was dispatched:
call.py._post()anddispatch()takereasoning_effort, posted as{"reasoning": {"effort": …}}.- The critic is re-dispatched under the new tag
critic_pass1b— not the old one — so the rejected billed body cannot be overwritten (note (bgz)). effort: "low"on the first attempt, and a declared reserve,google/gemini-3.6-flash, which S074 established is the remedy that actually works on this seat when a raised cap does not. The reserve takes no part in either coding stage, so the design's independence condition holds.- Declared worst case revised from $0.84 to $1.10, and the $0.260301 already spent on nothing is inside that figure rather than excused from it.
Pass 1 of the pre-run critic then returned NEEDS-AMENDMENT, eight findings, four BLOCKING, and
all eight are accepted. Body: runs/critic_pass1b_kimi-k3_try1.txt (P4, provider Fireworks, stop,
76.0 s, $0.05136 — a fifth of the failed dispatch, for a complete verdict). Amendments A2–A7 below were
made before any coding seat was dispatched.
A2 — the glosses handed the seats the answer (critic finding 1, BLOCKING, and it was fatal).
Stage 2 claims to show no English. It showed something worse: the gloss of each item paraphrased the
Russian word's own morphology, so I05 read "the spirit of the forest", I07 "the spirit of the
water", I08 quoted the formula itself, and I10 said the plant "bursts locks and opens what is shut".
A seat asked whether a transparent calque is available for an item whose gloss is the transparent
calque is not judging calque-availability; it is reading the answer off the question. Every gloss is
rewritten to describe the referent — what the being is believed to do, where it is met — with no
paraphrase of the word's parts. materials/items.json carries the rule and the record of what was
changed.
A3 — the sham control was not a control for the failure mode that matters (finding 4). Five
trivial nouns test only false-YES. Three culture-bound positive calibrators are added — паперть
(the step outside a village church), полати (the boarded sleeping platform above a peasant stove),
староста (the peasant head of a village community) — items as culture-bound as the class but with a
plainly available English rendering. All eight fillers are expected NO. F2 now fires at two or more
filler YES-majorities, not one, so a single seat's slip on a single filler cannot void the primary.
A4 — the strategy list order is shuffled per seat (finding 5). transparent calque was second in
a fixed list in every payload. The list is now ordered by sha256(seat + strategy). The residual leak
the critic names — that showing one item with three renderings reveals a translation-comparison study
— is accepted and not fixed: removing it means padding the payload with renderings from other
passages, and the strategy coding is a description of visible English, not a judgment.
A5 — a cell may hold two strategies (finding 3, BLOCKING). One rendering in the set uses two
different handlings at different tokens of the same item. Forcing one label would have made that cell
wrong by construction and would have hidden the within-item drift, which is the very thing under study.
Seats may answer X / Y; score.py represents every cell as a set, counts both members toward a
translator's distinct-strategy total, and counts an item as divergent in P3 when the two sets differ.
Split cells are reported separately as split_items.
A6 — a minimum for each conditioning cell (finding 2, BLOCKING). P2 divides by cells that could be
empty. New failure criterion F5: fewer than 4 YES-majority class items or fewer than 4
NO-majority class items → P2 and P4 return NO VERDICT. With 14 items, a gap computed off one or two
YES items is one miskey wide, and the design said nothing about it.
A7 — P2 is restated as a pooled estimate with the dependence caveat attached (finding 6). The original threshold — "gap ≥ 0.40 for both published translators" — is a convergence claim, and §3(a) forbids resting anything on these two converging. P2 is now decided on the pooled published estimate (28 translator×item cells), with both individual gaps reported beside it and the interpretive sentence stating that this is one dependent pair, not two observations. The divergent items (P3) are the evidential core, as the critic asked.
A8 — the fixed order is called fixed, and its balance is checked (finding 8). The A/B/C assignment
is fixed by sha256(item_id), not randomised. Checked: no permutation takes more than 4 of 14
items, and no translator occupies one slot more than 7 of 14 times (garnett in slot B). Recorded
in materials/order.json.
11. The result → option map, registered (critic finding 7, BLOCKING for the decision)
A frozen default with no frozen decision rule lets the choice among B–E be made after the numbers. This table is registered before dispatch. Combinations not listed return "no recommendation", and the decision page then records that rather than improvising one.
| F1 / F2 / F5 | P2 (pooled) | P1 | option the decision page opens with as its lead |
|---|---|---|---|
| any fired | — | — | no recommendation on the clause; the page reports the instrument failure and the descriptive P1/P3 figures only |
| clear | MET | met | A — no change. Drift is caused; the census measured licensed drift |
| clear | MET | not met | A, with the note that this class is unlike the census's seven |
| clear | NOT MET | met | B — amend, and the entry states the measured base rate of within-class non-uniformity. C and E are recorded on the page as the live alternatives with what each would need |
| clear | NOT MET | not met | no recommendation; the class did not behave like the census's and the property did not predict |
D (retire) is not reachable from this run and the page will say so: one class in one text cannot
retire a sense, and consistency holds the strongest cross-corpus census on the list.
Registered, on critic pass 2 finding 6: P3 and P4 are reported as context and do NOT gate the option. If P4 fails while P2 passes, option A carries the caveat that the cause locus is not independently recoverable from the source by outside readers. If P3 fails while P2 passes, the causal reading rests on one pooled dependent pair, and the page states that 28 cells are not 28 observations.
12. Second critic pass, and the amendments it forced
Pass 2 on the amended design returned NEEDS-AMENDMENT, nine findings, three BLOCKING, all nine
accepted (runs/critic_pass2_kimi-k3_try1.txt, P4, stop, 106.0 s, $0.11721).
A12 — the primary was wrong-signed by construction (finding 1, BLOCKING, and this is the one that
mattered). P2 measured departure from the translator's modal strategy conditioned on
BLOCKS-CALQUE. That only detects cause when the modal strategy is the calque. On the frozen §4
coding Garnett's modal is substitute, not calque — so for Garnett, blocked items are exactly
where the modal substitute is used, the gap goes to zero or negative, and the run would have
reported "not caused" for the translator the census is actually about. P2 is restated as the direct,
modality-free question: does calque-availability predict whether a calque is used?
P2a (primary):
P(calque | calque available) − P(calque | calque blocked), pooled over the two published translators, threshold ≥ 0.40.
The old modal-departure measure is retained and reported as P2b, context only. score.py computes
both and records P2_decided_on: P2a.
A9 — the slash convention collided with the translators' own punctuation (finding 2, BLOCKING).
a deep pit / pool and water-nymph / water-sprite look identical in the payload and are different
things. Genuine two-occurrence splits are now marked with ‖ (U+2016), the convention is stated in
the stage-1 prompt, and the parser splits on it. The lead's I14 is corrected to a single rendering: its
head noun does not change.
A10 — полати was a mis-registered calibrator (finding 3, BLOCKING). It was added as a
culture-bound item with a "plainly available" English rendering; it has no uncontested English
compound, so two seats answering YES — defensibly — could have voided the primary on a wrong
expectation rather than on an instrument failure. Withdrawn and replaced by приказчик. Every
filler's expectation is now justified individually in materials/items.json, not asserted as a batch.
A11 — all ten strategy labels are defined in the payload (finding 4), and two glosses lose their
markedness cues (finding 7). Five labels were listed and never defined, which would have compressed
the coding onto the five that were. And A2's claim to have removed the gloss leak is corrected here
to "reduced, not removed": the referent glosses still carry the calque's spatial half (among trees,
beneath a river, a particular building), which a seat that can read the Russian word can join up.
What is removed is the markedness cue — I12 no longer says the word is non-standard and footnoted,
I04 no longer says the name is a diminutive — because both invited YES on grounds other than
calque-availability.
A13 — P4's rubric is frozen, and P4 is demoted (finding 8). P4 compares the seats' BLOCKS-CALQUE
majority against the site-local-reason column exactly as frozen in the translator's log, machine-read
from LOG_REASON in score.py, with no re-reading of the log's prose and no discretion after the
bodies are opened. The two are not the same construct — the log asks was there a site-local
reason, the seats are asked is a calque blocked — so P4 is a convergence check between two related
questions, not a validation, and per §11 it does not gate the option.
Accepted risk, recorded rather than fixed (finding 9). Fourteen of the twenty-two stage-2 items are folk-belief vocabulary and eight are mundane, so a seat can see which class is under study. This leaks the class, not the direction of the hypothesis. No interpretive sentence may claim the seats were naive to the class's specialness.
No third pass. Two passes is this project's practice, and pass 2's own closing judgment is that the three blocking findings are fixable without unfreezing anything that matters. They were fixed before any coding seat was dispatched.