Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260808b-discordant-marking/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260808b-discordant-marking
statusfrozen
created2026-08-08
updated2026-08-08
sensesstyle-correspondence, voice, cultural-mediation, accuracy
provisionaltrue
internal-judgment-onlyfalse
linkswiki/arms/ARM-marking-work.md, framework/v0.1/README.md, wiki/findings/results/RS-20260807d-marking-work.md, wiki/findings/results/RS-20260806e-published-loss.md, wiki/findings/results/RS-20260806f-persona-crossing.md, workshop/translations/tolstyi-i-tonkii/R04-v1/translation.md, workshop/translations/smert-chinovnika/R04-v1/translation.md, workshop/regimes/R20-r1-forced.md, config/models.md, config/budget.md

E-20260808b — where a deference form fights its own content, does a published English rendering lose the standing?

ARM-marking-work step 2 (T5). Frozen 2026-08-08 UTC, session S133, before any API call of any kind. The two lead renderings and their registered predictions were committed at b5af03d before this file was written.


1. Why this run exists, and why step 2 could not be written without it

ARM-marking-work step 2 is scoped as a writing step: state prediction 1′ on the admission condition step 1 measured, and write it into a framework/v0.2. Two things block writing it.

(a) The admission condition has one measurement, on one text, in one language. RS-20260807d established the population prediction 1 needs — utterances where the source marks the speaker–addressee relation and the content pulls against the standing the form claims — at 4 of 51, Fleiss κ = 0.786, in Ōgai's courtroom. A 1′ written on a base rate from a single Japanese text would be the fifth version of prediction 1 written on evidence too thin to carry it.

(b) The procedure a 1′ would name is known to be broken for this material. RS-20260807d §4.3: the E-20260802e yardstick — a written statement of the relation, which blind readers match the English against — leaks at 9 of 12 statements on dialogue, and structurally, because a relation between two speakers is enacted through the speech act, so any statement of it entails the content. Address marking lives in dialogue. Writing 1′ on a procedure that cannot be run on the material 1′ is about would repeat the mistake four attempts have already made.

This run supplies both. Its yardstick is numeric profile crossing, not statement matching: readers of the source and readers of the English answer the same question on the same scale, and the statistic is the discrepancy between them. That instrument shape is not new here — it is RS-20260806f / RS-20260807's narrator crossing, which fired 16 of 16 twice — and it cannot leak, because nothing is written down for anyone to match against.

One sentence on what this teaches about translating literature (continue-prompt.md §4.5): it measures where in a story a grammatical mark of social standing carries something the words do not, and whether a published English translator loses it there.

2. Materials

Two complete stories by Chekhov, both 1883, both public domain, both read whole in Russian.

«Толстый и тонкий» «Смерть чиновника»
copy-text ПСС 1975 via ru.wikisource ФЭБ / ПСС via ru.wikisource
Russian words 540 713
spoken utterances 10 16

Why these. Both stories are about the phenomenon. In the first, the thin man switches from ты to вы, to «ваше превосходительство» and to the словоерс (the deference clitic -с) in mid-conversation, while his content stays childhood-friendship. In the second, Chervyakov's forms are maximally deferential and his content is a man refusing to stop pestering a general. The corpus is deliberately enriched: RS-20260807d found the phenomenon at ~8% of marked utterances, and a design wanting power cannot sample randomly. Enrichment is declared, and it bounds the base rate this run may report (§8 limit 1).

The site inventory

26 spoken utterances, attributions stripped, hand-built from the copy-texts and the frozen renderings before any seat was called: materials/sites.json. Interior monologue is excluded — it has no addressee. Nothing was selected: every utterance in both stories is in.

Arms

arm what it is provenance
SRC the Russian utterance copy-text
GARNETT Constance Garnett 1922, The Tales of Chekhov vol. 13 (Gutenberg #13414) published, independent, evidence class X1a
IND-PLAIN an independent hand renders both stories whole, no brief beyond "translate" generated, stage 2
IND-FORCED the same hand, both stories whole, under the R20 brief quoted verbatim from framework/v0.1 §2 generated, stage 2
LEAD-CLOSE T-tolstyi-i-tonkii-R04-v1 / T-smert-chinovnika-R04-v1 descriptive only, excluded from every primary by the rule in §6
WRONG 6 sites, direction-reversed by the frozen table in materials/wrong-table.md positive control
REPEAT 6 GARNETT items duplicated into a different block precision control

IND-PLAIN exists because IND-FORCED alone would confound the brief with the hand. The contrast that answers prediction 1 is IND-FORCED − IND-PLAIN within one hand.

The contamination gate changed this design before anything was dispatched

Measured after the two lead renderings were frozen and committed, with tools/dependence_check.py:

cell shared 7-grams 12 15 longest run
lead R04 ~ Garnett, «Толстый и тонкий», whole 59 16 3 16
lead R04 ~ Garnett, «Смерть чиновника», whole 105 33 20 27
two independent published hands (Garnett 1920 ~ Koteliansky & Murry 1915 on «Пари», length-matched) 38–54 7 2 16
lead R04 ~ Garnett, dialogue only, «Толстый и тонкий» 19 0 0 11 — clean
lead R04 ~ Garnett, dialogue only, «Смерть чиновника» 39 6 2 16 — dependent

The lead's close rendering is not independent of Garnett's, and the 16-token dialogue run in «Смерть чиновника» falls inside SC13, one of the five sites the lead's own frozen log registered as discordant. Under the standing rule (CLAUDE.md, and note (bhb)) the lead is therefore not an arm this run's validity may rest on: LEAD-CLOSE is rated and reported and is excluded from P1, P2 and P4 by a rule written here, before dispatch. The close arm the primaries are read on is Garnett, a published hand, in evidence class X1a.

3. Questions

  1. Does the discordance census transfer? Is the concordant/discordant distinction that RS-20260807d measured at κ 0.786 in Japanese reproducible at agreement in Russian, by readers shown only the source?
  2. Is discordance where a close rendering loses the standing? This is the admission condition 1′ needs, and no run has tested it.
  3. Does the R1 move recover it there? This is prediction 1, on an admission condition defined on the source and the target rather than on a reader's recovery.

4. Procedure

Stages run in order. Every raw body is written to runs/ before anything is parsed.

Stage 0 — pre-run critic. One non-panel call (nvidia/nemotron-3-ultra-550b-a55b), given this design whole. BLOCKING findings are answered before Stage 1 dispatches; every finding is recorded in critic.md with accept/overrule and a reason.

Stage 1 — the source census. Three seats — P1, P3, P5 — each shown the 26 Russian utterances, in a fixed shuffled order, with no English of any kind, no speaker labels, no story title, no author. Per utterance, three answers:

Stage 2 — generation. One hand, P4 (moonshotai/kimi-k3), which is a seat in neither stage. Two dispatches per story: IND-PLAIN (translate the numbered paragraphs into English) and IND-FORCED (the same, plus the R20 brief quoted verbatim). The hand is told nothing about the study, nothing about discordance, and is given no site list. Paragraph numbering is returned as given and checked mechanically; a body whose numbering does not round-trip is discarded and re-dispatched.

Stage 3 — the English rating. Three seats — P2, QWEN (qwen/qwen3.7-max), GLM (z-ai/glm-5.2), disjoint from Stage 1 — each shown 116 English items in three blocks, in a fixed shuffled order that mixes arms and works, with no Russian, no arm label, no speaker label, no story. Per item, two answers:

Stage 4 — recognition, reported and non-gating. One non-seat call shown ten GARNETT items and asked to name work and author. Reported; it does not gate, and §8 limit 4 states why recognition is conservative for P1.

Order within every block is fixed by sha256(design-id | item-id), computed in build_blocks.py and committed before dispatch.

5. The statistic

For each site i and arm a:

A site is DISCORDANT iff at least 2 of the 3 Stage 1 seats return fit: discordant; CONCORDANT otherwise. A site is MARKED iff at least 2 of 3 return mark: yes. Only MARKED sites enter any primary.

6. Predictions, registered

statement bar
P1 (primary) For GARNETT, mean D at DISCORDANT sites exceeds mean D at CONCORDANT sites ≥ 0.75 scale points and exact permutation P ≤ 0.05 over all relabelings of MARKED sites into the two observed class sizes
P2 At DISCORDANT sites, IND-FORCED has lower mean D than IND-PLAIN ≥ 0.75
P3 The lead's frozen log (L10 in both artifacts, 7 of 26 sites) matches the seats' DISCORDANT set ≥ 5 of 7, with the exact hypergeometric P reported beside it
P4 At CONCORDANT sites, IND-FORCED does not beat IND-PLAIN |Δ| < 0.50
P5 (descriptive) The carrier: wording rate is higher for IND-FORCED than for GARNETT direction only
P6 (descriptive) Fleiss κ on the Stage 1 concordant/discordant call is comparable to RS-20260807d's 0.786 reported, no bar

LEAD-CLOSE appears in P5 and in every descriptive table and in no primary (§2).

What the lead expects and is putting on the record so it can be wrong: P1 holds, P2 holds, P3 fails at 4 of 7 (the S128 translator over-predicted his own source by a factor of two and this one will too), P4 holds, P5 holds.

7. Failure criteria, which withhold rather than adjust

fires when consequence
F1 fewer than 6 MARKED sites are DISCORDANT P1 and P2 withheld; the census is still reported
F2 mean D(WRONG) does not exceed mean D(GARNETT) on the same six sites by ≥ 1.00 the English instrument is not shown to respond to standing; all primaries withheld
F3 mean absolute difference between the two presentations of a REPEAT pair exceeds 0.75 precision insufficient; all primaries withheld
F4 Stage 1 between-seat SD of standing at DISCORDANT sites exceeds that at CONCORDANT sites by > 0.75 P1 is confounded with source-side ambiguity and is reported descriptively only
F5 fewer than 90% of cells return all primaries withheld
F6 fewer than 18 of 26 sites are MARKED by 2 of 3 seats the corpus is not what the design claims; all primaries withheld

No bar in §6 or §7 is moved after it fires. Any amendment is written into critic.md or an Amendments section with a timestamp, and amendments after a stage has returned are marked as such.

8. Limits, written before the run

  1. The corpus is deliberately enriched. Any discordance rate here is a rate in two stories chosen because they are about deference, and is not a base rate for Russian prose. The base rate on record stays RS-20260807d's 4 of 51.
  2. Two stories, one author, one year, one pair. Chekhov 1883, RU→EN.
  3. One published close hand. Garnett is the only published English of these two stories the project can reach; there is no second published arm and therefore no published-pair contrast of the kind RS-20260806e had.
  4. Recognition is uncontrolled and is conservative for P1. Both stories are famous. A seat that recognises «Смерть чиновника» knows Chervyakov is a clerk addressing a general, and would therefore rate standing closer to the source at exactly the discordant sites — which shrinks the effect P1 predicts. Stage 4 measures how much recognition there is.
  5. The lead wrote the WRONG inversions knowing the hypothesis. They are a positive control on the instrument, not an arm under test, and they are frozen in materials/wrong-table.md before dispatch.
  6. IND-PLAIN and IND-FORCED are one hand. Whatever that hand's register defaults are, they cancel in the within-hand contrast and do not cancel in any comparison against Garnett.
  7. Tier D is NOT PASSED. No seat ranks anything and nothing here is a quality judgment; every sentence of the result is provisional.

9. Spend

Pre-flight worst case built from max_tokens, per note (abc) — not from expected output length.

stage calls model max_tokens worst case
0 critic 1 nemotron-3-ultra (non-panel) 16,000 $0.06
1 census 3 P1, P3, P5 6,000 $0.13
2 generation 4 P4 8,000 $0.50
3 rating 9 P2, QWEN, GLM 8,000 $0.60
4 recognition 1 mistral-medium-3-5 (non-panel) 2,000 $0.01
re-dispatch allowance (note (bhf), fired in five consecutive sessions) — — — $0.30
declared ceiling $1.60

Today's headroom before this run: $4.266891680 of $5.00 (S132 spent $0.733108320). Key usage is snapshotted before Stage 0 and after Stage 4; per-request usage.cost is primary and the key delta is the cross-check.

10. Verification

verify.py imports nothing from analyse.py. It re-parses every body from runs/ with an independent extractor, recomputes S, E, D, every class assignment, both κ values and the permutation null by exhaustive enumeration, re-derives the WRONG strings from the frozen table, and confirms that the GARNETT items in the blocks are byte-identical to the quoted spans in materials/sites.json. Mutation tests: at least four, each shown to flip an outcome.


Amendments

Written before Stage 1 dispatched; no seat of any kind had been called. Both come from the pre-run critic (critic.md), which returned NEEDS-AMENDMENT with 2 BLOCKING and 6 ADVISORY, all eight accepted.

A1 (from BLOCKING 1) — the extraction step is removed rather than specified. Stage 2's input is the story as numbered paragraphs in which every dialogue paragraph is replaced by its frozen ru span from materials/sites.json, attributions removed; narration paragraphs are untouched and in place. The mapping is one-to-one in text order and was checked mechanically before dispatch (TT 10 ↔ 10, SC 16 ↔ 16). The English returned at numbered paragraph n is the site span, and the lead cuts nothing out of a generated body. Without this, the lead would have chosen the boundaries of the two arms that carry P2, after seeing the English and knowing the hypothesis.

A2 (from BLOCKING 2) — a target-side variance guard, F7. Added to §7:

F7 — Stage 3 between-seat SD of standing at DISCORDANT sites exceeds that at CONCORDANT sites by > 0.75, in the arm a primary is read on → that primary is reported descriptively only.

F4 guards the source side and there was no guard on the target side; D would inflate at discordant sites from rater disagreement alone.

A3 (from ADVISORY 3) — P5 is labelled exploratory as well as descriptive. No mechanism claim rests on carrier.

A4 (from ADVISORY 6) — P3's hypergeometric P is reported with the caveat that the lead had the title, the author, the speaker attributions and the whole story, and the Stage 1 seats had none of them.

§8 limit 8, added by A1

The independent hand does not see the attribution clauses inside dialogue paragraphs, where Garnett did. The narration around them is untouched and names the speakers throughout both stories, so the loss is small, but it is a real asymmetry between the generated arms and the published one and no comparison of IND-* against GARNETT is a primary.