Repository path: workshop/experiments/E-20260808b-discordant-marking/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260808b-discordant-marking |
| status | frozen |
| created | 2026-08-08 |
| updated | 2026-08-08 |
| senses | style-correspondence, voice, cultural-mediation, accuracy |
| provisional | true |
| internal-judgment-only | false |
| links | wiki/arms/ARM-marking-work.md, framework/v0.1/README.md, wiki/findings/results/RS-20260807d-marking-work.md, wiki/findings/results/RS-20260806e-published-loss.md, wiki/findings/results/RS-20260806f-persona-crossing.md, workshop/translations/tolstyi-i-tonkii/R04-v1/translation.md, workshop/translations/smert-chinovnika/R04-v1/translation.md, workshop/regimes/R20-r1-forced.md, config/models.md, config/budget.md |
E-20260808b — where a deference form fights its own content, does a published English rendering lose the standing?
ARM-marking-work step 2 (T5). Frozen 2026-08-08 UTC, session S133, before any API call of
any kind. The two lead renderings and their registered predictions were committed at b5af03d
before this file was written.
1. Why this run exists, and why step 2 could not be written without it
ARM-marking-work step 2 is scoped as a writing step: state prediction 1′ on the admission
condition step 1 measured, and write it into a framework/v0.2. Two things block writing it.
(a) The admission condition has one measurement, on one text, in one language.
RS-20260807d established the population prediction 1 needs — utterances where the source marks
the speaker–addressee relation and the content pulls against the standing the form claims — at
4 of 51, Fleiss κ = 0.786, in Ōgai's courtroom. A 1′ written on a base rate from a single
Japanese text would be the fifth version of prediction 1 written on evidence too thin to carry it.
(b) The procedure a 1′ would name is known to be broken for this material.
RS-20260807d §4.3: the E-20260802e yardstick — a written statement of the relation, which
blind readers match the English against — leaks at 9 of 12 statements on dialogue, and
structurally, because a relation between two speakers is enacted through the speech act, so any
statement of it entails the content. Address marking lives in dialogue. Writing 1′ on a
procedure that cannot be run on the material 1′ is about would repeat the mistake four attempts
have already made.
This run supplies both. Its yardstick is numeric profile crossing, not statement matching:
readers of the source and readers of the English answer the same question on the same scale, and
the statistic is the discrepancy between them. That instrument shape is not new here — it is
RS-20260806f / RS-20260807's narrator crossing, which fired 16 of 16 twice — and it cannot
leak, because nothing is written down for anyone to match against.
One sentence on what this teaches about translating literature (continue-prompt.md §4.5):
it measures where in a story a grammatical mark of social standing carries something the words do
not, and whether a published English translator loses it there.
2. Materials
Two complete stories by Chekhov, both 1883, both public domain, both read whole in Russian.
| «Толстый и тонкий» | «Смерть чиновника» | |
|---|---|---|
| copy-text | ПСС 1975 via ru.wikisource | ФЭБ / ПСС via ru.wikisource |
| Russian words | 540 | 713 |
| spoken utterances | 10 | 16 |
Why these. Both stories are about the phenomenon. In the first, the thin man switches from
ты to вы, to «ваше превосходительство» and to the словоерс (the deference clitic
-с) in mid-conversation, while his content stays childhood-friendship. In the second,
Chervyakov's forms are maximally deferential and his content is a man refusing to stop pestering a
general. The corpus is deliberately enriched: RS-20260807d found the phenomenon at ~8% of
marked utterances, and a design wanting power cannot sample randomly. Enrichment is declared, and
it bounds the base rate this run may report (§8 limit 1).
The site inventory
26 spoken utterances, attributions stripped, hand-built from the copy-texts and the frozen
renderings before any seat was called: materials/sites.json. Interior monologue is excluded — it
has no addressee. Nothing was selected: every utterance in both stories is in.
Arms
| arm | what it is | provenance |
|---|---|---|
SRC |
the Russian utterance | copy-text |
GARNETT |
Constance Garnett 1922, The Tales of Chekhov vol. 13 (Gutenberg #13414) | published, independent, evidence class X1a |
IND-PLAIN |
an independent hand renders both stories whole, no brief beyond "translate" | generated, stage 2 |
IND-FORCED |
the same hand, both stories whole, under the R20 brief quoted verbatim from framework/v0.1 §2 |
generated, stage 2 |
LEAD-CLOSE |
T-tolstyi-i-tonkii-R04-v1 / T-smert-chinovnika-R04-v1 |
descriptive only, excluded from every primary by the rule in §6 |
WRONG |
6 sites, direction-reversed by the frozen table in materials/wrong-table.md |
positive control |
REPEAT |
6 GARNETT items duplicated into a different block |
precision control |
IND-PLAIN exists because IND-FORCED alone would confound the brief with the hand. The
contrast that answers prediction 1 is IND-FORCED − IND-PLAIN within one hand.
The contamination gate changed this design before anything was dispatched
Measured after the two lead renderings were frozen and committed, with
tools/dependence_check.py:
| cell | shared 7-grams | 12 | 15 | longest run |
|---|---|---|---|---|
lead R04 ~ Garnett, «Толстый и тонкий», whole |
59 | 16 | 3 | 16 |
lead R04 ~ Garnett, «Смерть чиновника», whole |
105 | 33 | 20 | 27 |
| two independent published hands (Garnett 1920 ~ Koteliansky & Murry 1915 on «Пари», length-matched) | 38–54 | 7 | 2 | 16 |
lead R04 ~ Garnett, dialogue only, «Толстый и тонкий» |
19 | 0 | 0 | 11 — clean |
lead R04 ~ Garnett, dialogue only, «Смерть чиновника» |
39 | 6 | 2 | 16 — dependent |
The lead's close rendering is not independent of Garnett's, and the 16-token dialogue run in
«Смерть чиновника» falls inside SC13, one of the five sites the lead's own frozen log
registered as discordant. Under the standing rule (CLAUDE.md, and note (bhb)) the lead is
therefore not an arm this run's validity may rest on: LEAD-CLOSE is rated and reported and is
excluded from P1, P2 and P4 by a rule written here, before dispatch. The close arm the
primaries are read on is Garnett, a published hand, in evidence class X1a.
3. Questions
- Does the discordance census transfer? Is the concordant/discordant distinction that
RS-20260807dmeasured at κ 0.786 in Japanese reproducible at agreement in Russian, by readers shown only the source? - Is discordance where a close rendering loses the standing? This is the admission condition
1′needs, and no run has tested it. - Does the R1 move recover it there? This is prediction 1, on an admission condition defined on the source and the target rather than on a reader's recovery.
4. Procedure
Stages run in order. Every raw body is written to runs/ before anything is parsed.
Stage 0 — pre-run critic. One non-panel call (nvidia/nemotron-3-ultra-550b-a55b), given this
design whole. BLOCKING findings are answered before Stage 1 dispatches; every finding is recorded
in critic.md with accept/overrule and a reason.
Stage 1 — the source census. Three seats — P1, P3, P5 — each shown the 26 Russian utterances, in a fixed shuffled order, with no English of any kind, no speaker labels, no story title, no author. Per utterance, three answers:
mark— Does the Russian wording itself mark the speaker's standing relative to the person addressed by a grammatical form — pronoun choice, verb form, a particle, a formula of address?yes/no, and if yes, name the form.standing— Considering only these words, how does the speaker present his or her own social standing in relation to the person addressed? −3 far below · −2 clearly below · −1 slightly below · 0 level, or the words carry no indication · +1 slightly above · +2 clearly above · +3 far above.fit— Does what the speaker is saying agree with the standing the wording claims, or pull against it?concordant/discordant.
Stage 2 — generation. One hand, P4 (moonshotai/kimi-k3), which is a seat in neither
stage. Two dispatches per story: IND-PLAIN (translate the numbered paragraphs into English) and
IND-FORCED (the same, plus the R20 brief quoted verbatim). The hand is told nothing about the
study, nothing about discordance, and is given no site list. Paragraph numbering is returned as
given and checked mechanically; a body whose numbering does not round-trip is discarded and
re-dispatched.
Stage 3 — the English rating. Three seats — P2, QWEN (qwen/qwen3.7-max), GLM
(z-ai/glm-5.2), disjoint from Stage 1 — each shown 116 English items in three blocks, in a
fixed shuffled order that mixes arms and works, with no Russian, no arm label, no speaker label,
no story. Per item, two answers:
standing— the identical seven-point question and identical wording as Stage 1.carrier— What is your rating based on?wording(forms of address, particles, verb forms, politeness formulas) /content(what is being said, however it is worded) /nothing(the words carry no indication).
Stage 4 — recognition, reported and non-gating. One non-seat call shown ten GARNETT items
and asked to name work and author. Reported; it does not gate, and §8 limit 4 states why
recognition is conservative for P1.
Order within every block is fixed by sha256(design-id | item-id), computed in build_blocks.py
and committed before dispatch.
5. The statistic
For each site i and arm a:
- S(i) = mean of the three Stage 1
standingratings. - E(i,a) = mean of the three Stage 3
standingratings. - D(i,a) = |E(i,a) − S(i)| — the transfer error, in scale points.
A site is DISCORDANT iff at least 2 of the 3 Stage 1 seats return fit: discordant;
CONCORDANT otherwise. A site is MARKED iff at least 2 of 3 return mark: yes. Only
MARKED sites enter any primary.
6. Predictions, registered
| statement | bar | |
|---|---|---|
P1 (primary) |
For GARNETT, mean D at DISCORDANT sites exceeds mean D at CONCORDANT sites |
≥ 0.75 scale points and exact permutation P ≤ 0.05 over all relabelings of MARKED sites into the two observed class sizes |
P2 |
At DISCORDANT sites, IND-FORCED has lower mean D than IND-PLAIN |
≥ 0.75 |
P3 |
The lead's frozen log (L10 in both artifacts, 7 of 26 sites) matches the seats' DISCORDANT set |
≥ 5 of 7, with the exact hypergeometric P reported beside it |
P4 |
At CONCORDANT sites, IND-FORCED does not beat IND-PLAIN |
|Δ| < 0.50 |
P5 (descriptive) |
The carrier: wording rate is higher for IND-FORCED than for GARNETT |
direction only |
P6 (descriptive) |
Fleiss κ on the Stage 1 concordant/discordant call is comparable to RS-20260807d's 0.786 |
reported, no bar |
LEAD-CLOSE appears in P5 and in every descriptive table and in no primary (§2).
What the lead expects and is putting on the record so it can be wrong: P1 holds, P2 holds,
P3 fails at 4 of 7 (the S128 translator over-predicted his own source by a factor of two and
this one will too), P4 holds, P5 holds.
7. Failure criteria, which withhold rather than adjust
| fires when | consequence | |
|---|---|---|
F1 |
fewer than 6 MARKED sites are DISCORDANT | P1 and P2 withheld; the census is still reported |
F2 |
mean D(WRONG) does not exceed mean D(GARNETT) on the same six sites by ≥ 1.00 |
the English instrument is not shown to respond to standing; all primaries withheld |
F3 |
mean absolute difference between the two presentations of a REPEAT pair exceeds 0.75 |
precision insufficient; all primaries withheld |
F4 |
Stage 1 between-seat SD of standing at DISCORDANT sites exceeds that at CONCORDANT sites by > 0.75 |
P1 is confounded with source-side ambiguity and is reported descriptively only |
F5 |
fewer than 90% of cells return | all primaries withheld |
F6 |
fewer than 18 of 26 sites are MARKED by 2 of 3 seats | the corpus is not what the design claims; all primaries withheld |
No bar in §6 or §7 is moved after it fires. Any amendment is written into critic.md or an
Amendments section with a timestamp, and amendments after a stage has returned are marked as
such.
8. Limits, written before the run
- The corpus is deliberately enriched. Any discordance rate here is a rate in two stories
chosen because they are about deference, and is not a base rate for Russian prose. The base
rate on record stays
RS-20260807d's 4 of 51. - Two stories, one author, one year, one pair. Chekhov 1883, RU→EN.
- One published close hand. Garnett is the only published English of these two stories the
project can reach; there is no second published arm and therefore no published-pair contrast of
the kind
RS-20260806ehad. - Recognition is uncontrolled and is conservative for
P1. Both stories are famous. A seat that recognises «Смерть чиновника» knows Chervyakov is a clerk addressing a general, and would therefore ratestandingcloser to the source at exactly the discordant sites — which shrinks the effectP1predicts. Stage 4 measures how much recognition there is. - The lead wrote the
WRONGinversions knowing the hypothesis. They are a positive control on the instrument, not an arm under test, and they are frozen inmaterials/wrong-table.mdbefore dispatch. IND-PLAINandIND-FORCEDare one hand. Whatever that hand's register defaults are, they cancel in the within-hand contrast and do not cancel in any comparison against Garnett.- Tier D is NOT PASSED. No seat ranks anything and nothing here is a quality judgment; every
sentence of the result is
provisional.
9. Spend
Pre-flight worst case built from max_tokens, per note (abc) — not from expected output length.
| stage | calls | model | max_tokens |
worst case |
|---|---|---|---|---|
| 0 critic | 1 | nemotron-3-ultra (non-panel) | 16,000 | $0.06 |
| 1 census | 3 | P1, P3, P5 | 6,000 | $0.13 |
| 2 generation | 4 | P4 | 8,000 | $0.50 |
| 3 rating | 9 | P2, QWEN, GLM | 8,000 | $0.60 |
| 4 recognition | 1 | mistral-medium-3-5 (non-panel) | 2,000 | $0.01 |
| re-dispatch allowance (note (bhf), fired in five consecutive sessions) | — | — | — | $0.30 |
| declared ceiling | $1.60 |
Today's headroom before this run: $4.266891680 of $5.00 (S132 spent $0.733108320). Key usage
is snapshotted before Stage 0 and after Stage 4; per-request usage.cost is primary and the key
delta is the cross-check.
10. Verification
verify.py imports nothing from analyse.py. It re-parses every body from runs/ with an
independent extractor, recomputes S, E, D, every class assignment, both κ values and the
permutation null by exhaustive enumeration, re-derives the WRONG strings from the frozen
table, and confirms that the GARNETT items in the blocks are byte-identical to the quoted spans
in materials/sites.json. Mutation tests: at least four, each shown to flip an outcome.
Amendments
Written before Stage 1 dispatched; no seat of any kind had been called. Both come from the
pre-run critic (critic.md), which returned NEEDS-AMENDMENT with 2 BLOCKING and 6 ADVISORY, all
eight accepted.
A1 (from BLOCKING 1) — the extraction step is removed rather than specified. Stage 2's input
is the story as numbered paragraphs in which every dialogue paragraph is replaced by its frozen
ru span from materials/sites.json, attributions removed; narration paragraphs are untouched
and in place. The mapping is one-to-one in text order and was checked mechanically before dispatch
(TT 10 ↔ 10, SC 16 ↔ 16). The English returned at numbered paragraph n is the site span, and
the lead cuts nothing out of a generated body. Without this, the lead would have chosen the
boundaries of the two arms that carry P2, after seeing the English and knowing the hypothesis.
A2 (from BLOCKING 2) — a target-side variance guard, F7. Added to §7:
F7— Stage 3 between-seat SD ofstandingat DISCORDANT sites exceeds that at CONCORDANT sites by > 0.75, in the arm a primary is read on → that primary is reported descriptively only.
F4 guards the source side and there was no guard on the target side; D would inflate at
discordant sites from rater disagreement alone.
A3 (from ADVISORY 3) — P5 is labelled exploratory as well as descriptive. No mechanism
claim rests on carrier.
A4 (from ADVISORY 6) — P3's hypergeometric P is reported with the caveat that the lead
had the title, the author, the speaker attributions and the whole story, and the Stage 1 seats had
none of them.
§8 limit 8, added by A1
The independent hand does not see the attribution clauses inside dialogue paragraphs, where
Garnett did. The narration around them is untouched and names the speakers throughout both
stories, so the loss is small, but it is a real asymmetry between the generated arms and the
published one and no comparison of IND-* against GARNETT is a primary.