Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: wiki/findings/results/RS-20260808b-discordance-fails.md · rendered 2026-09-09

Page metadata (front matter)
typeresult
idRS-20260808b-discordance-fails
statusactive
created2026-08-08
updated2026-08-08
sensesstyle-correspondence, voice, cultural-mediation, accuracy
provisionaltrue
internal-judgment-onlyfalse
linksworkshop/experiments/E-20260808b-discordant-marking/design.md, workshop/experiments/E-20260808b-discordant-marking/critic.md, wiki/arms/ARM-marking-work.md, framework/v0.2/README.md, framework/v0.1/README.md, wiki/findings/results/RS-20260807d-marking-work.md, wiki/findings/results/RS-20260806e-published-loss.md, wiki/findings/results/RS-20260805h-content-or-marking.md, workshop/translations/tolstyi-i-tonkii/R04-v1/translation.md, workshop/translations/smert-chinovnika/R04-v1/translation.md, config/models.md, config/budget.md

RS-20260808b — the admission condition prediction 1′ was going to be written on does not survive contact with a second language

ARM-marking-work step 2 (T5). E-20260808b-discordant-marking, session S133, 2026-08-08 UTC. The arm closes resolved at 2 of 2 — by the second of its two completion conditions.

F1 fires and the primary is withheld. Prediction 1 is now withheld a FIFTH time, and this run is the one that says it should not be re-worded but retired. What withholds it is not a shortage of power: it is that the admission condition itself did not reproduce. The concordant/discordant call that three independent readers of Ōgai's Japanese made at Fleiss κ = 0.786 comes back, on two Chekhov stories chosen because they are made of the phenomenon, at κ = 0.0584 — three readers naming six utterances between them and agreeing jointly on none.

Tier D is NOT PASSED. No seat ranked anything, nothing was judged for quality, every sentence here is provisional.


1. The headline, in five numbers

24 of 26 utterances that 2 of 3 readers of the Russian call grammatically marked for the speaker–addressee relation, at Fleiss κ = 0.7214 — the markedness census reproduces
1 of 24 of those that 2 of 3 call DISCORDANT, at Fleiss κ = 0.0584 — the discordance census does not
0.048 mean transfer error, in scale points out of 7, at the seven most heavily marked utterances in the corpus, for Constance Garnett's published English — which deletes the deference clitic at every one of them
1.4444 the WRONG control's margin: the English instrument detects a reversed standing at 6 of 6, so the null above is a measurement and not a blind ruler
0.2222 mean disagreement between two presentations of the same English span in different API calls. The instrument is precise

The sentence this licenses. On 26 utterances of Chekhov, three readers of the Russian agree closely on where the grammar marks standing and not at all on where the content fights it; and at the sites where the marking is heaviest, four independent English renderings — including a published one that carries none of the source's clitics — convey the standing to blind readers of the English alone at a mean error of one twentieth of a scale point. The population framework/v0.1 prediction 1 has been hunting for five sessions is not there.


2. What was done

Two complete Chekhov stories of 1883, both public domain, both read whole in Russian: «Толстый и тонкий» (540 words) and «Смерть чиновника» (713 words). Chosen because they are the canonical Russian texts of deference-against-content — in the first, the thin man switches from ты to вы, «ваше превосходительство» and the словоерс in mid-conversation while his content stays childhood friendship; in the second, Chervyakov's forms are maximally deferential and his content is a man who will not stop pestering a general.

The lead translated both stories whole, R06 draft frozen and committed at 9dcf063 before the R04 revision, both committed at b5af03d before the design was written and before Garnett's English was opened. Both logs carry a registered prediction of which utterances are discordant — seven of twenty-six.

26 spoken utterances, attributions stripped, every utterance in both stories, nothing selected. Interior monologue excluded: no addressee.

The instrument is numeric profile crossing, not statement matching, because RS-20260807d §4.3 established that the E-20260802e relation-statement yardstick leaks structurally on dialogue. Readers of the Russian and readers of the English answer the same seven-point question in the same words — considering only these words, how does the speaker present his or her own social standing in relation to the person addressed? — and the statistic is D = |E − S|, the discrepancy. Nothing is written down for anyone to match against, so nothing can leak. Three source seats (P1, P3, P5) saw only Russian; three target seats (P2, qwen3.7-max, glm-5.2), disjoint from them, saw only English.

Arms: GARNETT (Constance Garnett 1922, published, evidence class X1a), IND-PLAIN and IND-FORCED (one independent hand, both stories whole, the second under the R20 brief quoted verbatim), LEAD-CLOSE (descriptive only — see §3), plus WRONG (6, direction-reversed by a frozen table) and REPEAT (6 duplicated spans in different API calls).

The contamination gate changed the design before anything was dispatched. Measured after the lead's renderings were frozen: lead ~ Garnett is 105 / 33 / 20 / longest run 27 on «Смерть чиновника» whole, against 38–54 / 7 / 2 / run 16 for two genuinely independent published hands (Garnett 1920 ~ Koteliansky & Murry 1915 on «Пари», length-matched). Restricted to dialogue the first story is clean (19 / 0 / 0 / run 11) and the second is not (39 / 6 / 2 / run 16), with the sixteen-token run falling inside SC13, one of the five sites the lead's own log had registered as discordant. So the lead was demoted out of every primary by a rule written into the design before dispatch, and the close arm the primaries are read on is Garnett's.

The pre-run critic returned NEEDS-AMENDMENT with 2 BLOCKING and 6 ADVISORY; all eight were accepted and none overruled. Its BLOCKING 1 is why P2 means anything: the design as frozen would have had the lead, knowing the hypothesis, cutting each site's words out of a generated whole-story translation by eye. Amendment A1 removed the extraction step instead of specifying it — the generation input presents dialogue paragraphs as their frozen spans, so the English at paragraph n is the site, byte for byte. Its BLOCKING 2 added F7, a target-side variance guard the design did not have.


3. Why the primary is withheld

P1 — for GARNETT, mean D at DISCORDANT sites exceeds mean D at CONCORDANT sites by ≥ 0.75, at exact permutation P ≤ 0.05 — is withheld by F1. Only one marked site is DISCORDANT by the registered 2-of-3 rule, against a floor of six.

The floor was not moved, and this is not a near miss. The three source seats:

seat discordant calls which
P1 4 SC06, SC10, SC12, TT06
P3 0 —
P5 3 SC12, TT07, TT09

Six utterances are named by somebody and none by everybody. One seat found the phenomenon nowhere in either story. Pairwise raw agreement is high (0.808–0.885) because the base rate is near zero, which is exactly the condition under which κ collapses: κ = 0.0584 against RS-20260807d's 0.786 on the same three-question instrument, in Japanese. P6 is the comparison and it fails as badly as a comparison can.

By contrast the markedness half of the same call, from the same seats on the same items, reproduces: 24 of 26 at κ = 0.7214, and the two unmarked ones («Ничего, ничего…» and Chervyakov's grumble to his wife) are the two that carry no address form at all. The seats can see the grammar. They cannot see the fight.

The descriptive figures, labelled as such

arm D at the 1 discordant site D at the 23 concordant D all 24 marked
GARNETT 0.6667 0.3551 0.3681
IND-PLAIN 1.0000 0.3913 0.4167
IND-FORCED 0.3333 0.3478 0.3472
LEAD-CLOSE 1.0000 0.3333 0.3611

P1's gap is +0.3116 at exact permutation P = 0.2083 over all 24 relabelings — the test could have reached P = 0.0417 and did not. P2's reduction is 0.6667 against a bar of 0.75 on a single site, and a single site is not a measurement; F7 fires on IND-FORCED and LEAD-CLOSE there for the same reason. No reading is taken from either number in either direction.

P4 holds: at the 23 concordant sites the R20 brief moves the transfer error by 0.0435 against a bar of 0.50. Applying R1 where the content already carries the relation buys nothing — the fifth demonstration of that in this project and the first with an independent hand writing both arms.


4. What the run does establish

4.1 The most heavily marked utterances are the ones English carries best

Seven utterances score S = −2.67 on the source side — the floor of the scale, three readers agreeing that the speaker places himself far below the person addressed. They are TT08, TT10, SC01, SC05, SC09, SC11, SC13: every «ваше превосходительство», every «ваше —ство», every stacked словоерс in both stories.

arm mean D at those seven sites
GARNETT 0.0476
IND-PLAIN 0.0476
IND-FORCED 0.0476
LEAD-CLOSE 0.1429

Garnett deletes the словоерс at all seven — «Очень приятно-с!» is delighted!, «Хи-хи-с» is He--he!, «я чихнул-с» is I sneezed, «Что-с?» is What? — and the standing arrives anyway, at a mean error of one twentieth of a scale point. The reason is not subtle and it is the same reason RS-20260806e found at Kleist and RS-20260805c at Prus: «ваше превосходительство» is a noun phrase and English has one. The clitic rides on top of a device that transfers whole.

This is the fifth pair in which the project has looked for sites where a competent English rendering loses a grammatically marked relation, and the fifth in which it has not found them (FR→EN S101, PL→EN S111 and S116, DE→EN S121–S122, JA→EN S128, and now RU→EN).

4.2 Discordance is a property of the trajectory, not of the utterance

The design put each utterance to the seats alone, without narration, speaker or story — which is the only way to keep the source-side yardstick from being written by someone who has read the English. Under that presentation, «Я, ваше превосходительство… Очень приятно-с!» is called concordant by 3 of 3. And they are right: on its own, that sentence is a man deferring, and nothing in it pulls the other way. It is discordant only if you know that the same man said «Миша! Друг детства!» ninety seconds earlier.

RS-20260807d §4.5 had already seen the shape without naming it: the five utterances the Japanese translator over-predicted were ones where "the tension he felt is in the situation rather than in the sentence's own grammar-against-content." This run makes that the whole finding. A per-utterance admission condition cannot pick out a phenomenon that lives in a text's trajectory, and RS-20260807d's κ = 0.786 was earned on a courtroom scene where the standing is restated in every line — a text where the situation is inside the sentence.

The consequence for framework/v0.2 is direct: prediction 1′ was to be written on source marks + content discordant + target device available, and the middle term is not a per-utterance property that independent readers can identify. §5 is what was written instead.

4.3 The largest transfer errors are a source-side instrument artifact, not a translation loss

site S E (Garnett) D what it is
SC10 −0.33 +2.00 2.33 «Какие пустяки… Вам что угодно?» — the general dismissing Chervyakov
SC06 −0.33 +1.67 2.00 «Ах, полноте… а вы всё о том же!» — the general again
SC02 0.00 +1.00 1.00 «Ничего, ничего…»

The two biggest discrepancies in the run are places where the Russian readers put the speaker below his addressee and the English readers put him above — and the English readers are right. The cause is visible in the seats' own form field: at both sites P1 and P3 named «вам» / «вы» as the marking form and scored it negative. Russian вы marks distance and respect toward the addressee; it does not claim lowness for the speaker, and a general saying what is it you want? to a petitioner is using it from above.

So the scale as worded conflates deference-shown with lowness-claimed, and the conflation is carried entirely by the source side, where an isolated polite sentence has no other cue. Removing those two sites drops Garnett's all-marked error from 0.3681 to about 0.20. This is a limit on this instrument, it was not anticipated by the design or by the critic, and it is the first thing a successor design has to fix. It does not touch §4.1, where source and target agree at the floor of the scale.

4.4 One site where the clitic did carry something, and only the lead kept it

SC15 is Chervyakov's whole reply to being shouted at: «Что-с?» — one interrogative pronoun and one deference particle.

English E D against S = −1.00
GARNETT What? 0.00 1.00
IND-PLAIN What? 0.00 1.00
IND-FORCED What? 0.00 1.00
LEAD-CLOSE What, sir? −0.67 0.33

The one utterance in the corpus where the clitic is doing the work by itself is the one where deleting it measurably costs something — and the arm that kept it is the only one that carries the standing. It is n = 1, the lead is out of every primary, and the appropriate weight is "worth looking at next time," not "shown." But it is the sharpest illustration in the run of what R1 is about, and it appears at the site where the source has nothing else to lean on.

4.5 The R1 brief moves the wording, and the wording was not the bottleneck

P5 (descriptive and exploratory, per amendment A3): the rate at which target seats say their rating came from the wording rather than the content.

arm wording content nothing
GARNETT 0.648 0.254 0.099
IND-PLAIN 0.625 0.222 0.153
IND-FORCED 0.732 0.225 0.042
LEAD-CLOSE 0.764 0.222 0.014

The brief does what it says — it puts more of the signal into the words, and it nearly eliminates nothing answers (0.153 → 0.042 within the same hand). And it changes the transfer error by 0.0435. The marking arrived either way.

4.6 The lead's registered prediction cannot be scored, and that is worth saying plainly

P3 registered seven of twenty-six sites as discordant, before any seat was called. One is in the seats' set — because the seats' set has exactly one member. The hypergeometric P is 0.2692 and the maximum achievable hit count was 1, so P3 is not a test of anything and no inference is drawn from either its failure or its P. Amendment A4's caveat stands on top of that: the lead had the title, the author, the attributions and the whole story; the seats had none of them — which, in the light of §4.2, is precisely the difference that produced the divergence.


5. Controls

F2 WRONG mean D 1.8333 against GARNETT's 0.3889 on the same six sites, gap 1.4444 passes; the English instrument responds to a reversed standing
F3 REPEAT mean difference
natural repeat SC14/SC16 are byte-identical in the Russian and all four arms; disagreement 0.00–0.33, source side 0.00 second, unplanned precision estimate
F4 source-side SD 0.943 discordant vs 0.455 concordant does not fire (and rests on one site)
F5 cells 346 of 348 = 99.43% passes
F6 markedness 24 of 26 passes
F7 target-side SD fires on IND-FORCED and LEAD-CLOSE at the single discordant site noted; both are already descriptive-only
recognition one non-seat call named Chekhov, "very confident", and got the title wrong see limit 4

F2 is necessary and not sufficient (critic ADVISORY 8): it shows the seats catch a reversal manufactured in Garnett-style prose, not that they would catch a subtler loss.

Recognition is real and it is conservative for P1. A reader who knows the story knows who is the general; that pulls E toward S at exactly the sites where discordance would have mattered, shrinking the effect the design predicted. It cannot manufacture the null in §4.1, where E and S agree at the floor of the scale in all four arms.


6. Limits

  1. Two stories, one author, one year, one pair, one published hand. RU→EN, Chekhov 1883, Garnett 1922.
  2. The corpus is deliberately enriched and no rate here is a base rate for Russian prose. The base rate on record stays RS-20260807d's 4 of 51.
  3. §4.3 is a defect in this run's own scale, not a finding about translation, and it inflates every all-marked mean reported here.
  4. Recognition is uncontrolled, measured on one non-seat call, and conservative for P1.
  5. LEAD-CLOSE is contaminated — 27 contiguous tokens with Garnett on «Смерть чиновника» whole, 16 inside the dialogue — and is in no primary. §4.4's observation is subject to that.
  6. The generated arms did not see the attribution clauses that Garnett saw (amendment A1's declared cost); the surrounding narration, which names the speakers, was untouched.
  7. One seat's block-1 body returned site ids without the arm suffix. Its line count equals the block length and its site order matches the block position for position, so the ids were recovered positionally with that equality asserted in both analyse.py and verify.py. Two cells were lost outright (one seat, block 2) and are the whole of the 0.57% shortfall.
  8. F1 fired on a floor of six that the design set for power, not for meaning. Had the floor been three, P1 would still have been withheld — the union of all three seats' discordant calls is six sites and their intersection is empty, so no 3-of-3 rule and no 2-of-3 rule reaches a usable set on this corpus.

7. What this hands to framework/v0.2

Written, not just named — framework/v0.2/README.md is this session's other product.


8. Spend

$0.851843373 against a declared $1.60 (53%). Key-usage reconciliation is exact to 1e-9: opening 77.420865591, closing 78.272708963, delta 0.851843372, per-request sum 0.851843373.

(The opening snapshot is $0.2554 above config/budget.md's closing figure for S132; the key is used outside this project and per-request costs are primary. Recorded, not explained.)

Waste row: $0.420768053, 49.4% of the session, five bodies, all preserved in runs/discarded/ and none overwritten (note (bhd)).

Lead translation is $0 and is never ledgered. Two complete stories rendered twice, both logs, the 26-site inventory, the contamination cells and all analysis and verification cost nothing.

Verification: verify.py, 143 checks, 0 failures, 5 mutation tests, 5 caught. It imports nothing from analyse.py, re-parses every body with its own key–value extractor, recomputes both Fleiss κ and the permutation null by exhaustive enumeration, re-derives the six WRONG strings from the frozen table, and checks that every GARNETT, LEAD and REPEAT string dispatched in a block is byte-identical to the frozen span and that every IND-* string round-trips against the numbered paragraph of the generation body it came from. Two mutation tests failed on their first run and were mis-sized, not wrong — setting all six WRONG ratings to −3 still leaves a gap of 1.22 because three of the six sites sit near −3 already, and moving one REPEAT copy to the ceiling reaches 0.722 against a 0.75 bar. Both were replaced with correctly sized forms, and both the reason and the replacement are in the file.