Repository path: wiki/findings/results/RS-20260808b-discordance-fails.md · rendered 2026-09-09
Page metadata (front matter)
| type | result |
|---|---|
| id | RS-20260808b-discordance-fails |
| status | active |
| created | 2026-08-08 |
| updated | 2026-08-08 |
| senses | style-correspondence, voice, cultural-mediation, accuracy |
| provisional | true |
| internal-judgment-only | false |
| links | workshop/experiments/E-20260808b-discordant-marking/design.md, workshop/experiments/E-20260808b-discordant-marking/critic.md, wiki/arms/ARM-marking-work.md, framework/v0.2/README.md, framework/v0.1/README.md, wiki/findings/results/RS-20260807d-marking-work.md, wiki/findings/results/RS-20260806e-published-loss.md, wiki/findings/results/RS-20260805h-content-or-marking.md, workshop/translations/tolstyi-i-tonkii/R04-v1/translation.md, workshop/translations/smert-chinovnika/R04-v1/translation.md, config/models.md, config/budget.md |
RS-20260808b — the admission condition prediction 1′ was going to be written on does not survive contact with a second language
ARM-marking-work step 2 (T5). E-20260808b-discordant-marking, session S133, 2026-08-08 UTC.
The arm closes resolved at 2 of 2 — by the second of its two completion conditions.
F1 fires and the primary is withheld. Prediction 1 is now withheld a FIFTH time, and this run
is the one that says it should not be re-worded but retired. What withholds it is not a shortage
of power: it is that the admission condition itself did not reproduce. The
concordant/discordant call that three independent readers of Ōgai's Japanese made at Fleiss
κ = 0.786 comes back, on two Chekhov stories chosen because they are made of the phenomenon, at
κ = 0.0584 — three readers naming six utterances between them and agreeing jointly on none.
Tier D is NOT PASSED. No seat ranked anything, nothing was judged for quality, every sentence
here is provisional.
1. The headline, in five numbers
| 24 of 26 | utterances that 2 of 3 readers of the Russian call grammatically marked for the speaker–addressee relation, at Fleiss κ = 0.7214 — the markedness census reproduces |
| 1 of 24 | of those that 2 of 3 call DISCORDANT, at Fleiss κ = 0.0584 — the discordance census does not |
| 0.048 | mean transfer error, in scale points out of 7, at the seven most heavily marked utterances in the corpus, for Constance Garnett's published English — which deletes the deference clitic at every one of them |
| 1.4444 | the WRONG control's margin: the English instrument detects a reversed standing at 6 of 6, so the null above is a measurement and not a blind ruler |
| 0.2222 | mean disagreement between two presentations of the same English span in different API calls. The instrument is precise |
The sentence this licenses. On 26 utterances of Chekhov, three readers of the Russian agree
closely on where the grammar marks standing and not at all on where the content fights it; and
at the sites where the marking is heaviest, four independent English renderings — including a
published one that carries none of the source's clitics — convey the standing to blind readers of
the English alone at a mean error of one twentieth of a scale point. The population
framework/v0.1 prediction 1 has been hunting for five sessions is not there.
2. What was done
Two complete Chekhov stories of 1883, both public domain, both read whole in Russian: «Толстый и тонкий» (540 words) and «Смерть чиновника» (713 words). Chosen because they are the canonical Russian texts of deference-against-content — in the first, the thin man switches from ты to вы, «ваше превосходительство» and the словоерс in mid-conversation while his content stays childhood friendship; in the second, Chervyakov's forms are maximally deferential and his content is a man who will not stop pestering a general.
The lead translated both stories whole, R06 draft frozen and committed at 9dcf063 before
the R04 revision, both committed at b5af03d before the design was written and before Garnett's
English was opened. Both logs carry a registered prediction of which utterances are discordant
— seven of twenty-six.
26 spoken utterances, attributions stripped, every utterance in both stories, nothing selected. Interior monologue excluded: no addressee.
The instrument is numeric profile crossing, not statement matching, because RS-20260807d §4.3
established that the E-20260802e relation-statement yardstick leaks structurally on dialogue.
Readers of the Russian and readers of the English answer the same seven-point question in the same
words — considering only these words, how does the speaker present his or her own social standing
in relation to the person addressed? — and the statistic is D = |E − S|, the discrepancy.
Nothing is written down for anyone to match against, so nothing can leak. Three source seats
(P1, P3, P5) saw only Russian; three target seats (P2, qwen3.7-max, glm-5.2), disjoint from
them, saw only English.
Arms: GARNETT (Constance Garnett 1922, published, evidence class X1a), IND-PLAIN and
IND-FORCED (one independent hand, both stories whole, the second under the R20 brief quoted
verbatim), LEAD-CLOSE (descriptive only — see §3), plus WRONG (6, direction-reversed by a
frozen table) and REPEAT (6 duplicated spans in different API calls).
The contamination gate changed the design before anything was dispatched. Measured after the lead's renderings were frozen: lead ~ Garnett is 105 / 33 / 20 / longest run 27 on «Смерть чиновника» whole, against 38–54 / 7 / 2 / run 16 for two genuinely independent published hands (Garnett 1920 ~ Koteliansky & Murry 1915 on «Пари», length-matched). Restricted to dialogue the first story is clean (19 / 0 / 0 / run 11) and the second is not (39 / 6 / 2 / run 16), with the sixteen-token run falling inside SC13, one of the five sites the lead's own log had registered as discordant. So the lead was demoted out of every primary by a rule written into the design before dispatch, and the close arm the primaries are read on is Garnett's.
The pre-run critic returned NEEDS-AMENDMENT with 2 BLOCKING and 6 ADVISORY; all eight were
accepted and none overruled. Its BLOCKING 1 is why P2 means anything: the design as frozen
would have had the lead, knowing the hypothesis, cutting each site's words out of a generated
whole-story translation by eye. Amendment A1 removed the extraction step instead of specifying
it — the generation input presents dialogue paragraphs as their frozen spans, so the English at
paragraph n is the site, byte for byte. Its BLOCKING 2 added F7, a target-side variance
guard the design did not have.
3. Why the primary is withheld
P1 — for GARNETT, mean D at DISCORDANT sites exceeds mean D at CONCORDANT sites by
≥ 0.75, at exact permutation P ≤ 0.05 — is withheld by F1. Only one marked site is
DISCORDANT by the registered 2-of-3 rule, against a floor of six.
The floor was not moved, and this is not a near miss. The three source seats:
| seat | discordant calls | which |
|---|---|---|
| P1 | 4 | SC06, SC10, SC12, TT06 |
| P3 | 0 | — |
| P5 | 3 | SC12, TT07, TT09 |
Six utterances are named by somebody and none by everybody. One seat found the phenomenon
nowhere in either story. Pairwise raw agreement is high (0.808–0.885) because the base rate is
near zero, which is exactly the condition under which κ collapses: κ = 0.0584 against
RS-20260807d's 0.786 on the same three-question instrument, in Japanese. P6 is the
comparison and it fails as badly as a comparison can.
By contrast the markedness half of the same call, from the same seats on the same items, reproduces: 24 of 26 at κ = 0.7214, and the two unmarked ones («Ничего, ничего…» and Chervyakov's grumble to his wife) are the two that carry no address form at all. The seats can see the grammar. They cannot see the fight.
The descriptive figures, labelled as such
| arm | D at the 1 discordant site |
D at the 23 concordant |
D all 24 marked |
|---|---|---|---|
GARNETT |
0.6667 | 0.3551 | 0.3681 |
IND-PLAIN |
1.0000 | 0.3913 | 0.4167 |
IND-FORCED |
0.3333 | 0.3478 | 0.3472 |
LEAD-CLOSE |
1.0000 | 0.3333 | 0.3611 |
P1's gap is +0.3116 at exact permutation P = 0.2083 over all 24 relabelings — the test
could have reached P = 0.0417 and did not. P2's reduction is 0.6667 against a bar of 0.75
on a single site, and a single site is not a measurement; F7 fires on IND-FORCED and
LEAD-CLOSE there for the same reason. No reading is taken from either number in either
direction.
P4 holds: at the 23 concordant sites the R20 brief moves the transfer error by 0.0435
against a bar of 0.50. Applying R1 where the content already carries the relation buys nothing —
the fifth demonstration of that in this project and the first with an independent hand writing both
arms.
4. What the run does establish
4.1 The most heavily marked utterances are the ones English carries best
Seven utterances score S = −2.67 on the source side — the floor of the scale, three readers
agreeing that the speaker places himself far below the person addressed. They are
TT08, TT10, SC01, SC05, SC09, SC11, SC13: every «ваше превосходительство», every
«ваше —ство», every stacked словоерс in both stories.
| arm | mean D at those seven sites |
|---|---|
GARNETT |
0.0476 |
IND-PLAIN |
0.0476 |
IND-FORCED |
0.0476 |
LEAD-CLOSE |
0.1429 |
Garnett deletes the словоерс at all seven — «Очень приятно-с!» is delighted!, «Хи-хи-с» is
He--he!, «я чихнул-с» is I sneezed, «Что-с?» is What? — and the standing arrives anyway, at
a mean error of one twentieth of a scale point. The reason is not subtle and it is the same reason
RS-20260806e found at Kleist and RS-20260805c at Prus: «ваше превосходительство» is a noun
phrase and English has one. The clitic rides on top of a device that transfers whole.
This is the fifth pair in which the project has looked for sites where a competent English rendering loses a grammatically marked relation, and the fifth in which it has not found them (FR→EN S101, PL→EN S111 and S116, DE→EN S121–S122, JA→EN S128, and now RU→EN).
4.2 Discordance is a property of the trajectory, not of the utterance
The design put each utterance to the seats alone, without narration, speaker or story — which is the only way to keep the source-side yardstick from being written by someone who has read the English. Under that presentation, «Я, ваше превосходительство… Очень приятно-с!» is called concordant by 3 of 3. And they are right: on its own, that sentence is a man deferring, and nothing in it pulls the other way. It is discordant only if you know that the same man said «Миша! Друг детства!» ninety seconds earlier.
RS-20260807d §4.5 had already seen the shape without naming it: the five utterances the Japanese
translator over-predicted were ones where "the tension he felt is in the situation rather than
in the sentence's own grammar-against-content." This run makes that the whole finding. A
per-utterance admission condition cannot pick out a phenomenon that lives in a text's trajectory,
and RS-20260807d's κ = 0.786 was earned on a courtroom scene where the standing is restated in
every line — a text where the situation is inside the sentence.
The consequence for framework/v0.2 is direct: prediction 1′ was to be written on source
marks + content discordant + target device available, and the middle term is not a per-utterance
property that independent readers can identify. §5 is what was written instead.
4.3 The largest transfer errors are a source-side instrument artifact, not a translation loss
| site | S |
E (Garnett) |
D |
what it is |
|---|---|---|---|---|
SC10 |
−0.33 | +2.00 | 2.33 | «Какие пустяки… Вам что угодно?» — the general dismissing Chervyakov |
SC06 |
−0.33 | +1.67 | 2.00 | «Ах, полноте… а вы всё о том же!» — the general again |
SC02 |
0.00 | +1.00 | 1.00 | «Ничего, ничего…» |
The two biggest discrepancies in the run are places where the Russian readers put the speaker
below his addressee and the English readers put him above — and the English readers are right.
The cause is visible in the seats' own form field: at both sites P1 and P3 named «вам» / «вы»
as the marking form and scored it negative. Russian вы marks distance and respect toward the
addressee; it does not claim lowness for the speaker, and a general saying what is it you want?
to a petitioner is using it from above.
So the scale as worded conflates deference-shown with lowness-claimed, and the conflation is carried entirely by the source side, where an isolated polite sentence has no other cue. Removing those two sites drops Garnett's all-marked error from 0.3681 to about 0.20. This is a limit on this instrument, it was not anticipated by the design or by the critic, and it is the first thing a successor design has to fix. It does not touch §4.1, where source and target agree at the floor of the scale.
4.4 One site where the clitic did carry something, and only the lead kept it
SC15 is Chervyakov's whole reply to being shouted at: «Что-с?» — one interrogative pronoun
and one deference particle.
| English | E |
D against S = −1.00 |
|
|---|---|---|---|
GARNETT |
What? | 0.00 | 1.00 |
IND-PLAIN |
What? | 0.00 | 1.00 |
IND-FORCED |
What? | 0.00 | 1.00 |
LEAD-CLOSE |
What, sir? | −0.67 | 0.33 |
The one utterance in the corpus where the clitic is doing the work by itself is the one where
deleting it measurably costs something — and the arm that kept it is the only one that carries
the standing. It is n = 1, the lead is out of every primary, and the appropriate weight is
"worth looking at next time," not "shown." But it is the sharpest illustration in the run of what
R1 is about, and it appears at the site where the source has nothing else to lean on.
4.5 The R1 brief moves the wording, and the wording was not the bottleneck
P5 (descriptive and exploratory, per amendment A3): the rate at which target seats say
their rating came from the wording rather than the content.
| arm | wording | content | nothing |
|---|---|---|---|
GARNETT |
0.648 | 0.254 | 0.099 |
IND-PLAIN |
0.625 | 0.222 | 0.153 |
IND-FORCED |
0.732 | 0.225 | 0.042 |
LEAD-CLOSE |
0.764 | 0.222 | 0.014 |
The brief does what it says — it puts more of the signal into the words, and it nearly eliminates nothing answers (0.153 → 0.042 within the same hand). And it changes the transfer error by 0.0435. The marking arrived either way.
4.6 The lead's registered prediction cannot be scored, and that is worth saying plainly
P3 registered seven of twenty-six sites as discordant, before any seat was called. One is in
the seats' set — because the seats' set has exactly one member. The hypergeometric P is
0.2692 and the maximum achievable hit count was 1, so P3 is not a test of anything and no
inference is drawn from either its failure or its P. Amendment A4's caveat stands on top of
that: the lead had the title, the author, the attributions and the whole story; the seats had none
of them — which, in the light of §4.2, is precisely the difference that produced the divergence.
5. Controls
F2 WRONG |
mean D 1.8333 against GARNETT's 0.3889 on the same six sites, gap 1.4444 |
passes; the English instrument responds to a reversed standing |
F3 REPEAT |
mean | difference |
| natural repeat | SC14/SC16 are byte-identical in the Russian and all four arms; disagreement 0.00–0.33, source side 0.00 |
second, unplanned precision estimate |
F4 source-side SD |
0.943 discordant vs 0.455 concordant | does not fire (and rests on one site) |
F5 cells |
346 of 348 = 99.43% | passes |
F6 markedness |
24 of 26 | passes |
F7 target-side SD |
fires on IND-FORCED and LEAD-CLOSE at the single discordant site |
noted; both are already descriptive-only |
| recognition | one non-seat call named Chekhov, "very confident", and got the title wrong | see limit 4 |
F2 is necessary and not sufficient (critic ADVISORY 8): it shows the seats catch a reversal
manufactured in Garnett-style prose, not that they would catch a subtler loss.
Recognition is real and it is conservative for P1. A reader who knows the story knows who is
the general; that pulls E toward S at exactly the sites where discordance would have
mattered, shrinking the effect the design predicted. It cannot manufacture the null in §4.1, where
E and S agree at the floor of the scale in all four arms.
6. Limits
- Two stories, one author, one year, one pair, one published hand. RU→EN, Chekhov 1883, Garnett 1922.
- The corpus is deliberately enriched and no rate here is a base rate for Russian prose. The
base rate on record stays
RS-20260807d's 4 of 51. - §4.3 is a defect in this run's own scale, not a finding about translation, and it inflates every all-marked mean reported here.
- Recognition is uncontrolled, measured on one non-seat call, and conservative for
P1. LEAD-CLOSEis contaminated — 27 contiguous tokens with Garnett on «Смерть чиновника» whole, 16 inside the dialogue — and is in no primary. §4.4's observation is subject to that.- The generated arms did not see the attribution clauses that Garnett saw (amendment
A1's declared cost); the surrounding narration, which names the speakers, was untouched. - One seat's block-1 body returned site ids without the arm suffix. Its line count equals the
block length and its site order matches the block position for position, so the ids were
recovered positionally with that equality asserted in both
analyse.pyandverify.py. Two cells were lost outright (one seat, block 2) and are the whole of the 0.57% shortfall. F1fired on a floor of six that the design set for power, not for meaning. Had the floor been three,P1would still have been withheld — the union of all three seats' discordant calls is six sites and their intersection is empty, so no 3-of-3 rule and no 2-of-3 rule reaches a usable set on this corpus.
7. What this hands to framework/v0.2
Written, not just named — framework/v0.2/README.md is this session's other product.
- Prediction 1 is RETIRED rather than re-worded, after five attempts, five pairs and four distinct obstacles, the last of which is that its admission condition is not a per-utterance property independent readers can identify.
- A new statement replaces it in §3, which is what the evidence supports: across five pairs the project has not produced a population of sites where a competent English rendering loses a grammatically marked relation, so R1 addresses a problem whose frequency is, on this evidence, low.
- §8 Q-b gains its third and fourth measurements and its answer changes shape: not how much survives, but how little there was to lose.
- §4 gains the clitic-versus-noun-phrase contrast from §4.1 and §4.4 — the sharpest device-category evidence the release has, because the two categories sit inside one utterance.
- §7 gains a strengthened RU→EN row, and a new open question Q-d on trajectory-level marking, which is where §4.2 points.
8. Spend
$0.851843373 against a declared $1.60 (53%). Key-usage reconciliation is exact to 1e-9: opening 77.420865591, closing 78.272708963, delta 0.851843372, per-request sum 0.851843373.
(The opening snapshot is $0.2554 above config/budget.md's closing figure for S132; the key is
used outside this project and per-request costs are primary. Recorded, not explained.)
Waste row: $0.420768053, 49.4% of the session, five bodies, all preserved in
runs/discarded/ and none overwritten (note (bhd)).
- Note (bhf), sixth consecutive session:
deepseek-v4-proreturned zero content atfinish_reason: lengthon the census; the cap-raise remedy (note (bhq)) worked at the first attempt, $0.0089 lost. - Then three of four generation bodies from
moonshotai/kimi-k3returned zero content, at $0.365763 — the whole completion budget spent on hidden reasoning — while the fourth, from the same model on the same prompt shape, came back clean. Note (bkh)'s the failure is prompt-shaped, not seat-shaped does not fit: here the same prompt shape succeeded once and failed three times on one model in one batch. The seat was changed per note (bhf) rule (iii) andmistralai/mistral-medium-3-5produced all four arms cleanly for $0.042, an eighth of what the failures cost. - The fourth kimi body was clean and was discarded deliberately, so that
IND-PLAINandIND-FORCEDcome from one hand. Mixing hands would have confounded the arm with the author, which is the contrastP2exists to make.
Lead translation is $0 and is never ledgered. Two complete stories rendered twice, both logs, the 26-site inventory, the contamination cells and all analysis and verification cost nothing.
Verification: verify.py, 143 checks, 0 failures, 5 mutation tests, 5 caught. It imports
nothing from analyse.py, re-parses every body with its own key–value extractor, recomputes both
Fleiss κ and the permutation null by exhaustive enumeration, re-derives the six WRONG strings
from the frozen table, and checks that every GARNETT, LEAD and REPEAT string dispatched in a
block is byte-identical to the frozen span and that every IND-* string round-trips against the
numbered paragraph of the generation body it came from. Two mutation tests failed on their first
run and were mis-sized, not wrong — setting all six WRONG ratings to −3 still leaves a gap of
1.22 because three of the six sites sit near −3 already, and moving one REPEAT copy to the
ceiling reaches 0.722 against a 0.75 bar. Both were replaced with correctly sized forms, and both
the reason and the replacement are in the file.