Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: framework/v0.1/README.md · rendered 2026-09-09

Page metadata (front matter)
typeprogram
idframework-v01
statusactive
created2026-08-02
updated2026-08-08
sensesstyle-correspondence, accuracy, voice, cultural-mediation, naturalness, consistency, purpose-fit
provisionaltrue
linksframework/v0.2/README.md, wiki/findings/results/RS-20260808b-discordance-fails.md, wiki/findings/results/RS-20260806e-published-loss.md, wiki/arms/ARM-r1-census.md, workshop/translations/kohlhaas/R20-v2/translation.md, wiki/findings/results/RS-20260804-yardstick.md, framework/README.md, framework/closure.md, framework/traceability-inventory.md, wiki/findings/results/RS-20260802e-displaced-marking.md, wiki/findings/results/RS-20260803c-occupied-slot.md, wiki/findings/results/RS-20260728j-classb-marking.md, wiki/findings/results/RS-20260802d-class-line-carry.md, wiki/findings/results/RS-20260728i-coverage-independent.md, wiki/findings/results/RS-20260730i-candidate-reach.md, wiki/findings/results/RS-20260729d-decision-grain.md, wiki/findings/results/RS-20260802-tierD-verdict.md, wiki/arms/ARM-framework-v01.md, wiki/decisions/resolved/D-20260724-04-pair-relative-sense-weights.md, PROJECT.md, config/models.md

Framework v0.1

SUCCEEDED BY framework/v0.2/README.md (2026-08-08, S133), which does NOT restate this page. v0.2 changes exactly one thing and it is a subtraction: §3 prediction 1 below is RETIRED, and the statement S1 replaces it. §2's recommendation R1 is unchanged, and §5, §6 and the whole evidence base here carry over unaltered. Read this page, then v0.2 for the delta.

The first release, drafted whole at S096. It contains one operational recommendation addressed to a translator, fourteen candidates that are not recommendations and are listed here by name with the reason for each, a coverage statement carrying no scalar, a pair-by-pair declaration, registered predictions, two open questions with measurements attached, and a changelog. The rest of this page is about why there is one recommendation and not fifteen.

Standing. provisional: true (charter §2.4 — Tier D is NOT PASSED). This is an internal versioned artifact under framework/, the named deliverable of ARM-framework-v01. It is not a publication: the repository is private, external release and publicity are Tom's alone (charter §7, §10), and nothing here changes that.


1. What the evidence permits, stated before the recommendation

Tier D was run to completion on 2026-08-02 and did not pass (RS-20260802-tierD-verdict). The release gate as the charter states it — "not before Tier D has run and at least one disciplined regime comparison has results" — is met on both limbs, but the failure removes one evidence class and it is the class most of a translation framework would ordinarily rest on:

class what it is standing
X1a a published translation read against its source, independently checked admissible
X1b the same, single-reader admissible, unverified
X2 machine-measured from stored texts, recomputed admissible
X3 panel-scored INADMISSIBLE — no jury verdict carries evidential weight

So v0.1 may not contain a single recommendation that depends on anyone's judgment of whether a translation is good. Every claim below is about what a translation does or what is available, never about what is better.

2. The recommendation

R1 — displaced marking

Where the source marks a relation or attitude by a grammatical form the target lacks, the absence of a same-category counterpart is not the absence of the marking.

Do not record the loss on the ground that the category is missing. Render the site again under a brief that requires the marking to appear, and let it fall wherever the target does mark such things — on a pronoun, a verb, an address noun, a courtesy formula, an adjective, the argument structure of the clause. Record the loss only if that second attempt fails.

Shape. Prescriptive, addressed to a translator, at the grain of a single decision. This is the first such recommendation this project has admitted. Fourteen earlier candidates were sorted at S035 and are listed in §5; a fifteenth, C15, was written at S056 and refused because two readers did not apply it the same way.

Pairs it is evidenced on (D-20260724-04): IT→EN, FI→SV, FI→EN, JA→EN, RU→EN, LZH→EN, ES→EN. Seven pairs, six source languages, two target languages. Untested on every other pair, and in particular untested with English as source — the project's standing gap since S015. The full pair-by-pair declaration is §7.

What the marking costs, measured. An independent seat rating naturalness as English prose only — never told what the study was about — scored the forced renderings at 4.17 against 4.83 for both the filed rendering and the length-matched decoy. The cost is not the cost of rewriting: the decoy is equally changed and equally long and costs nothing. It is the cost of the marking, about two-thirds of a point on a five-point scale. S052 found the same thing qualitatively at three sites of one paragraph — hypothetical → assertion, counterfactual → general truth, certainty → supposition, all scoring as unlicensed addition under D-20260727-08.

So R1 tells a translator the option exists. It does not tell them to take it, and it cannot: that judgment is X3 and X3 is inadmissible. A translator who applies R1 and then declines the marking has used it correctly.

Where R1 does not reach — and this clause is QUALIFIED as of 2026-08-04 (S106). As released, it read: the one site of six that failed was Spanish grammatical gender, where the relation is metalinguistic — the femaleness is automatic rather than chosen — and no English category encodes automaticity; R1 is evidenced for relations that are social, attitudinal or perspectival, and refuted for relations about the grammar itself.

The whole of that refutation was one site, and the relation it turned on was the lead's reading of what the Spanish conveys. RS-20260804g put the same site to a seat shown the Spanish and no English at all, and it does not read automaticity there: it reads "an intimate and nurturing bond, as if his delusions had given birth to cherished life." Against that relation the R1 rendering scores 3 of 3, and so does an independent re-rendering that avoids the word daughters altogether. So the clause now reads: whether a relation at a site is metalinguistic has, in this project, been the translator's own call, and R1's one documented refutation does not survive an independent statement of what the source conveys. R1 is not thereby shown to reach metalinguistic relations — there is no longer a tested site anyone independent calls metalinguistic.

And one site where R1's move is available and insufficient, new at S106. At Turgenev's «Роза» the formal вы between lovers has an English counterpart in the courtesy formula, and RS-20260802e scored the compensation 3 of 3 against a relation the lead stated as "the manner reserved for people who are not on intimate terms". An independent reader states the same passage as "a painful restraint, as if emotional closeness still lags behind their commitment" — and the compensation falls to 1 of 3, with a second, differently-briefed re-rendering at 0 of 3. The device carried the source's grammatical function and not the passage's effect. That distinction is new to this release and R1's text does not currently make it.

Traceability

Human-anchored, at two sites, in two language pairs — this is the part that does not depend on any machine:

Machine-read, at six further sites — RS-20260802e, this release's own evidence, in four more source languages: three independent seats, blind to provenance and in both rotations, recovered the relation from the forced rendering at 5 of 6 sites, from the filed rendering at 0 of 6, and from a length-matched unmarked decoy at 0 of 6.

REPRODUCED UNDER AN INDEPENDENT YARDSTICK, 2026-08-04 (S106) — this is §8 Q-c's answer. RS-20260804g re-ran those sites with the FROZEN, FORCED and DECOY spans byte-identical and the statements of the relation written by a seat shown no English of any kind — no rendering, no gloss, no log. One site (S1) left the denominator when its statement twice failed a leak screen committed with the design. On the five that remain: FORCED 4 of 5, FROZEN 0 of 5, DECOY 0 of 5 — a separation of +4, exactly the value RS-20260802e reached on those same five sites. DECOY was graded YES in 0 of 120 judgements across two stages. The lead registered the opposite expectation before the run and was wrong.

Two things that reproduction does not carry, and the release states them next to it. (i) The sites changed hands: the R1 rendering lost S2 and gained S6, so the count reproduces and the site-level pattern does not — see the two paragraphs below. (ii) Precision is not separated from authorship: the independent statements are in places more specific than the lead's, and nothing in that run distinguishes the yardstick's author changed from the yardstick got sharper. The rule is followable at κ 0.630–0.774 between independent readers, against the 0.452 at which C15 was refused. The two numbers are never summed. A grading YES licenses three independent readers recovered the relation from this English and licenses nothing about human readers. No human reader is available to this project: Tom is never an experimental subject (charter §9).

And the project's own logs corroborate R1 without ever having stated it. The census behind RS-20260802e found 8 sites where a grammatical marking was declared wholly lost and 8 where the same kind of gap was repaired by a compensation in a different category — the move R1 prescribes, made by the same hand that elsewhere declared it impossible, and never once connected. Two of the eight Class B labels are contested by the independent readers (RS-20260802e §5.1), so the ratio is stated as 8 : 6–8 and not as a clean eight-all.

3. Predictions, stated so a later session can score them

Registered before any application of v0.1 to fresh text, and none of them is discharged yet:

  1. On fresh Class A sites in a new pair, R1 recovers a marking at more than half. Score by the E-20260802e procedure. The failure it should not survive: recovery at or below a third. Status: still open on a FRESH pair, and now RE-TESTED CLEAN on the original one. E-20260803c (DE→EN) hid the source and is therefore not the named procedure. E-20260804 (FR→EN, S101) IS the named procedure and ran to completion on eight fresh Class A sites — and three of its own pre-registered failure criteria fired, so it is descriptive only (RS-20260804-yardstick). E-20260804g (S106) re-ran the original sites under an independent yardstick, fired nothing, and recovered at 4 of 5 (RS-20260804g) — but those are not fresh sites in a new pair, so prediction 1 is not discharged by it either. What it does is remove the yardstick as the explanation of the FR→EN failure, which §8's Q-c is now the record of. E-20260805c (PL→EN, S111) is the third attempt and its admission gate admitted 0 of 8 sites (RS-20260805c): the filed close translation already conveyed the relation an independent reader stated from the Polish alone. F5 fired and P1p was withheld a third time. STATUS AS OF 2026-08-05 (S116): the prediction is TESTABLE by the named procedure, and its obstacle is the census rather than the ruler. RS-20260805h / E-20260805h put six minimal pairs to the same seats against the same frozen relation statements — every proposition of the Polish held fixed, the marking device removed — and recovery fell from 0.861 to 0.222 and from 6 recovered sites to 1, with POSITIVE at 36 of 36 and WRONG at 0 of 36 in the same bodies. RS-20260805c §4(b)'s reading — that the procedure scores content and not marking — is REFUTED; the unbriefed paraphrase that recovered at 8 of 8 was never marking-free, because a paraphraser keeps the tone while changing the words. So prediction 1 remains OPEN and UNDISCHARGED after three attempts, and what no run has yet produced is a population of sites where a competent translation actually lost the marking. That run's own gate fired at 1 of 6 and its primary is withheld; the figures above are descriptive and are stated as such on the result page. STATUS AS OF 2026-08-06 (S121–S122): WITHHELD A FOURTH TIME, and the obstacle has moved a third time — off the census and onto the prediction's own admission condition. E-20260806-published-loss / RS-20260806e went to published practice in evidence class X1a: Kleist's «Michael Kohlhaas», DE→EN, nine vocative-free asymmetric-address sites selected from the German alone, against two published translations of the whole tale (Oxenford 1844, King 1914). The census this prediction has been waiting for exists on the textual record — King 1914 carries no English device of address at 9 of 9 sites, Oxenford at 4 of 9. And the admitted set is ONE. Under the registered admission rule — both published translations failing to convey the relation to blind seats — F1 fired at 1 site against a floor of 5 and P2 was withheld. King's device-free English was judged to convey the relation at 8 of 9 sites, with POSITIVE at 1.000, WRONG at 0.028, the direction-reversed MISMATCH control at 0.037 and every gate clean. So the reason prediction 1 cannot be reached is no longer that translators do not lose the device; it is that losing the device mostly does not lose the relation — these scenes state the standing in what the speakers say. This is the fourth attempt and the third distinct obstacle, and the honest reading is that the prediction is mis-specified rather than unlucky: an admission condition defined as the published English does not convey the relation selects for sites whose content is relation-neutral, and the only site that qualified here (A2) is the one where the speaker's content works against the deference his address claims. A version of prediction 1 that could be discharged has to define its admission on the source's marking and the target's device, not on a reader's recovery. Rewriting it is not this release's act; RS-20260806e §7 is the record. STATUS AS OF 2026-08-08 (S133): RETIRED, and framework/v0.2 §2 is the record. A fifth attempt (RS-20260808b, RU→EN, two whole Chekhov stories) tested the admission condition S128 proposed — source marks + content discordant + target device available — and the middle term did not reproduce: three readers of the Russian called the concordant/discordant question at Fleiss κ = 0.0584 against S128's 0.786, naming six utterances between them and agreeing jointly on none, while the markedness half of the same call from the same seats reproduced at 0.7214. F1 fired at 1 site against a floor of 6 and the floor was not moved. Discordance is a property of a text's trajectory, not of an utterance, so no per-utterance admission condition can select for it. Prediction 1 is therefore retired rather than re-worded, and framework/v0.2 §2's S1 states what the five attempts did establish: across five pairs the project has not produced a population of sites at which a competent English rendering loses a grammatically marked relation.
  2. R1 will fail on relations that are metalinguistic rather than social or attitudinal. S6 (grammatical gender: female, and automatically so) failed exactly here. Prediction: any future site whose relation is about a marking being automatic rather than chosen fails too. Status: NOT FALSIFIABLE BY THE NAMED PROCEDURE, established 2026-08-04. E-20260804 A3: a relation that is metalinguistic — that a marking is automatic rather than chosen — can only be conveyed by metalanguage, which is exactly what the procedure's POSITIVE arm is. The procedure therefore cannot distinguish "R1 fails on metalinguistic relations" from "this instrument cannot score metalinguistic relations". The prediction as worded here is a defect, and at the one site tried the explicit gloss was graded NO by 2 of 3 readers on the ground that it destroyed the effect it stated. A future test of prediction 2 needs a different instrument, not more sites.
  3. Applying R1 will increase logged decisions per span and will not reduce declared losses to zero. R1 changes where a translator stops, not what the target can do. Status: untested.
  4. The cost of a recovered marking will be an unlicensed addition more often than not (S052: 3 of 3). If a later session finds recoveries that are cost-free at a majority of sites, this release over-stated the trade. Status: untested.

4. What v0.1 is NOT evidenced for, stated as plainly as what it is

Device categories are not interchangeable, and the first contrast between two of them is against the release (2026-08-05, S116). R1 offers six places for a marking to fall — a pronoun, a verb, an address noun, a courtesy formula, an adjective, the argument structure of the clause — as alternatives. RS-20260805h put a pronoun substitution (you → one) and three adjective substitutions (little removed) to two independent readers asked only whether the two Englishes state the same facts: the pronoun substitution changed no fact for either reader; the adjective substitutions changed a fact for both, at three sites of three. The English periphrasis a minor little clerk predicates size of the clerk where the Polish suffix -ek does not. The suffix modulates a relation; the adjective states a fact. This is one contrast on one pair and it isolates two of the six categories — but it is evidence that the list should not be read as a free choice.

Second contrast, and it says the list is not a menu at all (2026-08-06, S121–S122, DE→EN). E-20260806-published-loss excluded the pronoun from its forced arm before translating, on the argument that thou against you is one-sided: it can mark the down-address and English has no pronoun for the up-address. The one published translator in that run who carried Kleist's asymmetry used exactly that device (Oxenford 1844, at four sites, all four Luther→Kohlhaas). The lead then rendered the whole scene under the excluded device to find out what it costs (T-kohlhaas-R20-v2, contamination: high, no evaluation designed or made), and three things came out of it that this section did not know:

  1. The categories are alternatives, not ingredients. Carrying the marking on the pronoun made the address nouns redundant rather than additional: the forced rendering's seven vocative mans collapsed to zero, and sir fell from eleven occurrences to six, the six being exactly the places the German writes hochwürdiger Herr. A "device budget" is the wrong model — applying one category does not leave the others free to apply on top of it.
  2. The pronoun cannot reach the up-address, and that is where the marking was needed. English has no deferential second-person pronoun, so the device leaves the Ihr direction with plain you — and the up-address site A2 is the only site in the whole run where an English rendering without a device failed to convey the relation. The device that worked in published practice is the one that cannot reach the one site that mattered.
  3. A device can be wrong in a way §2's clause does not catch. R1 and R20 forbid a device that asserts a fact the source does not. The archaic pronoun asserts no fact; it asserts a register — Kleist's Luther speaks ordinary 1810 German and the English thou dates him. 41 pronoun tokens dragged 15 archaic verb forms in with them. The clause should reach register and does not.

And the census's own figure is a site-level count of a scene-level decision. Oxenford uses the archaic pronoun 48 times in the Luther scene and zero times in the Herse scene, where the same German asymmetry runs between master and groom. His device tracks Luther, not the marking. A device the translator applies to a whole relationship cannot be counted turn by turn, which is a limit on every census of this kind, including this project's.

5. The fourteen candidates that are not recommendations, written out

framework/traceability-inventory.md holds the evidence rows; this is the release's own statement of why none of them is addressed to a translator as an instruction, which is what a framework release owes a reader. None is withdrawn. Every one is untested as a prescription.

Four are prescriptive about running this project, not about translating. They belong in a methods appendix and would be meaningless advice to a translator who is not also running an experiment:

# candidate why it is not in §2
10 Measure the lead's contamination before selecting material addressed to an experimenter; class X2
11 Check a published pair for dependence before using it as a baseline addressed to an experimenter; class X2
13 Measure whether a comparison's arms are identifiable without reading addressed to an experimenter; class X2
14 Never pool a paired comparison across length-sign strata addressed to an experimenter; a specification, framework/control-arm-spec.md

Nine are a vocabulary, a taxonomy and a set of diagnostic questions — real findings, and not one of them tells anybody to do anything:

# candidate shape what it actually gives a translator
1 Declare the target register; "natural" is not a single corpus procedural a question to answer before starting, not a way to answer it
2 Ask what the fluency cost, not whether it is too fluent diagnostic a better question than Venuti's, with no procedure attached
3 The handling set for culture-bound items is eight-wide; choose knowingly taxonomic an inventory of options and no rule for choosing among them
5 Grammar-borne meaning transcodes into lexis and loses systematicity descriptive a fact about what happens, not an instruction — and R1 is the one place this became an instruction
6 The false friend does its damage inside a drift window descriptive (mechanism) a warning with no operational test at a site
7 Published translators do not handle a forked class uniformly descriptive, normatively unsettled evidence that practice varies; the project cannot say which way is right without X3
8 Archaism buys inheritability back — foreignisation is a capability, not a stance descriptive folded into #6's evidence; never separately claimed
9 Long source periods are split in published practice descriptive, n = 1 refused on evidence, not on shape
4 Declare the pair; sense weights are pair-relative ratified rule D-20260724-04 binding on this project's own evaluations; it is a decision, not a recommendation

One is X3 and inadmissible. C12 — one self-revision pass buys naturalness without moving accuracy — rests entirely on panel scoring, and Tier D's failure makes it inadmissible. It is also refused on a second, independent ground: it was invoked zero times in 63 logged decisions, so reviving it would change nothing a translator does.

And one was written, tested and refused. C15 (S056) failed the followability bar: two independent readers applied it at κ 0.452 while applying a deliberately groundless rule of the same shape at κ 0.933 (RS-20260729d). That refusal is the reason R1 was tested for followability before admission and cleared it at κ 0.630–0.774.

6. Coverage, written out and carrying no scalar

The absorbed backlog row (S082) asked for this section to be written without a k, and the reason is that the scalar does not reproduce. RS-20260728i measured the proportion of real translation decisions the candidate set reaches, on identical data, with two rater pools: 0.867 and 0.289. A number that moves by a factor of three between pools is not an estimate of anything, and quoting either would be quoting a rater pool.

What does reproduce is qualitative, across three independent measurements, and it is the harder finding:

  1. RS-20260728i — 126 independent classifications of 63 real logged decisions. The candidate set decided zero. Not few. Zero. Every candidate that fired fired as a description of what the translator had already done.
  2. RS-20260730i — 45 PT→EN decision sites, a third language pair, every live rendering written down at the moment of decision, three raters given all fourteen rows. The reach did not change in kind.
  3. RS-20260729d — the followability failure. A rule can be in the set, be judged relevant, and still not be applied the same way by two readers. Coverage counted relevance; it never counted agreement.

So the release states coverage as a sentence and not as a number: the fourteen candidates in §5 name and describe a large share of what a translator is doing, and decide none of it. R1 is the first row that decides anything, and it decides one thing at one kind of site. How large that share is, is not something this project can currently measure to a figure that survives a change of raters.

7. Pairs, declared row by row under D-20260724-04

D-20260724-04 requires a release to declare the pairs it is evidenced on, because sense weights are pair-relative. Evidenced means at least one site in that pair entered R1's evidence base; everything else is untested.

pair standing for R1 what carries it
IT→EN evidenced, human-anchored RS-20260728j, Dole 1896, 2 of 3 sites
FI→SV evidenced, human-anchored RS-20260802d, Hertzberg 1886, kin-term carry
FI→EN evidenced RS-20260802d; and RS-20260803b priced the residual loss at ≈2.2 scale points
JA→EN evidenced, machine-read RS-20260802e site S1 (stacked humbling auxiliaries)
RU→EN evidenced, machine-read RS-20260802e sites S2, S3, S4
LZH→EN evidenced, machine-read RS-20260802e site S5 (humble 1sg)
ES→EN evidenced — and R1 REFUTED at the one site tested RS-20260802e site S6, metalinguistic gender
DE→EN untested, and the second attempt failed at the admission gate rather than at the test RS-20260803c ran on this pair and is void; E-20260806-published-loss (S121–S122) then ran the named procedure against two published translations and F1 fired at 1 admitted site of 9, so P2 was withheld and no R1 site on this pair has been tested; §3 prediction 1, §8 Q-b
FR→EN untested — and it is the pair where the procedure itself came apart RS-20260804 ran the named procedure to completion, F1, F2 and F5 all fired, and the filed translation scored above the forced one; §8 Q-c
PL→EN untested E-20260805c (S111) ran the named procedure and its gate admitted 0 of 8 sites; E-20260805h (S116) then showed the procedure discriminates on these sites, so the pair carries no tested R1 site and the reason is a census that has never been non-empty
EN→anything NO LONGER untested, and it never was the same question. R1's problem — a grammatical distinction the source makes and the target cannot — DOES NOT ARISE in this direction: 0 sites at A-cask-forced-choice's 26 and 0 at A-brown-calaveras-address's 7, because English marks none of this grammatically. What the direction produces instead is the mirror-image problem: an OBLIGATORY TARGET CATEGORY THE SOURCE DOES NOT FILL, and the translator must invent the filling. On how translators fill it the evidence is now two censuses and nine hands: NINE OF NINE wrote the relation SYMMETRICALLY where the English marks nothing, and the SAME hands wrote it ASYMMETRICALLY where the English marks rank. Whether the symmetric grid loses anything is untested and has failed to be tested TWICE, on two different obstacles. RS-20260809d (S143): seven published hands, 3 languages, 56 years, Poe; P1 withheld on a speaker-halo gate. RS-20260810 (S148): Harte, one published Polish hand + the lead's French, and the marked/unmarked contrast inside one story; the successor run stopped at its recognition gate, 2 of 5 seats naming the story and 5 of 5 naming the author. §8 Q-f
every other pair untested —

8. Two open questions, each with a measurement attached and neither of them a condition on R1

This section exists because E-20260803c was designed to amend R1 and its own pre-run critic established that it could not (critic.md, pass 1, seven BLOCKING findings). The amendment was abandoned before the run; what follows is measurement and an open question, which is what the evidence supports.

Q-a — is a compensation blocked when the source has already spent the device?

S095 hit one case (RS-20260803b): Canth's child says «Äiti, … Nouskaa ylös!», and the licensed kin-term compensation was unavailable because Canth had already written the kin term. E-20260803c put that to twelve second-person sites in Keller, with one mechanical repair applied identically everywhere.

The run is VOID — its registered reading check fired — and the hypothesis fails independently of the void: gain +0.833 where the address-noun slot was free against +0.633 where it was taken, difference +0.200, exact permutation over 792 splits p = 0.2487, and the whole gap rides on two sites of twelve. At three of the five occupied sites the repair scored 1.000 — a second address noun stacked on the source's own conveyed the relation to every grader.

So the release records the opposite of what it set out to record: on the one device tested, an occupied slot is not a blocked slot, and R1 gains no condition. Q-a stays open for the five device categories R1 names that have never been isolated.

Q-b — how much of an address system survives a good translation at all?

The same run's descriptive figures are the more interesting ones, and they are not licensed — they are a single work, a single contaminated translator, machine graders, and a void run:

recovery of the German address relation, 12 sites
the lead's close R04 rendering 0.167
Wolf von Schierbrand's published 1919 English 0.111
the same close rendering plus one inserted address noun 0.917

A professional published translation recovered the relation at one site in twelve, and at that one site it was copying a courtesy formula Keller had written rather than compensating for anything. If that pattern holds anywhere else, the question it raises is not whether translators can mark these relations but whether they do — and that is a question about published practice, answerable in evidence class X1a, without any jury.

Q-b's second measurement — the first with two translators, and the two halves of it disagree (2026-08-06, S121–S122)

E-20260806-published-loss / RS-20260806e asked exactly that question of Kleist's «Michael Kohlhaas» at nine vocative-free sites where the German marks the speaker–addressee relation and English has one you.

Oxenford 1844 King 1914
carries an English device of address (textual count, no panel) 4 of 9 0 of 9
judged to convey the relation (3 blind seats × 2 orderings, ≥ 4 of 6 YES) 8 of 9 8 of 9

Published practice splits on the device and does not split on the relation. So Q-b's question — how much of an address system survives a good translation at all — has, on this work, two defensible answers that point opposite ways, and the difference between them is what you decide to count. The device count is checkable by anyone who reads the two Englishes. The recovery figure is three panel models' reading and licenses nothing about a human reader (RS-20260806e §6, limit 1).

Q-b stays open, and it is now open on a sharper question than it was: not do translators mark these relations, but where in a text does the marking do work that the scene does not already do? The nine sites answer at one of them — A2, where the speaker's content contradicts the deference his address claims — and one site is a hypothesis, not a finding.

Q-c — does R1's separation survive a yardstick the lead did not write? — ANSWERED 2026-08-04 (S106): YES

Opened at S101 by RS-20260804-yardstick; answered at S106 by RS-20260804g-yardstick-holds.

The answer is yes, on the sites the question was about. With the FROZEN, FORCED and DECOY spans byte-identical to 2026-08-02 and the relation statements written by a seat shown no English at all, the separation reproduces at +4 on five sites, the same value the original run reached on those five. The repair went further than Q-c asked for: Q-c proposed giving the yardstick seat the frozen gloss, and the pre-run critic established that the glosses state the relation outright — at one site the gloss is the filed rendering — so the yardstick seats were given no English at all.

What this closes, and what it opens instead. It closes the suspicion that R1's warrant was an artifact of its author writing his own yardstick. It opens a narrower question that E-20260804's own headline can no longer answer: that run found no separation on FR→EN, and the yardstick is no longer the explanation. Whatever produced the FR→EN null was the language, the sites, or the translator's state of knowledge, and RS-20260804-yardstick §4's "the yardstick was the variable" is superseded as an explanation of it.

And the reproduction is not site-for-site, which is the part a future session should not lose: the count held while two of five sites changed verdict in opposite directions. §2 carries both.


The record as it stood before the answer, kept because the question was worth the shape it had:

R1's machine-read evidence is RS-20260802e: three independent seats recovered the relation from the FORCED rendering at 5 of 6 sites and from the FROZEN rendering at 0 of 6. The seats scored each rendering against a written statement of the relation, and the lead wrote that statement.

E-20260804 repaired exactly that, on a pre-run critic's BLOCKING finding raised in both of two independent passes: seat P5, shown the French, a gloss and the construction and no English at all, wrote every relation statement, and the lead's were discarded. On the same procedure otherwise:

7 scored sites, per-judgement YES rate
the filed close translation (FROZEN) 0.810
an independent seat's plain translation (NEUTRAL) 0.810
the R1 forced re-rendering (FORCED) 0.738
a paraphrase marking nothing (DECOY) 0.524
an explicit statement of a different relation (WRONG) 0.071

WRONG at 0.071 says the readers discriminate. They are not saying yes to everything, so the absence of a FORCED advantage is not an instrument that has stopped working.

This does not refute R1. The language, the sites and the translator's own state of knowledge all differ from E-20260802e as well, and E-20260804 is descriptive only by its own F1. What it establishes is that the independence of the yardstick has never been varied deliberately in this project, and that the one run in which it was varied found no separation. Until a design varies only that, R1's separation is not known to survive it.

What a repair would look like, so a later session does not have to invent it: re-run E-20260802e's own six sites, unchanged in every other respect, with the relation statements authored by a seat that has seen no rendering — and with POSITIVE authored from that statement rather than before it, which is the mistake E-20260804 made and paid F1 for.

9. Changelog