Repository path: framework/v0.1/README.md · rendered 2026-09-09
Page metadata (front matter)
Framework v0.1
SUCCEEDED BY
framework/v0.2/README.md(2026-08-08, S133), which does NOT restate this page. v0.2 changes exactly one thing and it is a subtraction: §3 prediction 1 below is RETIRED, and the statementS1replaces it. §2's recommendation R1 is unchanged, and §5, §6 and the whole evidence base here carry over unaltered. Read this page, then v0.2 for the delta.
The first release, drafted whole at S096. It contains one operational recommendation addressed to a translator, fourteen candidates that are not recommendations and are listed here by name with the reason for each, a coverage statement carrying no scalar, a pair-by-pair declaration, registered predictions, two open questions with measurements attached, and a changelog. The rest of this page is about why there is one recommendation and not fifteen.
Standing. provisional: true (charter §2.4 — Tier D is NOT PASSED). This is an internal
versioned artifact under framework/, the named deliverable of ARM-framework-v01. It is not a
publication: the repository is private, external release and publicity are Tom's alone
(charter §7, §10), and nothing here changes that.
1. What the evidence permits, stated before the recommendation
Tier D was run to completion on 2026-08-02 and did not pass (RS-20260802-tierD-verdict). The
release gate as the charter states it — "not before Tier D has run and at least one disciplined
regime comparison has results" — is met on both limbs, but the failure removes one evidence
class and it is the class most of a translation framework would ordinarily rest on:
| class | what it is | standing |
|---|---|---|
| X1a | a published translation read against its source, independently checked | admissible |
| X1b | the same, single-reader | admissible, unverified |
| X2 | machine-measured from stored texts, recomputed | admissible |
| X3 | panel-scored | INADMISSIBLE — no jury verdict carries evidential weight |
So v0.1 may not contain a single recommendation that depends on anyone's judgment of whether a translation is good. Every claim below is about what a translation does or what is available, never about what is better.
2. The recommendation
R1 — displaced marking
Where the source marks a relation or attitude by a grammatical form the target lacks, the absence of a same-category counterpart is not the absence of the marking.
Do not record the loss on the ground that the category is missing. Render the site again under a brief that requires the marking to appear, and let it fall wherever the target does mark such things — on a pronoun, a verb, an address noun, a courtesy formula, an adjective, the argument structure of the clause. Record the loss only if that second attempt fails.
Shape. Prescriptive, addressed to a translator, at the grain of a single decision. This is the
first such recommendation this project has admitted. Fourteen earlier candidates were sorted at
S035 and are listed in §5; a fifteenth, C15, was written at S056 and refused because two
readers did not apply it the same way.
Pairs it is evidenced on (D-20260724-04): IT→EN, FI→SV, FI→EN, JA→EN, RU→EN, LZH→EN, ES→EN.
Seven pairs, six source languages, two target languages. Untested on every other pair, and in
particular untested with English as source — the project's standing gap since S015. The full
pair-by-pair declaration is §7.
What the marking costs, measured. An independent seat rating naturalness as English prose
only — never told what the study was about — scored the forced renderings at 4.17 against
4.83 for both the filed rendering and the length-matched decoy. The cost is not the cost of
rewriting: the decoy is equally changed and equally long and costs nothing. It is the cost of the
marking, about two-thirds of a point on a five-point scale. S052 found the same thing
qualitatively at three sites of one paragraph — hypothetical → assertion, counterfactual → general
truth, certainty → supposition, all scoring as unlicensed addition under D-20260727-08.
So R1 tells a translator the option exists. It does not tell them to take it, and it cannot: that judgment is X3 and X3 is inadmissible. A translator who applies R1 and then declines the marking has used it correctly.
Where R1 does not reach — and this clause is QUALIFIED as of 2026-08-04 (S106). As released, it read: the one site of six that failed was Spanish grammatical gender, where the relation is metalinguistic — the femaleness is automatic rather than chosen — and no English category encodes automaticity; R1 is evidenced for relations that are social, attitudinal or perspectival, and refuted for relations about the grammar itself.
The whole of that refutation was one site, and the relation it turned on was the lead's reading of
what the Spanish conveys. RS-20260804g put the same site to a seat shown the Spanish and no
English at all, and it does not read automaticity there: it reads "an intimate and nurturing
bond, as if his delusions had given birth to cherished life." Against that relation the R1 rendering
scores 3 of 3, and so does an independent re-rendering that avoids the word daughters
altogether. So the clause now reads: whether a relation at a site is metalinguistic has, in this
project, been the translator's own call, and R1's one documented refutation does not survive an
independent statement of what the source conveys. R1 is not thereby shown to reach
metalinguistic relations — there is no longer a tested site anyone independent calls metalinguistic.
And one site where R1's move is available and insufficient, new at S106. At Turgenev's «Роза» the
formal вы between lovers has an English counterpart in the courtesy formula, and RS-20260802e
scored the compensation 3 of 3 against a relation the lead stated as "the manner reserved for people
who are not on intimate terms". An independent reader states the same passage as "a painful
restraint, as if emotional closeness still lags behind their commitment" — and the compensation
falls to 1 of 3, with a second, differently-briefed re-rendering at 0 of 3. The device
carried the source's grammatical function and not the passage's effect. That distinction is new to
this release and R1's text does not currently make it.
Traceability
Human-anchored, at two sites, in two language pairs — this is the part that does not depend on any machine:
RS-20260728j(X1b, IT→EN). Verga's out-of-sequence conditionals. The frozen log: "there is nothing to repair with — English has no marked conditional to reach for." True of the category. Dole 1896 marked two of the three sites, by changing grammatical category — conditional to present tense, conditional to were-subjunctive — and a forced re-rendering made blind reached the same construction at both.RS-20260802d(X1b, FI→SV / FI→EN). Canth's honorific plural, child to mother. The frozen register: "no compensation available that would not be a fabrication." Hertzberg 1886 — into Swedish, a language that has the pronominal honorific — dropped it and carried the deference on a kin term, "Blir mamma länge borta?", a device English also has. Six independent machine attempts on the same passage marked nothing; the human translator marked it.
Machine-read, at six further sites — RS-20260802e, this release's own evidence, in four more
source languages: three independent seats, blind to provenance and in both rotations, recovered
the relation from the forced rendering at 5 of 6 sites, from the filed rendering at 0 of 6, and from
a length-matched unmarked decoy at 0 of 6.
REPRODUCED UNDER AN INDEPENDENT YARDSTICK, 2026-08-04 (S106) — this is §8 Q-c's answer.
RS-20260804g re-ran those sites with the FROZEN, FORCED and DECOY spans byte-identical and the
statements of the relation written by a seat shown no English of any kind — no rendering, no
gloss, no log. One site (S1) left the denominator when its statement twice failed a leak screen
committed with the design. On the five that remain: FORCED 4 of 5, FROZEN 0 of 5, DECOY 0 of 5 —
a separation of +4, exactly the value RS-20260802e reached on those same five sites. DECOY was
graded YES in 0 of 120 judgements across two stages. The lead registered the opposite
expectation before the run and was wrong.
Two things that reproduction does not carry, and the release states them next to it. (i) The
sites changed hands: the R1 rendering lost S2 and gained S6, so the count reproduces and the
site-level pattern does not — see the two paragraphs below. (ii) Precision is not separated from
authorship: the independent statements are in places more specific than the lead's, and nothing in
that run distinguishes the yardstick's author changed from the yardstick got sharper. The rule is followable at κ 0.630–0.774 between
independent readers, against the 0.452 at which C15 was refused. The two numbers are never
summed. A grading YES licenses three independent readers recovered the relation from this English
and licenses nothing about human readers. No human reader is available to this project: Tom is never
an experimental subject (charter §9).
And the project's own logs corroborate R1 without ever having stated it. The census behind
RS-20260802e found 8 sites where a grammatical marking was declared wholly lost and 8 where the
same kind of gap was repaired by a compensation in a different category — the move R1 prescribes,
made by the same hand that elsewhere declared it impossible, and never once connected. Two of the
eight Class B labels are contested by the independent readers (RS-20260802e §5.1), so the ratio
is stated as 8 : 6–8 and not as a clean eight-all.
3. Predictions, stated so a later session can score them
Registered before any application of v0.1 to fresh text, and none of them is discharged yet:
- On fresh Class A sites in a new pair, R1 recovers a marking at more than half. Score by the
E-20260802eprocedure. The failure it should not survive: recovery at or below a third. Status: still open on a FRESH pair, and now RE-TESTED CLEAN on the original one.E-20260803c(DE→EN) hid the source and is therefore not the named procedure.E-20260804(FR→EN, S101) IS the named procedure and ran to completion on eight fresh Class A sites — and three of its own pre-registered failure criteria fired, so it is descriptive only (RS-20260804-yardstick).E-20260804g(S106) re-ran the original sites under an independent yardstick, fired nothing, and recovered at 4 of 5 (RS-20260804g) — but those are not fresh sites in a new pair, so prediction 1 is not discharged by it either. What it does is remove the yardstick as the explanation of the FR→EN failure, which §8's Q-c is now the record of.E-20260805c(PL→EN, S111) is the third attempt and its admission gate admitted 0 of 8 sites (RS-20260805c): the filed close translation already conveyed the relation an independent reader stated from the Polish alone.F5fired andP1pwas withheld a third time. STATUS AS OF 2026-08-05 (S116): the prediction is TESTABLE by the named procedure, and its obstacle is the census rather than the ruler.RS-20260805h/E-20260805hput six minimal pairs to the same seats against the same frozen relation statements — every proposition of the Polish held fixed, the marking device removed — and recovery fell from 0.861 to 0.222 and from 6 recovered sites to 1, withPOSITIVEat 36 of 36 andWRONGat 0 of 36 in the same bodies.RS-20260805c§4(b)'s reading — that the procedure scores content and not marking — is REFUTED; the unbriefed paraphrase that recovered at 8 of 8 was never marking-free, because a paraphraser keeps the tone while changing the words. So prediction 1 remains OPEN and UNDISCHARGED after three attempts, and what no run has yet produced is a population of sites where a competent translation actually lost the marking. That run's own gate fired at 1 of 6 and its primary is withheld; the figures above are descriptive and are stated as such on the result page. STATUS AS OF 2026-08-06 (S121–S122): WITHHELD A FOURTH TIME, and the obstacle has moved a third time — off the census and onto the prediction's own admission condition.E-20260806-published-loss/RS-20260806ewent to published practice in evidence class X1a: Kleist's «Michael Kohlhaas», DE→EN, nine vocative-free asymmetric-address sites selected from the German alone, against two published translations of the whole tale (Oxenford 1844, King 1914). The census this prediction has been waiting for exists on the textual record — King 1914 carries no English device of address at 9 of 9 sites, Oxenford at 4 of 9. And the admitted set is ONE. Under the registered admission rule — both published translations failing to convey the relation to blind seats —F1fired at 1 site against a floor of 5 andP2was withheld. King's device-free English was judged to convey the relation at 8 of 9 sites, withPOSITIVEat 1.000,WRONGat 0.028, the direction-reversedMISMATCHcontrol at 0.037 and every gate clean. So the reason prediction 1 cannot be reached is no longer that translators do not lose the device; it is that losing the device mostly does not lose the relation — these scenes state the standing in what the speakers say. This is the fourth attempt and the third distinct obstacle, and the honest reading is that the prediction is mis-specified rather than unlucky: an admission condition defined as the published English does not convey the relation selects for sites whose content is relation-neutral, and the only site that qualified here (A2) is the one where the speaker's content works against the deference his address claims. A version of prediction 1 that could be discharged has to define its admission on the source's marking and the target's device, not on a reader's recovery. Rewriting it is not this release's act;RS-20260806e§7 is the record. STATUS AS OF 2026-08-08 (S133): RETIRED, andframework/v0.2§2 is the record. A fifth attempt (RS-20260808b, RU→EN, two whole Chekhov stories) tested the admission condition S128 proposed — source marks + content discordant + target device available — and the middle term did not reproduce: three readers of the Russian called the concordant/discordant question at Fleiss κ = 0.0584 against S128's 0.786, naming six utterances between them and agreeing jointly on none, while the markedness half of the same call from the same seats reproduced at 0.7214.F1fired at 1 site against a floor of 6 and the floor was not moved. Discordance is a property of a text's trajectory, not of an utterance, so no per-utterance admission condition can select for it. Prediction 1 is therefore retired rather than re-worded, andframework/v0.2§2'sS1states what the five attempts did establish: across five pairs the project has not produced a population of sites at which a competent English rendering loses a grammatically marked relation. - R1 will fail on relations that are metalinguistic rather than social or attitudinal.
S6(grammatical gender: female, and automatically so) failed exactly here. Prediction: any future site whose relation is about a marking being automatic rather than chosen fails too. Status: NOT FALSIFIABLE BY THE NAMED PROCEDURE, established 2026-08-04.E-20260804A3: a relation that is metalinguistic — that a marking is automatic rather than chosen — can only be conveyed by metalanguage, which is exactly what the procedure's POSITIVE arm is. The procedure therefore cannot distinguish "R1 fails on metalinguistic relations" from "this instrument cannot score metalinguistic relations". The prediction as worded here is a defect, and at the one site tried the explicit gloss was graded NO by 2 of 3 readers on the ground that it destroyed the effect it stated. A future test of prediction 2 needs a different instrument, not more sites. - Applying R1 will increase logged decisions per span and will not reduce declared losses to zero. R1 changes where a translator stops, not what the target can do. Status: untested.
- The cost of a recovered marking will be an unlicensed addition more often than not (S052: 3 of 3). If a later session finds recoveries that are cost-free at a majority of sites, this release over-stated the trade. Status: untested.
4. What v0.1 is NOT evidenced for, stated as plainly as what it is
- Nothing about quality. Not whether a marked rendering is better than an unmarked one, not whether any translation here is good. Tier D is NOT PASSED.
- Nothing at the level of a whole translation. R1 fires at a site. On 126 independent
classifications of 63 real decisions, this project's earlier candidates decided zero
(
RS-20260728i), and that finding is not withdrawn by R1's admission — it is what R1 was built against. - No coverage figure. See §6, which is written out and carries no scalar, deliberately.
- Nothing about English as a source language. Six of the seven evidenced pairs have English as target. The one exception (Poe → Japanese) is a Class B corroboration, not a tested site.
- Nothing about human readers, per §2.
- Nothing about how many of R1's six licensed device categories actually work. All of the evidence above, and all of §8's, concerns compensations the lead or a published translator happened to reach for. Five of the six categories R1 names have never been isolated.
Device categories are not interchangeable, and the first contrast between two of them is
against the release (2026-08-05, S116). R1 offers six places for a marking to fall — a pronoun, a
verb, an address noun, a courtesy formula, an adjective, the argument structure of the clause — as
alternatives. RS-20260805h put a pronoun substitution (you → one) and three adjective
substitutions (little removed) to two independent readers asked only whether the two Englishes
state the same facts: the pronoun substitution changed no fact for either reader; the adjective
substitutions changed a fact for both, at three sites of three. The English periphrasis a minor
little clerk predicates size of the clerk where the Polish suffix -ek does not. The suffix
modulates a relation; the adjective states a fact. This is one contrast on one pair and it isolates
two of the six categories — but it is evidence that the list should not be read as a free choice.
Second contrast, and it says the list is not a menu at all (2026-08-06, S121–S122, DE→EN).
E-20260806-published-loss excluded the pronoun from its forced arm before translating, on the
argument that thou against you is one-sided: it can mark the down-address and English has no
pronoun for the up-address. The one published translator in that run who carried Kleist's
asymmetry used exactly that device (Oxenford 1844, at four sites, all four Luther→Kohlhaas). The
lead then rendered the whole scene under the excluded device to find out what it costs
(T-kohlhaas-R20-v2, contamination: high, no evaluation designed or made), and three things came
out of it that this section did not know:
- The categories are alternatives, not ingredients. Carrying the marking on the pronoun made the
address nouns redundant rather than additional: the forced rendering's seven vocative mans
collapsed to zero, and sir fell from eleven occurrences to six, the six being exactly the places
the German writes
hochwürdiger Herr. A "device budget" is the wrong model — applying one category does not leave the others free to apply on top of it. - The pronoun cannot reach the up-address, and that is where the marking was needed. English has
no deferential second-person pronoun, so the device leaves the
Ihrdirection with plain you — and the up-address siteA2is the only site in the whole run where an English rendering without a device failed to convey the relation. The device that worked in published practice is the one that cannot reach the one site that mattered. - A device can be wrong in a way §2's clause does not catch. R1 and
R20forbid a device that asserts a fact the source does not. The archaic pronoun asserts no fact; it asserts a register — Kleist's Luther speaks ordinary 1810 German and the English thou dates him. 41 pronoun tokens dragged 15 archaic verb forms in with them. The clause should reach register and does not.
And the census's own figure is a site-level count of a scene-level decision. Oxenford uses the archaic pronoun 48 times in the Luther scene and zero times in the Herse scene, where the same German asymmetry runs between master and groom. His device tracks Luther, not the marking. A device the translator applies to a whole relationship cannot be counted turn by turn, which is a limit on every census of this kind, including this project's.
5. The fourteen candidates that are not recommendations, written out
framework/traceability-inventory.md holds the evidence rows; this is the release's own statement of
why none of them is addressed to a translator as an instruction, which is what a framework
release owes a reader. None is withdrawn. Every one is untested as a prescription.
Four are prescriptive about running this project, not about translating. They belong in a methods appendix and would be meaningless advice to a translator who is not also running an experiment:
| # | candidate | why it is not in §2 |
|---|---|---|
| 10 | Measure the lead's contamination before selecting material | addressed to an experimenter; class X2 |
| 11 | Check a published pair for dependence before using it as a baseline | addressed to an experimenter; class X2 |
| 13 | Measure whether a comparison's arms are identifiable without reading | addressed to an experimenter; class X2 |
| 14 | Never pool a paired comparison across length-sign strata | addressed to an experimenter; a specification, framework/control-arm-spec.md |
Nine are a vocabulary, a taxonomy and a set of diagnostic questions — real findings, and not one of them tells anybody to do anything:
| # | candidate | shape | what it actually gives a translator |
|---|---|---|---|
| 1 | Declare the target register; "natural" is not a single corpus | procedural | a question to answer before starting, not a way to answer it |
| 2 | Ask what the fluency cost, not whether it is too fluent | diagnostic | a better question than Venuti's, with no procedure attached |
| 3 | The handling set for culture-bound items is eight-wide; choose knowingly | taxonomic | an inventory of options and no rule for choosing among them |
| 5 | Grammar-borne meaning transcodes into lexis and loses systematicity | descriptive | a fact about what happens, not an instruction — and R1 is the one place this became an instruction |
| 6 | The false friend does its damage inside a drift window | descriptive (mechanism) | a warning with no operational test at a site |
| 7 | Published translators do not handle a forked class uniformly | descriptive, normatively unsettled | evidence that practice varies; the project cannot say which way is right without X3 |
| 8 | Archaism buys inheritability back — foreignisation is a capability, not a stance | descriptive | folded into #6's evidence; never separately claimed |
| 9 | Long source periods are split in published practice | descriptive, n = 1 | refused on evidence, not on shape |
| 4 | Declare the pair; sense weights are pair-relative | ratified rule D-20260724-04 |
binding on this project's own evaluations; it is a decision, not a recommendation |
One is X3 and inadmissible. C12 — one self-revision pass buys naturalness without moving
accuracy — rests entirely on panel scoring, and Tier D's failure makes it inadmissible. It is
also refused on a second, independent ground: it was invoked zero times in 63 logged
decisions, so reviving it would change nothing a translator does.
And one was written, tested and refused. C15 (S056) failed the followability bar: two
independent readers applied it at κ 0.452 while applying a deliberately groundless rule of the
same shape at κ 0.933 (RS-20260729d). That refusal is the reason R1 was tested for
followability before admission and cleared it at κ 0.630–0.774.
6. Coverage, written out and carrying no scalar
The absorbed backlog row (S082) asked for this section to be written without a k, and the
reason is that the scalar does not reproduce. RS-20260728i measured the proportion of real
translation decisions the candidate set reaches, on identical data, with two rater pools: 0.867 and
0.289. A number that moves by a factor of three between pools is not an estimate of anything, and
quoting either would be quoting a rater pool.
What does reproduce is qualitative, across three independent measurements, and it is the harder finding:
RS-20260728i— 126 independent classifications of 63 real logged decisions. The candidate set decided zero. Not few. Zero. Every candidate that fired fired as a description of what the translator had already done.RS-20260730i— 45 PT→EN decision sites, a third language pair, every live rendering written down at the moment of decision, three raters given all fourteen rows. The reach did not change in kind.RS-20260729d— the followability failure. A rule can be in the set, be judged relevant, and still not be applied the same way by two readers. Coverage counted relevance; it never counted agreement.
So the release states coverage as a sentence and not as a number: the fourteen candidates in §5 name and describe a large share of what a translator is doing, and decide none of it. R1 is the first row that decides anything, and it decides one thing at one kind of site. How large that share is, is not something this project can currently measure to a figure that survives a change of raters.
7. Pairs, declared row by row under D-20260724-04
D-20260724-04 requires a release to declare the pairs it is evidenced on, because sense weights are
pair-relative. Evidenced means at least one site in that pair entered R1's evidence base;
everything else is untested.
| pair | standing for R1 | what carries it |
|---|---|---|
| IT→EN | evidenced, human-anchored | RS-20260728j, Dole 1896, 2 of 3 sites |
| FI→SV | evidenced, human-anchored | RS-20260802d, Hertzberg 1886, kin-term carry |
| FI→EN | evidenced | RS-20260802d; and RS-20260803b priced the residual loss at ≈2.2 scale points |
| JA→EN | evidenced, machine-read | RS-20260802e site S1 (stacked humbling auxiliaries) |
| RU→EN | evidenced, machine-read | RS-20260802e sites S2, S3, S4 |
| LZH→EN | evidenced, machine-read | RS-20260802e site S5 (humble 1sg) |
| ES→EN | evidenced — and R1 REFUTED at the one site tested | RS-20260802e site S6, metalinguistic gender |
| DE→EN | untested, and the second attempt failed at the admission gate rather than at the test |
RS-20260803c ran on this pair and is void; E-20260806-published-loss (S121–S122) then ran the named procedure against two published translations and F1 fired at 1 admitted site of 9, so P2 was withheld and no R1 site on this pair has been tested; §3 prediction 1, §8 Q-b |
| FR→EN | untested — and it is the pair where the procedure itself came apart |
RS-20260804 ran the named procedure to completion, F1, F2 and F5 all fired, and the filed translation scored above the forced one; §8 Q-c |
| PL→EN | untested |
E-20260805c (S111) ran the named procedure and its gate admitted 0 of 8 sites; E-20260805h (S116) then showed the procedure discriminates on these sites, so the pair carries no tested R1 site and the reason is a census that has never been non-empty |
| EN→anything | NO LONGER untested, and it never was the same question. R1's problem — a grammatical distinction the source makes and the target cannot — DOES NOT ARISE in this direction: 0 sites at A-cask-forced-choice's 26 and 0 at A-brown-calaveras-address's 7, because English marks none of this grammatically. What the direction produces instead is the mirror-image problem: an OBLIGATORY TARGET CATEGORY THE SOURCE DOES NOT FILL, and the translator must invent the filling. On how translators fill it the evidence is now two censuses and nine hands: NINE OF NINE wrote the relation SYMMETRICALLY where the English marks nothing, and the SAME hands wrote it ASYMMETRICALLY where the English marks rank. Whether the symmetric grid loses anything is untested and has failed to be tested TWICE, on two different obstacles. |
RS-20260809d (S143): seven published hands, 3 languages, 56 years, Poe; P1 withheld on a speaker-halo gate. RS-20260810 (S148): Harte, one published Polish hand + the lead's French, and the marked/unmarked contrast inside one story; the successor run stopped at its recognition gate, 2 of 5 seats naming the story and 5 of 5 naming the author. §8 Q-f |
| every other pair | untested |
— |
8. Two open questions, each with a measurement attached and neither of them a condition on R1
This section exists because E-20260803c was designed to amend R1 and its own pre-run critic
established that it could not (critic.md, pass 1, seven BLOCKING findings). The amendment was
abandoned before the run; what follows is measurement and an open question, which is what the
evidence supports.
Q-a — is a compensation blocked when the source has already spent the device?
S095 hit one case (RS-20260803b): Canth's child says «Äiti, … Nouskaa ylös!», and the licensed
kin-term compensation was unavailable because Canth had already written the kin term.
E-20260803c put that to twelve second-person sites in Keller, with one mechanical repair applied
identically everywhere.
The run is VOID — its registered reading check fired — and the hypothesis fails independently of the void: gain +0.833 where the address-noun slot was free against +0.633 where it was taken, difference +0.200, exact permutation over 792 splits p = 0.2487, and the whole gap rides on two sites of twelve. At three of the five occupied sites the repair scored 1.000 — a second address noun stacked on the source's own conveyed the relation to every grader.
So the release records the opposite of what it set out to record: on the one device tested, an occupied slot is not a blocked slot, and R1 gains no condition. Q-a stays open for the five device categories R1 names that have never been isolated.
Q-b — how much of an address system survives a good translation at all?
The same run's descriptive figures are the more interesting ones, and they are not licensed — they are a single work, a single contaminated translator, machine graders, and a void run:
| recovery of the German address relation, 12 sites | |
|---|---|
the lead's close R04 rendering |
0.167 |
| Wolf von Schierbrand's published 1919 English | 0.111 |
| the same close rendering plus one inserted address noun | 0.917 |
A professional published translation recovered the relation at one site in twelve, and at that one site it was copying a courtesy formula Keller had written rather than compensating for anything. If that pattern holds anywhere else, the question it raises is not whether translators can mark these relations but whether they do — and that is a question about published practice, answerable in evidence class X1a, without any jury.
Q-b's second measurement — the first with two translators, and the two halves of it disagree (2026-08-06, S121–S122)
E-20260806-published-loss / RS-20260806e asked exactly that question of Kleist's «Michael
Kohlhaas» at nine vocative-free sites where the German marks the speaker–addressee relation and
English has one you.
| Oxenford 1844 | King 1914 | |
|---|---|---|
| carries an English device of address (textual count, no panel) | 4 of 9 | 0 of 9 |
| judged to convey the relation (3 blind seats × 2 orderings, ≥ 4 of 6 YES) | 8 of 9 | 8 of 9 |
Published practice splits on the device and does not split on the relation. So Q-b's question —
how much of an address system survives a good translation at all — has, on this work, two
defensible answers that point opposite ways, and the difference between them is what you decide to
count. The device count is checkable by anyone who reads the two Englishes. The recovery figure is
three panel models' reading and licenses nothing about a human reader (RS-20260806e §6, limit 1).
Q-b stays open, and it is now open on a sharper question than it was: not do translators mark
these relations, but where in a text does the marking do work that the scene does not already
do? The nine sites answer at one of them — A2, where the speaker's content contradicts the
deference his address claims — and one site is a hypothesis, not a finding.
Q-c — does R1's separation survive a yardstick the lead did not write? — ANSWERED 2026-08-04 (S106): YES
Opened at S101 by RS-20260804-yardstick; answered at S106 by RS-20260804g-yardstick-holds.
The answer is yes, on the sites the question was about. With the FROZEN, FORCED and DECOY spans byte-identical to 2026-08-02 and the relation statements written by a seat shown no English at all, the separation reproduces at +4 on five sites, the same value the original run reached on those five. The repair went further than Q-c asked for: Q-c proposed giving the yardstick seat the frozen gloss, and the pre-run critic established that the glosses state the relation outright — at one site the gloss is the filed rendering — so the yardstick seats were given no English at all.
What this closes, and what it opens instead. It closes the suspicion that R1's warrant was an
artifact of its author writing his own yardstick. It opens a narrower question that E-20260804's
own headline can no longer answer: that run found no separation on FR→EN, and the yardstick is no
longer the explanation. Whatever produced the FR→EN null was the language, the sites, or the
translator's state of knowledge, and RS-20260804-yardstick §4's "the yardstick was the variable"
is superseded as an explanation of it.
And the reproduction is not site-for-site, which is the part a future session should not lose: the count held while two of five sites changed verdict in opposite directions. §2 carries both.
The record as it stood before the answer, kept because the question was worth the shape it had:
R1's machine-read evidence is RS-20260802e: three independent seats recovered the relation from the
FORCED rendering at 5 of 6 sites and from the FROZEN rendering at 0 of 6. The seats scored
each rendering against a written statement of the relation, and the lead wrote that statement.
E-20260804 repaired exactly that, on a pre-run critic's BLOCKING finding raised in both of two
independent passes: seat P5, shown the French, a gloss and the construction and no English at all,
wrote every relation statement, and the lead's were discarded. On the same procedure otherwise:
| 7 scored sites, per-judgement YES rate | |
|---|---|
| the filed close translation (FROZEN) | 0.810 |
| an independent seat's plain translation (NEUTRAL) | 0.810 |
| the R1 forced re-rendering (FORCED) | 0.738 |
| a paraphrase marking nothing (DECOY) | 0.524 |
| an explicit statement of a different relation (WRONG) | 0.071 |
WRONG at 0.071 says the readers discriminate. They are not saying yes to everything, so the absence of a FORCED advantage is not an instrument that has stopped working.
This does not refute R1. The language, the sites and the translator's own state of knowledge all
differ from E-20260802e as well, and E-20260804 is descriptive only by its own F1. What it
establishes is that the independence of the yardstick has never been varied deliberately in this
project, and that the one run in which it was varied found no separation. Until a design varies
only that, R1's separation is not known to survive it.
What a repair would look like, so a later session does not have to invent it: re-run
E-20260802e's own six sites, unchanged in every other respect, with the relation statements authored
by a seat that has seen no rendering — and with POSITIVE authored from that statement rather than
before it, which is the mistake E-20260804 made and paid F1 for.
9. Changelog
- v0.1, prediction 1 withheld a FOURTH time and its obstacle relocated a third (2026-08-06,
S121–S122),
ARM-r1-censussteps 1, 1b and 2 — and R1's text is again unchanged.RS-20260806e-published-loss/E-20260806-published-loss; $1.156621729 across two sessions against a declared $1.90; verifier 215 checks, 0 failures, three mutation tests, the outcomes move; 36 of 36 grading bodies, 105 items, cost cross-check closes exactly. The census this prediction has been waiting three sessions for exists — Kleist «Michael Kohlhaas» DE→EN, nine vocative-free asymmetric-address sites chosen from the German alone, two published translations of the whole tale, and King 1914 carries no English device of address at 9 of 9. And the graded limb contradicts the census: King's device-free English is judged to convey the relation at 8 of 9, withPOSITIVE1.000,WRONG0.028, the direction-reversed control 0.037, andF2/F3/F4/F5/F6/F7all silent.P1FAILS for both translators.F1fires at 1 admitted site against a floor of 5 andP2— which is prediction 1 — is WITHHELD.P5holds by counting (8 against 9) at an exact permutationP= 0.500, i.e. on no evidence, and is reported with itsPbeside it. The finding is that losing the device mostly did not lose the relation, because these scenes state the standing in what the speakers say; the one exception,A2, is the site where the speaker's content works against the deference his address claims. §3 prediction 1 gains its fourth status line and the reading that it is MIS-SPECIFIED rather than unlucky. §7's DE→EN row is rewritten and staysuntested. §8 Q-b gains its second measurement — the first with two translators — and the two halves of it point opposite ways. §4 gains its second device-category contrast, from a lead rendering of the whole scene under the device the run had excluded (T-kohlhaas-R20-v2): the categories are alternatives, not ingredients; the pronoun cannot reach the up-address, which is the only place anything was at stake; and a device can assert a register the source does not, which §2's no-added-fact clause does not reach. Nothing in §2 changes and no recommendation is added or withdrawn. - v0.1, prediction 1's obstacle relocated (2026-08-05, S116),
ARM-r1-fresh-pairstep 2, and R1's text is again unchanged.RS-20260805h-content-or-marking/E-20260805h; $0.105467919 against a declared $1.15; verifier 115 checks, 0 failures, three mutation tests, three caught; six of six bodies, 180 judgements, zero seat failures. The run's own gateF3fired at 1 of 6 and its primary is WITHHELD, which amendment A4 recorded before the grading dispatch; every figure it reports is descriptive and three descriptive predictions registered in advance all failed, against the lead's written expectation. What it establishes: theE-20260802eprocedure responds to the MARKING, not to the content — six minimal pairs with every proposition held fixed fell from 0.861 to 0.222, withPOSITIVE36 of 36 andWRONG0 of 36 in the same bodies, andREPEAT— byte-identical spans — disagreeing in 0 of 36 pairs, the first precision measurement this procedure has ever had.RS-20260805c§4(b) is refuted: the unbriefed paraphrase was never marking-free. §3 prediction 1 stays OPEN and UNDISCHARGED, with its obstacle moved from the instrument to the census — no run has yet produced sites where a competent translation lost the marking. §7 gains a PL→EN row,untested. §4 gains the device-category contrast, which is against the release: pronoun-borne and adjective-borne marking are not interchangeable. Two findings about the S111 census travel with it: one of its eight Class A sites (sprzedano) is a site where English has the device, and at another (— to już idź) the relation is entailed by the propositions, so no minimal pair exists. Nothing in §2 changes and no recommendation is added or withdrawn. - v0.1, Q-c answered (2026-08-04, S106),
ARM-r1-warrantstep 1, and R1's text is again unchanged.RS-20260804g-yardstick-holds/E-20260804g-yardstick-repair; $0.548190715; verifier 522 checks, 0 failures, three mutation tests, three caught; cost cross-check closes at −0.000000001. R1's machine-read separation REPRODUCES under a yardstick written by a seat shown no English at all: FORCED 4 of 5, FROZEN 0 of 5, DECOY 0 of 5, d = +4 against the same five sites' 2026-08-02 value of +4. Every registered prediction held and no failure criterion fired. §2 gains the reproduction, and next to it the two things it does not carry — the sites changed hands, and yardstick precision is not separated from yardstick authorship. §2's scope clause is QUALIFIED: refuted for relations about the grammar itself rested on one site whose relation an independent reader states as social and affective, and against that statement R1's rendering scores 3 of 3. §2 gains one site where R1's move is available and insufficient (S2: the courtesy formula carries the source's grammatical function and not the passage's effect, 1 of 3 for the compensation and 0 of 3 for a second re-rendering aimed at the independent statement). §8 Q-c is answered YES and closed, andRS-20260804-yardstick's "the yardstick was the variable" is superseded as the explanation of the FR→EN null. Prediction 1 is still undischarged — these are not fresh sites in a new pair. Nothing in §2's recommendation text changes, and no recommendation is added or withdrawn. - v0.1 (2026-08-02, S091). First release. R1 admitted. The release gate re-read: both limbs
returned, Tier D's failure removes evidence class X3 rather than the release. Coverage stated
without a scalar. Seven pairs declared; everything else
untested. - v0.1 stress-tested (2026-08-04, S101),
ARM-framework-v01step 3, and R1's text is again unchanged.RS-20260804-yardstick/E-20260804-displaced-marking-fr, FR→EN, Maupassant «La Nuit», eight fresh Class A sites, $0.252177609, verifier 217 checks 0 failures, cost cross-check exact. Prediction 1 is neither discharged nor falsified: F1, F2 and F5 fired and the run is descriptive only. Prediction 2 is established as not falsifiable by the named procedure. Prediction 3 held weakly on one text. §8 gains Q-c, which asks whether R1's separation survives a yardstick the lead did not write, because the one run that varied that found none. §7 gains an FR→EN row,untested. Nothing in §2 changes and no recommendation is added or withdrawn — F1 forbids it, and the arm's completion criterion never required a win. - v0.1 completed (2026-08-03, S096),
ARM-framework-v01step 2. Drafted whole. Added: §5, the fourteen candidates written out by name with the reason each is not a recommendation; §6, the coverage section written out, scalar-free, with the 0.867-against-0.289 non-reproduction stated as the reason; §7, the pair-by-pair declaration underD-20260724-04; §8, two open questions with measurements attached; prediction statuses in §3; and one line to §4 — five of R1's six licensed device categories have never been isolated. R1's text is unchanged. The occupancy conditionE-20260803cwas built to add was not added, and the reason is in §8: the run is void by its own gate and the hypothesis failed independently of the void.