Repository path: workshop/experiments/E-20260806-published-loss/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260806-published-loss |
| status | frozen |
| created | 2026-08-06 |
| updated | 2026-08-06 |
| senses | style-correspondence, voice, cultural-mediation, accuracy |
| provisional | true |
| links | wiki/arms/ARM-r1-census.md, framework/v0.1/README.md, wiki/findings/results/RS-20260805h-content-or-marking.md, wiki/findings/results/RS-20260805c-no-loss-to-repair.md, wiki/findings/results/RS-20260804g-yardstick-holds.md, wiki/findings/results/RS-20260802e-displaced-marking.md, workshop/regimes/R04-lead-close.md, workshop/regimes/R20-r1-forced.md, config/models.md, config/budget.md, wiki/base/consulted.md |
E-20260806-published-loss — what two published translators did with Kleist's address forms
ARM-r1-census step 1 (T5). Frozen 2026-08-06 before any panel call. The design is frozen in
the order the page is written: the contamination gate ran first, the sites were selected from the
German alone second, the lead's two renderings were written and committed third, and only then was
any published English read at any site.
1. Question
framework/v0.1 §3 prediction 1 has been withheld three times, and RS-20260805h located the
obstacle in the census: no run has produced sites where a competent translation actually lost the
marking, because every census so far was built out of the lead's own renderings. §8 Q-b names the
way out — ask what published translators do, in evidence class X1a, without any jury.
Q1 (the census, X1a). At sites where the German marks the speaker–addressee relation by an address form English lacks, do two independent published English translations carry the relation?
Q2 (framework/v0.1 prediction 1). At the sites where they do not, does a rendering made under
R1's brief recover it?
Nothing here asks whether any translation is good. Tier D is NOT PASSED; no verdict of quality is sought, made, or reportable.
2. Materials, all public domain and all freely reachable
| Source | Heinrich von Kleist, «Michael Kohlhaas» (1810), German, read in German (charter §4, A2). Project Gutenberg #6645, Ausgewählte Schriften. 34,095 words; stored whole at materials/source-de-full.txt |
| Published A | John Oxenford, Michael Kohlhaas, in Tales from the German, London 1844 (Gutenberg #32046; the volume's contents page attributes this tale to "J. O."). 35,279 words |
| Published B | Frances H. King, Michael Kohlhaas, in The German Classics of the Nineteenth and Twentieth Centuries, vol. IV, New York 1914 (Gutenberg #12060). 38,800 words |
| Scenes | the two long dialogue scenes: Herse's report (1,304 German words) and the Luther interview (1,421 German words). materials/scene-herse-de.txt, materials/scene-luther-de.txt |
Both translations are complete renderings of the whole tale, so the reading is complete rather than
excerpt-limited (charter §7, A8). Logged in wiki/base/consulted.md.
Why DE→EN. framework/v0.1 §7 declares the pair untested: the one run on it,
RS-20260803c, is void by its own gate, and §8 Q-b's only figure — a professional published
translation recovered the relation at one site in twelve — comes from that void run, on one work,
one translator, twelve sites. This run puts the same question to a different work, two translators
and an independent yardstick.
Why this text. Kleist's dialogue runs a four-way address system — du, Ihr, the third-person
Er, Sie — against English's single you, and in the two scenes chosen the asymmetries are
sustained and unmistakable: Luther says du to Kohlhaas throughout while Kohlhaas says Ihr to
Luther throughout, and Kohlhaas says du to his groom Herse while Herse says Ihr back. The
same form does opposite work in the two dyads — Luther's du puts a man beneath him, Kohlhaas's
du keeps a trusted servant close — which is the property no English pronoun has.
3. The contamination gate — run FIRST, and it changed the design
CLAUDE.md's standing rule: measure before choosing, as a selection gate, never as a diagnostic
inside a run. A 163-word probe span of narration (Gegen Mittag kam Herse … stand schon in drei
Stunden vor Erlabrunn), deliberately outside every candidate site and containing no direct speech
and no address form at all, was rendered by the lead from the German alone and put through
tools/dependence_check.py against both published translations of the same span.
| pair | shared 7-grams | 12-grams | 15-grams | longest common run | verdict |
|---|---|---|---|---|---|
| King 1914 ~ lead | 15 | 8 | 5 | 19 tokens | DEPENDENT? |
| King 1914 ~ Oxenford 1844 | 0 | 0 | 0 | 6 | clean |
| lead ~ Oxenford 1844 | 3 | 0 | 0 | 8 | clean |
The shared run is "a village cart kohlhaas sighed deeply at this news he asked whether the horses
had been fed and when" — nineteen consecutive tokens, one proper name in it. Inputs and figures:
materials/gate-*.txt, materials/gate-result.json.
What the gate did to the design, before anything was translated. The lead's rendering may not serve as the independent competent-English baseline the admission condition needs, because it is not independent of one of the two comparators. So the admission condition is carried by the two published translations, which are independent of each other, and the lead's close rendering is retained only as a within-translator contrast against its own forced rendering. This is the standing rule working as a selection gate rather than as a footnote: it removed an arm's job from it.
One prior exposure, declared. Before the gate was run, ~150 words of King's opening and ~60 words
of King's rendering of Luther's first outburst (dein Odem ist Pest) were printed to the console
while locating section boundaries. That outburst is not a site — it was excluded by the site
selector's length condition before the exposure was noticed, and the exclusion is not retro-fitted:
see materials/sites.py, condition (d).
4. Sites — selected from the German alone, by a selector fixed before the spans were read out
The selector, its conditions, its anchor strings and its own machine check are materials/sites.py;
the output is materials/sites.json. A site is one uninterrupted turn of direct speech in the
two scenes such that
- (a) it contains at least one second-person address pronoun directed at the other speaker;
- (b) the dyad's address is asymmetric — the two speakers use different forms for each other across the scene;
- (c) the turn contains no vocative of any kind — no name, no rank noun, no term of abuse — so the relation is carried by the grammatical form and nothing else;
- (d) the span is at least 15 German words.
NULL controls, matched: turns from the same two scenes and the same speakers, 40–120 German words, carrying no second-person address form at all. There is nothing at a NULL site for an English translation to lose, so a translation that fails there is failing the instrument and not the marking.
The selector's own check removed a site the experimenter had chosen, and it is not re-admitted.
Kohlhaas's Wohlan, … Verschafft mir … freies Geleit nach Dresden addresses Luther in the Ihr
form and carries no Ihr pronoun: the marking sits entirely in the imperative ending
(Verschafft against du-form Verschaff). Condition (a) is operationalised on pronouns, so the
site fails it. Loosening (a) after the fact would be choosing a site because of what it contains.
The consequence is registered here rather than discovered later: a pronoun-based census of German
address marking is a lower bound.
One machine flag was adjudicated by reading and recorded rather than suppressed: N3's single
capital Sie is sentence-initial (Sie guckten nun, wie Gänse, aus dem Dach vor) and is the
third-person plural referring to the horses. German orthography makes that form genuinely ambiguous
to a regex; the adjudication is in materials/sites.py, ADJUDICATED.
The 12 sites.
| id | stratum | scene | speaker → addressee | form | German words |
|---|---|---|---|---|---|
| A1 | asym | Luther | Luther → Kohlhaas | du |
40 |
| A2 | asym | Luther | Kohlhaas → Luther | Ihr |
70 |
| A4 | asym | Luther | Luther → Kohlhaas | du |
36 |
| A5 | asym | Luther | Luther → Kohlhaas | du |
91 |
| A6 | asym | Luther | Luther → Kohlhaas | du |
26 |
| A7 | asym | Herse | Kohlhaas → Herse | du |
15 |
| A8 | asym | Herse | Kohlhaas → Herse | du |
31 |
| A9 | asym | Herse | Herse → Kohlhaas | Ihr |
63 |
| A10 | asym | Herse | Kohlhaas → Herse | du |
17 |
| N1 | null | Luther | Kohlhaas → Luther | — | 62 |
| N2 | null | Herse | Herse → Kohlhaas | — | 70 |
| N3 | null | Herse | Herse → Kohlhaas | — | 72 |
Registered as a limit before the run: the nine asymmetry sites are four dyad directions but only two scenes of one work, and Luther→Kohlhaas supplies four of the nine.
5. Arms
| arm | what it is | written when |
|---|---|---|
OXEN |
John Oxenford's 1844 English at the site | 1844 |
KING |
Frances H. King's 1914 English at the site | 1914 |
CLOSE |
the lead's close rendering, regime R04, no marking brief, written from the German alone before any published English was read at any site |
stage 0 |
FORCE |
the lead's rendering under R1's brief, regime R20, same condition |
stage 0 |
POSITIVE |
an explicit English statement of the relation, authored from the yardstick statement | stage 2 |
WRONG |
an explicit English statement of a different relation | stage 2 |
NULL sites carry OXEN, KING, CLOSE, POSITIVE, WRONG and no FORCE — there is no
marking to force. 69 items: 9 × 6 + 3 × 5.
FORCE uses no archaic pronoun, and that is a registered choice, not an oversight. English does
possess a second-person pronoun contrast — thou against you — and R1's device list opens with
a pronoun. It is excluded here for a reason that is stated so a later session can overturn it:
the device is one-sided. Thou can mark the down-address (du) and there is no English pronoun
that marks the up-address (Ihr), because you is the modern default; a rendering using it would
mark four dyad-directions in one direction only and would change the register of the whole scene
rather than the site. FORCE therefore uses the other five categories R1 names — address noun,
courtesy formula, verb choice, syntax, the argument structure of the clause. What this run does
not test is the pronoun device, and §7 says so.
6. Procedure, in the order it was run
Stage 0 — translation (no API, no ledger entry). Both scenes rendered whole, twice, by the lead
from the German alone: R04 close (T-kohlhaas-R04-v1) and R20 forced (T-kohlhaas-R20-v1).
Translator's logs written and frozen before any evaluation was designed and before any published
English was read at any site (charter §3, A4; continue-prompt.md §5).
Stage 1 — the yardstick. One seat, shown the German span and nothing else — no English of any
kind, no gloss, no log, no description of the study — writes, for each of the 12 sites, one or two
sentences stating the relation the speaker's manner of address establishes between the two people.
Ordering matters and is the mistake E-20260804 paid F1 for: POSITIVE is authored from this
statement, never before it.
Leak screen, mandatory (E-20260804g §4). A statement naming the source's grammatical device
tells the graders the answer. Screened words, applied mechanically to the returned statements:
du, Sie, Ihr, Euch, pronoun, second person, second-person, formal, informal,
familiar form, polite form, address form, T-V, grammatical, verb ending, German.
A flagged statement is regenerated once, with the screen restated. A statement that fails twice
takes its site out of the run.
Stage 2 — the control texts. POSITIVE and WRONG are written by the lead from the yardstick
statements: POSITIVE states the yardstick's relation outright in English; WRONG states, equally
outright, a different relation of the same kind. Neither is a translation.
Stage 3 — the parity screen (S116 constraint, RS-20260805h §3). Two seats, neither a
grader, see CLOSE and FORCE for each asymmetry site unlabelled, with no relation statement, no
source and no description of the study, and are asked only whether the two passages state the same
facts, judging speech for propositional content and ignoring exact wording, tone, attitude, warmth,
sympathy and emphasis. (The wording is RS-20260805h A1's, adopted verbatim, because a screen that
asks about "the same words" fails the very difference under test.) A site where the two seats do not
both answer SAME leaves P2's denominator: R1's English device has asserted something the German did
not, which is framework/v0.1 §4's standing warning.
Stage 4 — grading, blind. Three seats × two hash-fixed orderings × three blocks. Each item is one English passage and one relation statement; the seat answers YES / NO: does this English convey that relation between the two speakers? Seats see no German, no arm labels, no authorship, no other arm, and no description of the study. Judgement is never parallelised across seats within a body.
7. Predictions, registered before any dispatch
| # | prediction | basis |
|---|---|---|
| P1 (primary; Q1, the census) | OXEN and KING each recover the relation at half or fewer of the nine asymmetry sites |
§8 Q-b's 0.111 on a void run; the release's own claim that English lacks the category |
P2 (Q2; this is framework/v0.1 prediction 1) |
on the admitted sites — those where both published translations fail — FORCE recovers at more than half. The failure it should not survive: at or below a third |
RS-20260802e 5 of 6 |
| P3 (control) | at the three NULL sites, OXEN and KING each recover at ≥ 2 of 3 |
there is nothing there to lose |
| P4 (instrument) | POSITIVE ≥ 0.90 per-judgement, WRONG ≤ 0.10 |
RS-20260805h: 1.000 and 0.000 |
| P5 (descriptive, registered against the lead's expectation) | CLOSE recovers at fewer sites than FORCE |
S111 found the opposite and it is the reason this arm exists |
The lead's written expectation, registered so it can be wrong: that the two published
translations will differ from each other — Oxenford, writing in 1844 with thou still live in his
literary register, is expected to mark more often than King in 1914 — and that FORCE will recover
at most of the admitted sites.
8. Failure criteria, registered before any dispatch
| # | fires when | consequence |
|---|---|---|
| F1 | fewer than 5 sites are admitted (both published translations recover) | P2 is WITHHELD. The census is empty again, and this time against published practice rather than the lead's own prose — which is a stronger negative and is reported as one |
| F2 | POSITIVE < 0.75 or WRONG > 0.25 |
the whole run is descriptive only |
| F3 | more than 3 of 12 yardstick statements fail the leak screen twice | the run is withheld; the yardstick is not fit |
| F4 | fewer than 5 admitted sites survive the parity screen | P2 is WITHHELD |
| F5 | OXEN or KING recovers at fewer than 2 of 3 NULL sites |
P1 is descriptive only — the instrument is not passing ordinary competent English, so a low published-translation rate at the asymmetry sites is not attributable to the marking |
| F6 | the two orderings disagree on more than 20% of items within a seat | position effects dominate; all primaries withheld |
9. What this run cannot establish, written before it is run
- Nothing about quality. No arm is scored against another for goodness; Tier D is NOT PASSED and
every sentence of the result is
provisional. - Nothing about human readers. The graders are panel models. A YES licenses three independent seats recovered the relation from this English and nothing about a human reader.
- Nothing about
CLOSEas an independent baseline. It is measured at 19 shared tokens againstKING, and it is in the run as a within-translator contrast only. - Nothing about the pronoun device, excluded by §5 with its reason.
- Nothing about German address marking carried in the verb, excluded by the selector, §4.
- One work, two scenes, two dyads, nine sites, and Luther→Kohlhaas supplies four of them.
- The relation statements are one seat's reading of the German. A second yardstick author is not
in this design's budget;
RS-20260804gestablished that changing the author does not change the separation on the sites it tested, and that is the only evidence on the point.
10. Budget
Pre-flight worst case built from max_tokens and not from expected output (note (abc)), with a
routing margin for the provider spread config/models.md records:
| stage | calls | worst case |
|---|---|---|
| pre-run adversarial critic | 1 | $0.30 |
| yardstick + one regeneration | 2 | $0.10 |
| parity screen | 2 | $0.10 |
| grading, 3 seats × 2 orderings × 3 blocks | 18 | $0.70 |
| total | 23 | $1.20 |
Against $3.568724493 of headroom on UTC day 2026-08-06. Lead translation is free and is never ledgered (charter §3, A4).
11. Amendments, made after the pre-run critic and BEFORE any grading call
critic.md, one pass, moonshotai/kimi-k3. Verdict NEEDS-AMENDMENT, three BLOCKING
findings and nine notes. All three BLOCKING findings are accepted in full, and six of the nine
notes are acted on. Nothing below was written after a grading datum existed; the yardstick had not
been called when these were written.
A1 (from BLOCKING B1) — the judgement→site aggregation rule, which the frozen design omitted and
every primary sits on. Registered here: a site recovers for an arm iff at least 4 of its 6
judgements (3 seats × 2 orderings) are YES. Admission for P2: both OXEN and KING fail
that rule at the site. P4 stays per-judgement, as written.
A2 (from BLOCKING B2) — the mismatched-statement control, which the frozen design lacked. Without
it, "OXEN recovers at k of 9" is compatible with seats that say YES to any fluent passage paired
with any plausible relation statement, and the leak screen pushes statements toward content-level
descriptions these denunciation scenes satisfy with no marking at all. Added: each of OXEN,
KING, CLOSE and FORCE at each of the nine asymmetry sites is also graded against another
site's statement, so 36 further items. The pairing is fixed here and is a direction reversal
in every case, which is the sharpest available form: within each scene, every down-address site
takes the up-address site's statement, and the up-address site takes a down-address one — Luther
scene A1, A4, A5, A6 → A2 and A2 → A1; Herse scene A7, A8, A10 → A9 and A9 → A7.
New failure criterion F7: if MISMATCH exceeds 0.25 per-judgement pooled, every translation-arm
rate in this run is descriptive only. Limit registered with it: a reversed statement differs from
its site's own in content as well as in direction, so a NO is not attributable to direction alone.
The perfectly matched control — same content, reversed relation — does not exist in these materials.
A3 (from BLOCKING B3) — three sentences of the frozen design are wrong and are corrected here
rather than quietly.
1. §9's licence sentence is withdrawn and replaced. Seats are handed the relation and asked
whether the English conveys it; nobody recovers anything. A YES licenses only: this seat judged
this English to convey the stated relation. The word recover is retained as the run's label
for that event and means nothing more.
2. F1's consequence is scoped. If the census is empty, the report says empty under this
instrument, with these statements — not published practice carries the marking. An empty census
produced by content-satisfiable statements would say nothing about the marking at all.
3. P2's claim is scoped to the relation-types actually present in the admitted set. If the
admitted sites turn out to exclude one direction of address, P2 says nothing about that
direction.
A4 (from N4) — block construction, unregistered in the frozen design. Six blocks, built so that
the six arms of any one site fall in six different blocks, and so that a passage's matched and
mismatched items never share a block. Item order inside a block is hash-fixed, two orderings.
Consequence: 36 grading calls, not 18, and §10's grading line becomes $1.40, total worst case
$1.90 against $3.568724493 of headroom. The frozen F6 is unchanged and still applies.
A5 (from N1) — the yardstick's role-direction check. Each returned statement is printed beside its
site's speaker and addressee and adjudicated for direction before stage 2. A statement that assigns
the roles the wrong way round takes its site out of the run; the adjudication is recorded in
materials/statements-check.md whether or not it removes anything. It is a reading, not a machine
test, and is recorded as such.
A6 (from N3, N5, N7, N8, N9) — five corrections carried into the report. (i) The parity screen
tells its seats to ignore the channel FORCE manipulates, so F4 passing shows little and is
not evidence that FORCE marks. (ii) The three grading seats are three different models from
three different labs (config/models.md P1, P2, P5), which is what "independent" claims here.
(iii) The dependence verdict covers narration; CLOSE and FORCE are dialogue and may be
King-tinged, and every occurrence of "written from the German alone" carries the declared-exposure
asterisk of §3. (iv) F1's parenthetical gloss in the frozen table reads "both published
translations recover" and should read "both published translations recover at the site, so it is
not admitted". (v) POSITIVE and WRONG are authored by the lead after it has read OXEN and
KING at every site; they are derived from the yardstick statement by rule, and the ordering is
disclosed here and in the result.
A7 — the critic's one error of fact, corrected because provenance is the thing at issue. The
critique states that this page "was frozen after published English was read", and infers that §7's
expectation about Oxenford is arm-informed. It is not. Commit 75cd449 contains this design,
both lead renderings and §7's expectation, and its materials/lead-renderings.json carries only
the CLOSE and FORCE arms — the published spans were added in a later commit. §7's expectation
that Oxenford will mark more often than King was written before a word of either translation's
dialogue had been read, and is registered. Everything else the critic says about the thou
observation is accepted without qualification: the §5 exclusion stands as registered and its
rationale may not be retro-fitted; the distribution of thou across the arms is descriptive;
no thou-conditioned recovery analysis is confirmatory; and the claim that Oxenford used thou
in order to mark the down-address is mind-reading a dead translator and will not be asserted.
Registered before the yardstick call, in light of the arms having been read: the lead now expects
P1 to fail for OXEN and hold for KING, and expects F1 to fire — because the four
Luther→Kohlhaas sites are the ones Oxenford marked, leaving at most five admitted. This expectation
is arm-informed and is therefore descriptive, not registered as a prediction. The frozen P1,
P2, P3, P4, P5 are unchanged and are what the run is scored on.