Repository path: workshop/experiments/E-20260802e-displaced-marking/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260802e-displaced-marking |
| status | frozen |
| created | 2026-08-02 |
| updated | 2026-08-02 |
| senses | style-correspondence, accuracy, voice, cultural-mediation, naturalness |
| provisional | true |
| links | workshop/experiments/E-20260802e-displaced-marking/materials/census.md, wiki/arms/ARM-framework-v01.md, framework/closure.md, framework/traceability-inventory.md, wiki/findings/results/RS-20260728j-classb-marking.md, wiki/findings/results/RS-20260802d-class-line-carry.md, wiki/findings/results/RS-20260729d-decision-grain.md, config/models.md, config/budget.md, wiki/goodness-senses.md |
E-20260802e — is "the target cannot mark this" a reliable judgment?
ARM-framework-v01 step 1 (T5). Frozen before any rendering was written, before any seat was
addressed, and before the census's classification was shown to anyone.
1. The question, and why it is not a question about this project
A translator working from a source that marks something grammatically — a polite pronoun, an honorific auxiliary, a diminutive suffix, a dative of experience — reaches a site where the target has no such category, writes "English has no way to do this", and records the loss. Is that judgment reliable?
It has been tested twice in this project's history, both times by opening an independent published translation of the same work, and both times it was wrong:
- S052 (
RS-20260728j): Verga's out-of-sequence conditionals. The frozen log said "there is nothing to repair with — English has no marked conditional to reach for." True of the category. False of the marking: a forced re-rendering marked all three sites by changing grammatical category, and Dole 1896 independently reached the same construction at two of them. - S090 (
RS-20260802d): Canth's honorific plural from a child to her mother. The frozen register said the erasure had "no compensation available that would not be a fabrication", and six cold machine attempts marked nothing. Hertzberg's 1886 Swedish — a language that has the pronominal honorific — dropped it and wrote "Blir mamma länge borta?", carrying the deference on a kin term, a device English also has.
Two instances are two instances. This experiment tests the remaining six sites in the project's entire record, across four source languages, and so can answer at the level of a population rather than an anecdote.
What it teaches about translating literature, in one sentence (the subject rule,
wiki/tracks.md): whether a translator's report of an unmarkable site is a fact about the language
pair or a fact about where the translator stopped looking — and therefore whether "before
recording the loss, look in another grammatical category" is advice worth giving anyone.
2. The candidate under test
R1 (displaced marking). Where the source marks a relation or attitude by a grammatical form the target lacks, the absence of a same-category counterpart is not the absence of the marking. Before recording the loss, render the site again under a brief that requires the marking to appear, letting it fall in any category.
R1 is the project's first prescriptive, translator-addressed candidate since C15, which was
written at S056 and refused — not for want of evidence but because two independent readers
applying it to 23 sites agreed at κ 0.452 while agreeing at κ 0.933 on a deliberately groundless
rule of the same shape (RS-20260729d). framework/closure.md §6 item 3 left the open question:
is that trade a property of warrant or of this rule's wording? C15's conditions were
evaluative (load-bearing or furniture; an exact equivalent). R1's are grammatical and checkable.
§5 below is the direct test, on the same failure criterion C15 was refused under.
3. Materials
materials/census.md, frozen: 126 unavailability claims → 61 naming a grammatical category → 8
Class A sites (marking declared wholly lost), of which 2 are already tested and refuted. The six
untested Class A sites are the test population, and they are the whole of it.
| site | pair | source device | frozen rendering |
|---|---|---|---|
| S1 | JA→EN | stacked humbling auxiliaries 拝見させて頂きたい |
should like to be allowed to look at it for a moment |
| S2 | RU→EN | formal вы between two lovers |
"What are you crying about?" I asked |
| S3 | RU→EN | dative-impersonal страшно Семёну |
Semyon is afraid |
| S4 | RU→EN | diminutives землица, юртёнки |
a scrap of frozen ground; little smoky yurts |
| S5 | LZH→EN | humble 1sg 僕/仆 |
I |
| S6 | ES→EN | gender agreement creaciones → hijas |
the children of his delirium |
Contamination. No published comparator is opened and none is needed: the claim under test is
existential — does a marked rendering exist — and a rendering that reproduces a published
translator's device is still an existence proof. Per CLAUDE.md's standing rule the absence of a
reachable comparator is declared on the artifact rather than passed over. Note (bhb) cuts the
other way here and is the design's friend: the lead matches itself at up to 37 contiguous tokens
across sessions, so a forced re-rendering that simply reproduces the frozen one will mark nothing
and count against R1. The known bias runs toward the null.
4. Procedure — stage 1, does a marked rendering exist
- The lead writes four English renderings per site, all frozen and committed before any seat is addressed: - FROZEN — the filed rendering, verbatim from the translation page. - FORCED — a new rendering under R1's brief: the marking must appear; it may fall in any grammatical category. Written from the source and the frozen log alone. - DECOY — a rendering that differs from FROZEN in wording but attempts no marking, and is length-matched to FORCED within 15% (verifier-checked). This is the control that can kill R1: it catches a grader who scores any changed or longer rendering as marked. S041/S046 measured this jury's length response and found it graded, not sign-only. - POSITIVE — an explicit metalinguistic gloss that unmistakably states the relation. If a grader misses this, the grader is not reading.
- Presentation order per site is fixed by
sha256(site id | seat id), not chosen by the lead, and each seat sees a different rotation. - Three non-Anthropic seats (P1
openai/gpt-5.6-terra, P2google/gemini-3.6-flash, P3x-ai/grok-4.5) receive: the source text, a literal gloss, the relation stated neutrally with the source's device never named, and the four renderings unlabelled. For each rendering: does an English reader with no access to the source get this relation from it — YES / NO, plus one line of reason. - Two orderings per seat (the rotation and its reverse) — six bodies. Slot preference was measured at 0.500–0.600 on this panel at S020 and is controlled by rotation, not assumed away.
5. Procedure — stage 2, is R1's antecedent followable
C15 died here, and R1 is not exempt. Eighteen frozen log excerpts — the 8 Class A sites, the 6
Class B sites, and 4 lexical-gap distractors that pass filter 1 and fail filter 2 — are given to the
same three seats, in an order fixed by hash, with the classification rule of census.md quoted
verbatim and no examples. Each seat returns A / B / NEITHER per excerpt.
Measured: three-way agreement with the lead's frozen census, and pairwise Cohen's κ between seats. The bar is pre-registered at κ ≥ 0.60, on the stated ground that a rule the project would adopt must be followed more consistently than the rule it already refused (C15: 0.452).
6. Predictions — registered, five, before any call
| # | prediction |
|---|---|
| P1 | (primary) At ≥ 4 of 6 sites the FORCED rendering is graded YES by ≥ 2 of 3 seats |
| P2 | At ≥ 5 of 6 sites the FROZEN rendering is graded NO by ≥ 2 of 3 seats — the logs' own loss claims reproduce |
| P3 | DECOY is graded YES by ≥ 2 of 3 seats at ≤ 1 site |
| P4 | POSITIVE is graded YES by ≥ 2 of 3 seats at 6 of 6 sites |
| P5 | Stage 2 pairwise κ ≥ 0.60 |
The lead's stated expectation, recorded so it can be wrong: P1 holds at 4 or 5 of 6, with S6 (grammatical gender) the most likely failure — "daughters" states the femininity but cannot state that it is agreement rather than choice, which is what the log actually claimed was lost. If S6 marks, the lead was wrong about the hardest case; if S1–S5 all fail, R1 is dead.
7. Failure criteria — pre-committed
- F1 — the grading instrument failed. POSITIVE graded NO at more than 1 site, or DECOY graded YES at more than 1 site. → the whole run is descriptive only; no candidate enters the framework on it, whatever P1 says.
- F2 — the census misread the logs. FROZEN graded YES at ≥ 3 of 6 sites. → P1 is uninterpretable (the sites were not losses to begin with) and the census's classification is the finding.
- F3 — admission. R1 enters
framework/v0.1/as an evidenced recommendation only if P1 holds AND F1 does not fire AND P5's κ ≥ 0.60. If P1 holds and κ fails, R1 enters markeduntested, with C15's verdict recorded as reproduced on a second rule — which would be the stronger finding, because it would answerclosure.md§6.3 in the direction that costs the project a release. - A null is a result. If R1 fails, the two prior refutations become two lucky sites,
V14aandE7stand as local errata, and the honest sentence is that the project has no prescriptive recommendation. That sentence goes intoframework/v0.1/and the arm closes on it.
8. What this design cannot show
- A marked rendering existing does not make it better. Every marking has a cost — S052 measured
three different ones at three sites (hypothetical → assertion, counterfactual → general truth,
certainty → supposition), all scoring as unlicensed addition under
D-20260727-08. This run measures availability, never quality, and Tier D is NOT PASSED so no quality claim is admissible anyway. - Three language models are not a survey of English readers (charter §4). A YES licenses three independent readers got the relation from this text; it does not license English marks this.
- The lead wrote all four renderings. Blinding is by presentation, not by authorship. The DECOY is the control on that, and F1 makes it binding rather than decorative.
- Six sites is the population, not a sample, so no sampling inference is available and none is drawn: the figure is a count over the project's own record.
9. Verification
analysis/verify.py, importing nothing from tools/: recomputes the census from the repository
(all three filters, the 126/61/13 counts), every quoted source string verbatim against the stored
files, the FORCED/DECOY length match at every site, the hash-derived rotations, all five predictions
and all three failure criteria from the raw stored bodies, κ by independent implementation, and the
billed total. At least three mutation tests, each asserting the bytes on disk changed and then
restoring.
10. Budget
Declared worst case $1.20, built from max_tokens and not from expected output (note (abc)),
with a 4× routing margin (note: the S022 measurement, and P2's reasoning tax measured three times —
note (bhf)).
| stage | calls | max_tokens |
worst case |
|---|---|---|---|
| pre-run critic | 1 | 24,000 | $0.20 |
| stage 1 — grading, 3 seats × 2 orderings | 6 | 4,000 | $0.55 |
| stage 2 — followability, 3 seats | 3 | 4,000 | $0.25 |
| retry reserve | — | — | $0.20 |
Today's headroom at session open: $2.915701162 of the $5.00 UTC cap, five sessions already run. The worst case fits with $1.7 to spare. Lead translation is free and is never ledgered.
11. Amendments, 2026-08-02, after the pre-run critic (critic.md) — all four accepted
A1 (BLOCKING). New stage 1b: P4 moonshotai/kimi-k3, a seat used nowhere else in this run,
rates FROZEN / FORCED / DECOY at all six sites for naturalness as English prose only, hash-
ordered, with no mention of relations, marking or sources. New failure criterion F4: if mean
DECOY naturalness is more than 1.0 below mean FORCED, the DECOY is not a fair control, F1 cannot be
trusted, and the run is descriptive only.
A2 (BLOCKING). (i) The loss filter was re-run over all 126 windows, not the 61, to surface
any Class A site whose window failed the grammatical lexicon. It found 19 windows, 6 of them new,
and no new Class A site with a grammatical marking: two are new Class B entries (koyhaa D42
subjectless verbs → fronted adverbial; levsha #8 derivational pattern → periphrasis, both now in
stage 2), one is a Class A site whose marking is typographic, not grammatical
(unsu-choun-nal #8, hanja glosses — outside R1's antecedent and recorded, not tested), and three
are NEITHER. (ii) The population claim is downgraded everywhere downstream. Not "the entire
record" — "every Class A site this project has identified, by mechanical extraction plus a manual
audit of all 126 windows; recall against sites phrased in ways no filter catches is not
established."
A3. F3 is re-cut so the two kinds of evidence are never summed. R1's warrant is human-anchored at two sites (Dole 1896, Hertzberg 1886 — published translators who made the move independently, read against their sources) and machine-read at the six tested here. A YES licenses three independent readers recovered the relation from this English and nothing about human readers. The release states the two numbers separately. A human reader is not available to this project — Tom is never an experimental subject (charter §9) — and that is stated as a limit, not talked around.
A4. DECOY must differ from FROZEN by ≥ 20% of tokens as well as matching FORCED within 15% on length. Both are computed on the span, because the design holds the surrounding context byte-identical across all four conditions so that only the span varies. Checked before any seat was addressed: all six sites pass (length gaps 0.000–0.133; DECOY-vs-FROZEN change 0.250–0.545). S2's first DECOY failed at 0.190 and was rewritten; the failure and the rewrite are recorded rather than the threshold moved.
Stage 2 grows from 18 items to 20 (8 Class A · 8 Class B · 4 lexical distractors) on A2(i).
Budget after amendments. One added call (stage 1b, max_tokens 4,000, P4 at $3.00/$15.00 per M
— the panel's most expensive seat). Declared worst case $1.20 → $1.40, still well inside the
day's $2.915701162 headroom.
12. Amendment A5 — mid-run, declared, and it is a capacity fix rather than an analysis choice
P2 google/gemini-3.6-flash returned finish_reason: length on all four stage-1 attempts (two
orderings × two tries) at max_tokens 4,000, spending 4,177–5,691 characters of hidden reasoning
on a task with no reasoning in it before reaching the answer lines. Under note (b) a length
body is a seat failure, never a partial answer, so all four were rejected by the runner.
Note (bhf) — measure a seat's appetite before depending on it — fires for the fourth recorded time, on the same slug S090 measured at 4–6× P1's cost for the same passage. It was read as a cost note and it is also a capacity note; that is the part this run adds.
What A5 changes: max_tokens for P2 alone, 4,000 → 16,000. The prompt is re-composed by the
same functions from the same frozen materials with the same hash-derived rotations and is
byte-identical; the question, the conditions and the scoring are untouched. Nothing about what is
measured changes.
What is recorded rather than absorbed. The four failed bodies billed $0.033369 + $0.0333675 +
$0.033369 + $0.0333675 = $0.1334730 and are kept on disk with their metadata. They are ledgered
in config/budget.md as spend against this session, because they were. A run that quietly
re-dispatches until it gets a body is not reporting its own cost.