Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260802e-displaced-marking/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260802e-displaced-marking
statusfrozen
created2026-08-02
updated2026-08-02
sensesstyle-correspondence, accuracy, voice, cultural-mediation, naturalness
provisionaltrue
linksworkshop/experiments/E-20260802e-displaced-marking/materials/census.md, wiki/arms/ARM-framework-v01.md, framework/closure.md, framework/traceability-inventory.md, wiki/findings/results/RS-20260728j-classb-marking.md, wiki/findings/results/RS-20260802d-class-line-carry.md, wiki/findings/results/RS-20260729d-decision-grain.md, config/models.md, config/budget.md, wiki/goodness-senses.md

E-20260802e — is "the target cannot mark this" a reliable judgment?

ARM-framework-v01 step 1 (T5). Frozen before any rendering was written, before any seat was addressed, and before the census's classification was shown to anyone.

1. The question, and why it is not a question about this project

A translator working from a source that marks something grammatically — a polite pronoun, an honorific auxiliary, a diminutive suffix, a dative of experience — reaches a site where the target has no such category, writes "English has no way to do this", and records the loss. Is that judgment reliable?

It has been tested twice in this project's history, both times by opening an independent published translation of the same work, and both times it was wrong:

Two instances are two instances. This experiment tests the remaining six sites in the project's entire record, across four source languages, and so can answer at the level of a population rather than an anecdote.

What it teaches about translating literature, in one sentence (the subject rule, wiki/tracks.md): whether a translator's report of an unmarkable site is a fact about the language pair or a fact about where the translator stopped looking — and therefore whether "before recording the loss, look in another grammatical category" is advice worth giving anyone.

2. The candidate under test

R1 (displaced marking). Where the source marks a relation or attitude by a grammatical form the target lacks, the absence of a same-category counterpart is not the absence of the marking. Before recording the loss, render the site again under a brief that requires the marking to appear, letting it fall in any category.

R1 is the project's first prescriptive, translator-addressed candidate since C15, which was written at S056 and refused — not for want of evidence but because two independent readers applying it to 23 sites agreed at κ 0.452 while agreeing at κ 0.933 on a deliberately groundless rule of the same shape (RS-20260729d). framework/closure.md §6 item 3 left the open question: is that trade a property of warrant or of this rule's wording? C15's conditions were evaluative (load-bearing or furniture; an exact equivalent). R1's are grammatical and checkable. §5 below is the direct test, on the same failure criterion C15 was refused under.

3. Materials

materials/census.md, frozen: 126 unavailability claims → 61 naming a grammatical category → 8 Class A sites (marking declared wholly lost), of which 2 are already tested and refuted. The six untested Class A sites are the test population, and they are the whole of it.

site pair source device frozen rendering
S1 JA→EN stacked humbling auxiliaries 拝見させて頂きたい should like to be allowed to look at it for a moment
S2 RU→EN formal вы between two lovers "What are you crying about?" I asked
S3 RU→EN dative-impersonal страшно Семёну Semyon is afraid
S4 RU→EN diminutives землица, юртёнки a scrap of frozen ground; little smoky yurts
S5 LZH→EN humble 1sg 僕/仆 I
S6 ES→EN gender agreement creaciones → hijas the children of his delirium

Contamination. No published comparator is opened and none is needed: the claim under test is existential — does a marked rendering exist — and a rendering that reproduces a published translator's device is still an existence proof. Per CLAUDE.md's standing rule the absence of a reachable comparator is declared on the artifact rather than passed over. Note (bhb) cuts the other way here and is the design's friend: the lead matches itself at up to 37 contiguous tokens across sessions, so a forced re-rendering that simply reproduces the frozen one will mark nothing and count against R1. The known bias runs toward the null.

4. Procedure — stage 1, does a marked rendering exist

  1. The lead writes four English renderings per site, all frozen and committed before any seat is addressed: - FROZEN — the filed rendering, verbatim from the translation page. - FORCED — a new rendering under R1's brief: the marking must appear; it may fall in any grammatical category. Written from the source and the frozen log alone. - DECOY — a rendering that differs from FROZEN in wording but attempts no marking, and is length-matched to FORCED within 15% (verifier-checked). This is the control that can kill R1: it catches a grader who scores any changed or longer rendering as marked. S041/S046 measured this jury's length response and found it graded, not sign-only. - POSITIVE — an explicit metalinguistic gloss that unmistakably states the relation. If a grader misses this, the grader is not reading.
  2. Presentation order per site is fixed by sha256(site id | seat id), not chosen by the lead, and each seat sees a different rotation.
  3. Three non-Anthropic seats (P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5) receive: the source text, a literal gloss, the relation stated neutrally with the source's device never named, and the four renderings unlabelled. For each rendering: does an English reader with no access to the source get this relation from it — YES / NO, plus one line of reason.
  4. Two orderings per seat (the rotation and its reverse) — six bodies. Slot preference was measured at 0.500–0.600 on this panel at S020 and is controlled by rotation, not assumed away.

5. Procedure — stage 2, is R1's antecedent followable

C15 died here, and R1 is not exempt. Eighteen frozen log excerpts — the 8 Class A sites, the 6 Class B sites, and 4 lexical-gap distractors that pass filter 1 and fail filter 2 — are given to the same three seats, in an order fixed by hash, with the classification rule of census.md quoted verbatim and no examples. Each seat returns A / B / NEITHER per excerpt.

Measured: three-way agreement with the lead's frozen census, and pairwise Cohen's κ between seats. The bar is pre-registered at κ ≥ 0.60, on the stated ground that a rule the project would adopt must be followed more consistently than the rule it already refused (C15: 0.452).

6. Predictions — registered, five, before any call

# prediction
P1 (primary) At ≥ 4 of 6 sites the FORCED rendering is graded YES by ≥ 2 of 3 seats
P2 At ≥ 5 of 6 sites the FROZEN rendering is graded NO by ≥ 2 of 3 seats — the logs' own loss claims reproduce
P3 DECOY is graded YES by ≥ 2 of 3 seats at ≤ 1 site
P4 POSITIVE is graded YES by ≥ 2 of 3 seats at 6 of 6 sites
P5 Stage 2 pairwise κ ≥ 0.60

The lead's stated expectation, recorded so it can be wrong: P1 holds at 4 or 5 of 6, with S6 (grammatical gender) the most likely failure — "daughters" states the femininity but cannot state that it is agreement rather than choice, which is what the log actually claimed was lost. If S6 marks, the lead was wrong about the hardest case; if S1–S5 all fail, R1 is dead.

7. Failure criteria — pre-committed

8. What this design cannot show

9. Verification

analysis/verify.py, importing nothing from tools/: recomputes the census from the repository (all three filters, the 126/61/13 counts), every quoted source string verbatim against the stored files, the FORCED/DECOY length match at every site, the hash-derived rotations, all five predictions and all three failure criteria from the raw stored bodies, κ by independent implementation, and the billed total. At least three mutation tests, each asserting the bytes on disk changed and then restoring.

10. Budget

Declared worst case $1.20, built from max_tokens and not from expected output (note (abc)), with a 4× routing margin (note: the S022 measurement, and P2's reasoning tax measured three times — note (bhf)).

stage calls max_tokens worst case
pre-run critic 1 24,000 $0.20
stage 1 — grading, 3 seats × 2 orderings 6 4,000 $0.55
stage 2 — followability, 3 seats 3 4,000 $0.25
retry reserve — — $0.20

Today's headroom at session open: $2.915701162 of the $5.00 UTC cap, five sessions already run. The worst case fits with $1.7 to spare. Lead translation is free and is never ledgered.


11. Amendments, 2026-08-02, after the pre-run critic (critic.md) — all four accepted

A1 (BLOCKING). New stage 1b: P4 moonshotai/kimi-k3, a seat used nowhere else in this run, rates FROZEN / FORCED / DECOY at all six sites for naturalness as English prose only, hash- ordered, with no mention of relations, marking or sources. New failure criterion F4: if mean DECOY naturalness is more than 1.0 below mean FORCED, the DECOY is not a fair control, F1 cannot be trusted, and the run is descriptive only.

A2 (BLOCKING). (i) The loss filter was re-run over all 126 windows, not the 61, to surface any Class A site whose window failed the grammatical lexicon. It found 19 windows, 6 of them new, and no new Class A site with a grammatical marking: two are new Class B entries (koyhaa D42 subjectless verbs → fronted adverbial; levsha #8 derivational pattern → periphrasis, both now in stage 2), one is a Class A site whose marking is typographic, not grammatical (unsu-choun-nal #8, hanja glosses — outside R1's antecedent and recorded, not tested), and three are NEITHER. (ii) The population claim is downgraded everywhere downstream. Not "the entire record" — "every Class A site this project has identified, by mechanical extraction plus a manual audit of all 126 windows; recall against sites phrased in ways no filter catches is not established."

A3. F3 is re-cut so the two kinds of evidence are never summed. R1's warrant is human-anchored at two sites (Dole 1896, Hertzberg 1886 — published translators who made the move independently, read against their sources) and machine-read at the six tested here. A YES licenses three independent readers recovered the relation from this English and nothing about human readers. The release states the two numbers separately. A human reader is not available to this project — Tom is never an experimental subject (charter §9) — and that is stated as a limit, not talked around.

A4. DECOY must differ from FROZEN by ≥ 20% of tokens as well as matching FORCED within 15% on length. Both are computed on the span, because the design holds the surrounding context byte-identical across all four conditions so that only the span varies. Checked before any seat was addressed: all six sites pass (length gaps 0.000–0.133; DECOY-vs-FROZEN change 0.250–0.545). S2's first DECOY failed at 0.190 and was rewritten; the failure and the rewrite are recorded rather than the threshold moved.

Stage 2 grows from 18 items to 20 (8 Class A · 8 Class B · 4 lexical distractors) on A2(i).

Budget after amendments. One added call (stage 1b, max_tokens 4,000, P4 at $3.00/$15.00 per M — the panel's most expensive seat). Declared worst case $1.20 → $1.40, still well inside the day's $2.915701162 headroom.

12. Amendment A5 — mid-run, declared, and it is a capacity fix rather than an analysis choice

P2 google/gemini-3.6-flash returned finish_reason: length on all four stage-1 attempts (two orderings × two tries) at max_tokens 4,000, spending 4,177–5,691 characters of hidden reasoning on a task with no reasoning in it before reaching the answer lines. Under note (b) a length body is a seat failure, never a partial answer, so all four were rejected by the runner.

Note (bhf) — measure a seat's appetite before depending on it — fires for the fourth recorded time, on the same slug S090 measured at 4–6× P1's cost for the same passage. It was read as a cost note and it is also a capacity note; that is the part this run adds.

What A5 changes: max_tokens for P2 alone, 4,000 → 16,000. The prompt is re-composed by the same functions from the same frozen materials with the same hash-derived rotations and is byte-identical; the question, the conditions and the scoring are untouched. Nothing about what is measured changes.

What is recorded rather than absorbed. The four failed bodies billed $0.033369 + $0.0333675 + $0.033369 + $0.0333675 = $0.1334730 and are kept on disk with their metadata. They are ledgered in config/budget.md as spend against this session, because they were. A run that quietly re-dispatches until it gets a body is not reporting its own cost.