Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260728f-nonlead-items/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260728f-nonlead-items
statusfrozen
created2026-07-28
updated2026-07-28
sensesstyle-correspondence, cultural-mediation
linkswiki/arms/ARM-sense-boundary.md, wiki/goodness-senses.md, workshop/experiments/E-20260727c-sense-boundary/design.md, wiki/findings/results/RS-20260727e-sense-boundary.md, wiki/decisions/resolved/D-20260727-09-sense-boundary-amendment.md, workshop/translations/wang-liulang/R04-v1/translation.md

E-20260728f — does R1's agreement gain survive items the lead did not write, in a language where clause (a) has nothing to point at?

FROZEN before paragraphs 2–4 of the held-out translation existed. Paragraph 1 existed only as the contamination selection gate's probe (workshop/translations/wang-liulang/R04-v1/contamination.md §4). The rule under test, R1, was frozen at S043 (c1cd2d3) and is not revisable by this experiment; the sham R0 is likewise carried over verbatim. What is new here is the materials and two of the predictions.


1. Question

RS-20260727e measured that three model outputs sort contested style-correspondence / cultural-mediation decisions far more consistently when rule R1 is stated (0.654 → 0.951 on socially indexical items), and that a length- and authority-matched sham moves agreement by −0.025. The ratifying vote of D-20260727-09 refused to adopt R1, and its first stated reason was that the experiment measured consistency with the lead's scheme rather than independent derivation. RS-20260727e §6.3 concedes the point in the same words: "what was measured is whether raters map lead-written descriptions onto labels more consistently with R1 than without it", and says it is not fixable inside that design, because a decision cannot be described without saying what was at stake in it.

This experiment removes the description. It asks:

(i) Does R1's agreement gain survive when no item contains any lead prose about the decision — when the item is the source span and the translator's rendering, and nothing else?

(ii) Does it survive in a language where R1's clause (a) has no referent? RS-20260727e §6.6 carried the pre-run critic's strongest unanswered objection to this arm's step 2: clause (a) presupposes a morphological contrast set that lexical, diachronic and discourse-level cases may not have. R1(a) sorts a marker to style-correspondence when its meaning comes from "the selection it makes among alternatives the source's own grammar or morphology offers at that same site." Classical Chinese has no politeness morphology at all. Its social marking is entirely lexical — 僕 for I, 君 and 兄 for you, a personal name for a title. So at every socially indexical site in this material, clause (a)'s test as written has nothing to point at, while clause (b)'s does.

That is not a defect of the material. It is the case S043's Japanese stratum was chosen to be the easiest possible instance of, and it is the case the objection names.

(iii) Is clause (c) — composite sites are decomposed, never double-scored — used by anyone? RS-20260727e §5 records that under R1 nobody decomposed anything, and that clause (c) is "the part of the rule this run says nothing about." The answer format is why: both was offered and composite was not. It is offered here, in every condition.

2. What is carried over verbatim and is not revisable here

3. Materials

Held-out text. 〈王六郎〉 from 《聊齋志異》 (Pu Songling, c. 1680), 1,533 characters, established from two witnesses with two emendations (workshop/translations/wang-liulang/provenance.md), translated by the lead in session under R04 as T-wang-liulang-R04-v1, with its single-pass draft frozen first as an R06 artifact and its translator's log frozen before this design's run. Contamination measured, not asserted (§4 of the contamination note).

Site extraction is not the lead's. One call to a model that is not a rater and not the critic — deepseek/deepseek-v4-pro (P5) — receives the frozen Chinese and the frozen English, paragraph-aligned, and returns 36–48 sites. Its brief, frozen here verbatim:

You are reading a classical Chinese tale and an English translation of it. Identify places where the translator had a real choice about how to render a word or expression — because its rendering is not settled by dictionary meaning alone. Return the source span verbatim (a contiguous substring of the Chinese, as short as it can be while still being identifiable), the enclosing source sentence verbatim, and the English words the translator actually used for it, verbatim.

Do not explain any choice. Do not evaluate anything. Do not classify anything. Return only the spans.

The extractor is told nothing about the two senses, about R1, or about what the sites will be used for. It supplies no label, no category and no description. It locates; it does not sort.

Item format, fixed, and it contains no lead prose about any decision:

<n>. 源: <enclosing source sentence, verbatim Chinese, with the site marked ⟦…⟧>
     訳: <the translator's English for that sentence, verbatim, with the corresponding words marked ⟦…⟧>

That is the whole item. There is no sentence saying what the translator had to decide, because that sentence is what RS-20260727e §6.3 says was doing the work.

Mechanical checks before dispatch, all of which abort the run if they fail:

  1. Every returned source span occurs verbatim in the frozen source. Spans that do not are dropped and the drop count is reported.
  2. Every returned English span occurs verbatim in the frozen translation. Same treatment.
  3. Duplicate spans are collapsed to one item.
  4. Leak check. No item text may contain any of: formal, form, style, stylistic, marked, culture, cultural, realia, honorific, politeness, deixis, register, correspondence, mediation, contrast, contrastive, denote, denotational, refer, referent — the S043 word ban, applied to the English half. Here it should be trivially satisfied, since the English half is running prose from a translation rather than a description of a decision; the check runs anyway and its result is reported, because "trivially satisfied" is a prediction.
  5. Items are presented in one fixed shuffled order (seed frozen below), the same order to every rater in every condition.

Stratification is mechanical and is frozen here, before the sites exist. Each item falls in stratum P if its source span contains at least one of the following characters or strings, and in stratum N otherwise:

君 僕 仆 卿 妾 臣 吾 余 予 爾 汝 某 兄 弟 郎 姥 媼 娘 翁 哥 我
足下 閣下 老身 小生 賤妾 鄙人 愚

These are classical Chinese first- and second-person reference and terms of address and self-designation. The list is applied by exact substring match by script, with no judgment at run time. It is not a partition into the two senses and must not be read as one: it separates sites where the source names a person from everything else, which is where clause (a)'s presupposition is at issue.

Raters: three panel models, three labs, all non-Anthropic — P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5. Blind to each other, to the lead's own reading, to the stratum of any item, and to which condition the others saw.

The lead's own assignments are not collected at all. At S043 they were frozen in advance and reported as a secondary non-evidential figure; here they are omitted, because the whole point of this run is that the lead's scheme is what is under suspicion, and a lead column invites the same reading a second time.

Shuffle seed: 20260728, random.Random(20260728).shuffle(rows) after sorting rows by source-span character offset.

4. Conditions

Four, run as independent stateless calls, three raters each — 12 calls.

Task, identical in all four. For each item, exactly one of:

style-correspondence · cultural-mediation · both · neither · composite

with both and composite distinguished in the instructions, in all four conditions:

both — the marked element belongs under both senses as a whole and cannot be split between them. composite — the marked element separates into more than one part, and the parts belong under different senses.

Five options, not four, and the difference from S043 is deliberate: composite is the only way clause (c) can be observed being used or not used, and offering it in every condition means the rule is not the only thing that makes decomposition available.

5. Measures

6. Predictions, pre-registered

7. Failure criteria, pre-committed

8. What a pass would and would not license

Would: "The agreement gain R1 produces at S043 is not an artefact of the lead having written the item descriptions, and it is not produced by a language whose politeness marking is morphological."

Would not: that R1 cuts in the right place. Charter §4 — panel agreement is not validation, and this instrument is evidence in the failing direction only. Three models agreeing could still reflect one shared conventional account of how people are addressed in classical Chinese.

Would not: anything about scoring translations. This classifies decision sites, not prose quality. Tier D is unpassed and nothing here changes that.

9. Known limitations, recorded before the run

  1. The English half of every item is still the lead's. The rendering is the artifact under discussion, not a description of the decision, and that is the whole reduction this design achieves — from 100% lead prose about the decision to the source verbatim plus the lead's actual English. It is a reduction and not an elimination: a rendering can telegraph its own classification (rendering 君 as "sir" is itself a legible choice). This is the residual form of RS-20260727e §6.3 and it is not removable by any design in which the lead is the translator.
  2. The lead chose the text, and the extractor's brief was written by the lead. The brief names no sense and returns no label, but a brief that asks for "places where the rendering is not settled by dictionary meaning alone" shapes which sites exist. Reported verbatim in §3 so it can be attacked.
  3. One language, one translator, one text. S043 was one language (Japanese) on the held-out side; this is one language (classical Chinese). Two single-language runs pointing the same way is better than one and is not generality.
  4. Three raters is three raters, with wide κ intervals on strata of this size. Reported with intervals.
  5. The composite option is new, so the response space is five-way here and four-way at S043. Agreement numbers are therefore not comparable across the two experiments in absolute terms; what is comparable is the within-run A→B difference, which is what Q2 is stated on.
  6. The extraction call could fail or return junk. The mechanical checks in §3 are the guard, and the drop count is reported rather than absorbed. If fewer than 24 items survive the checks, the run does not dispatch and the session reports that instead.

10. Amendments

Dated amendments made before any rating call was dispatched.

A1 (2026-07-28, before dispatch) — stratum P came back empty, and two predictions are therefore unevaluable rather than failed

The extractor returned 40 sites, all of which pass every mechanical check (0 dropped for any reason: no non-verbatim span, no duplicate, no leak). Not one of the 40 contains a character from the frozen L_P list. The extractor identified no site of person reference at all — not 君, not 僕, not 兄/弟, not 六郎, not 客 — in a 1,533-character tale that uses those forms twenty-odd times.

Consequence, stated before the run rather than after it:

The empty stratum is not treated as an accident to be worked around; it is reported as a result. Two readings are live and this run cannot separate them:

  1. The brief selected against the class. §3's brief asks for places where the rendering "is not settled by dictionary meaning alone", and 君 → you is settled by dictionary meaning in the only sense a dictionary records. If so, the fault is the lead's wording and a differently-worded brief would find the sites.
  2. The class is not visible to a reader who has not been pointed at it. If an independent reader working through a source and its translation does not register lexical honorifics as decisions at all, that bears on how much the seam matters in practice.

No second extraction was run. Re-briefing the extractor after seeing that the first brief produced an inconvenient item set is fitting the materials to the prediction, and it is the precise move the D-20260727-09 ratifying vote refused this rule's evidence for.

A2 (2026-07-28, before dispatch) — critic dispositions

Applied from the independent pre-run critic pass; see critic/dispositions.md for what was accepted, what was declined, and why. Ten findings accepted, three declined in writing.

A3 (2026-07-28, before dispatch) — what this run may and may not say, and A1 corrected

A1 was wrong about Q3 and the critic caught it (finding 1.2). Q3 is stated on stratum P — "the A→C rise on stratum P is less than half the A→B rise" — so moving it to the whole set makes it a different comparison chosen after the materials were seen. Q3 joins Q2 and Q4 as UNEVALUABLE. The sham therefore cannot perform its pre-registered role of attributing any rise to R1, and no causal attribution is available from this run at all.

Prohibited vocabulary, binding on the result page and on every downstream citation of it. This run may not be described as showing that R1's gain "survives the leak fix", that it "replicates S043", or that it "supports R1". The whole-set A→B and A→C differences are secondary and non-pre-registered and may only be stated in the form the critic proposed:

Among 40 extractor-selected sites, none of them in stratum P, condition B differed from condition A in raw agreement by X. This was not a pre-registered test and cannot adjudicate the contested seam or the leak-fix question.

Q5 is relabelled (critic 1.4): it measures use of the composite response category, not application of R1's clause (c), because the response schema collects a label and not a decomposition. Clause (c) remains untested, for the second run running.

Q1's pass is narrowed (critic 1.3). §6 justified the 0.60 threshold on N being the least contested stratum. N is now the whole item set and is a grab bag — ritual realia, offices, idioms, images, and ordinary lexical choices together. A pass therefore licenses only "raters can use this item format and produce non-random labels", and not "raters agree where the page already separates."

Why the run dispatches anyway, stated before the data exists. Not for the headline, which is gone. For one question this item set is unusually well suited to and which RS-20260727e §3 left open: there, under R1, chance-corrected agreement fell while raw agreement rose, because 92.5% of votes landed on one label. A largely out-of-scope item set is a strong test of whether R1 buys agreement by funnelling answers into a default — clause (d)'s residue in particular. The analysis therefore adds, before dispatch: full A→B and A→C label-transition matrices; a collapse diagnostic counting items that go from split under A to unanimous under B, and separately to unanimous neither; per-condition neither and composite rates; a sentence-clustered bootstrap (three source sentences carry two items each, so the item-level resample overstates independence — critic 1.6); a first-half/second-half position check; and all three rater pairs separately, which is also the check on the critic/rater overlap.

If the collapse diagnostic shows the agreement gain is mostly A-split → B-unanimous-neither, that is the finding, and it is a finding against R1's usefulness rather than for it.