Repository path: workshop/experiments/E-20260728f-nonlead-items/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260728f-nonlead-items |
| status | frozen |
| created | 2026-07-28 |
| updated | 2026-07-28 |
| senses | style-correspondence, cultural-mediation |
| links | wiki/arms/ARM-sense-boundary.md, wiki/goodness-senses.md, workshop/experiments/E-20260727c-sense-boundary/design.md, wiki/findings/results/RS-20260727e-sense-boundary.md, wiki/decisions/resolved/D-20260727-09-sense-boundary-amendment.md, workshop/translations/wang-liulang/R04-v1/translation.md |
E-20260728f — does R1's agreement gain survive items the lead did not write, in a language where clause (a) has nothing to point at?
FROZEN before paragraphs 2–4 of the held-out translation existed. Paragraph 1 existed only as the contamination selection gate's probe (workshop/translations/wang-liulang/R04-v1/contamination.md §4). The rule under test, R1, was frozen at S043 (c1cd2d3) and is not revisable by this experiment; the sham R0 is likewise carried over verbatim. What is new here is the materials and two of the predictions.
1. Question
RS-20260727e measured that three model outputs sort contested style-correspondence / cultural-mediation decisions far more consistently when rule R1 is stated (0.654 → 0.951 on socially indexical items), and that a length- and authority-matched sham moves agreement by −0.025. The ratifying vote of D-20260727-09 refused to adopt R1, and its first stated reason was that the experiment measured consistency with the lead's scheme rather than independent derivation. RS-20260727e §6.3 concedes the point in the same words: "what was measured is whether raters map lead-written descriptions onto labels more consistently with R1 than without it", and says it is not fixable inside that design, because a decision cannot be described without saying what was at stake in it.
This experiment removes the description. It asks:
(i) Does R1's agreement gain survive when no item contains any lead prose about the decision — when the item is the source span and the translator's rendering, and nothing else?
(ii) Does it survive in a language where R1's clause (a) has no referent? RS-20260727e §6.6 carried the pre-run critic's strongest unanswered objection to this arm's step 2: clause (a) presupposes a morphological contrast set that lexical, diachronic and discourse-level cases may not have. R1(a) sorts a marker to style-correspondence when its meaning comes from "the selection it makes among alternatives the source's own grammar or morphology offers at that same site." Classical Chinese has no politeness morphology at all. Its social marking is entirely lexical — 僕 for I, 君 and 兄 for you, a personal name for a title. So at every socially indexical site in this material, clause (a)'s test as written has nothing to point at, while clause (b)'s does.
That is not a defect of the material. It is the case S043's Japanese stratum was chosen to be the easiest possible instance of, and it is the case the objection names.
(iii) Is clause (c) — composite sites are decomposed, never double-scored — used by anyone? RS-20260727e §5 records that under R1 nobody decomposed anything, and that clause (c) is "the part of the rule this run says nothing about." The answer format is why: both was offered and composite was not. It is offered here, in every condition.
2. What is carried over verbatim and is not revisable here
- Rule R1, all four clauses, exactly as frozen at
workshop/experiments/E-20260727c-sense-boundary/design.md§2. Read programmatically from that file; not retyped. - Sham rule R0, exactly as dispatched at S043 (
E-20260727c-sense-boundary/run.py,SHAM_R0). Read programmatically from that file; not retyped. - The two sense definitions, read programmatically from
wiki/goodness-senses.md, i.e. as amended by the ratification ofD-20260727-09at S044. This differs from S043, where the definitions still carried the overlap the ratification removed, and it is a change in the control condition that must be stated: condition A here is a stronger control than condition A there. If R1 adds nothing over the narrowed page, that is a different and more interesting result than adding nothing over the old one. - The banner under which any rule is delivered, verbatim from S043: "APPLY THIS RULE. It takes precedence over your own reading of the two definitions above."
3. Materials
Held-out text. 〈王六郎〉 from 《聊齋志異》 (Pu Songling, c. 1680), 1,533 characters, established from two witnesses with two emendations (workshop/translations/wang-liulang/provenance.md), translated by the lead in session under R04 as T-wang-liulang-R04-v1, with its single-pass draft frozen first as an R06 artifact and its translator's log frozen before this design's run. Contamination measured, not asserted (§4 of the contamination note).
Site extraction is not the lead's. One call to a model that is not a rater and not the critic — deepseek/deepseek-v4-pro (P5) — receives the frozen Chinese and the frozen English, paragraph-aligned, and returns 36–48 sites. Its brief, frozen here verbatim:
You are reading a classical Chinese tale and an English translation of it. Identify places where the translator had a real choice about how to render a word or expression — because its rendering is not settled by dictionary meaning alone. Return the source span verbatim (a contiguous substring of the Chinese, as short as it can be while still being identifiable), the enclosing source sentence verbatim, and the English words the translator actually used for it, verbatim.
Do not explain any choice. Do not evaluate anything. Do not classify anything. Return only the spans.
The extractor is told nothing about the two senses, about R1, or about what the sites will be used for. It supplies no label, no category and no description. It locates; it does not sort.
Item format, fixed, and it contains no lead prose about any decision:
<n>. 源: <enclosing source sentence, verbatim Chinese, with the site marked ⟦…⟧>
訳: <the translator's English for that sentence, verbatim, with the corresponding words marked ⟦…⟧>
That is the whole item. There is no sentence saying what the translator had to decide, because that sentence is what RS-20260727e §6.3 says was doing the work.
Mechanical checks before dispatch, all of which abort the run if they fail:
- Every returned source span occurs verbatim in the frozen source. Spans that do not are dropped and the drop count is reported.
- Every returned English span occurs verbatim in the frozen translation. Same treatment.
- Duplicate spans are collapsed to one item.
- Leak check. No item text may contain any of: formal, form, style, stylistic, marked, culture, cultural, realia, honorific, politeness, deixis, register, correspondence, mediation, contrast, contrastive, denote, denotational, refer, referent — the S043 word ban, applied to the English half. Here it should be trivially satisfied, since the English half is running prose from a translation rather than a description of a decision; the check runs anyway and its result is reported, because "trivially satisfied" is a prediction.
- Items are presented in one fixed shuffled order (seed frozen below), the same order to every rater in every condition.
Stratification is mechanical and is frozen here, before the sites exist. Each item falls in stratum P if its source span contains at least one of the following characters or strings, and in stratum N otherwise:
君 僕 仆 卿 妾 臣 吾 余 予 爾 汝 某 兄 弟 郎 姥 媼 娘 翁 哥 我
足下 閣下 老身 小生 賤妾 鄙人 愚
These are classical Chinese first- and second-person reference and terms of address and self-designation. The list is applied by exact substring match by script, with no judgment at run time. It is not a partition into the two senses and must not be read as one: it separates sites where the source names a person from everything else, which is where clause (a)'s presupposition is at issue.
Raters: three panel models, three labs, all non-Anthropic — P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5. Blind to each other, to the lead's own reading, to the stratum of any item, and to which condition the others saw.
The lead's own assignments are not collected at all. At S043 they were frozen in advance and reported as a secondary non-evidential figure; here they are omitted, because the whole point of this run is that the lead's scheme is what is under suspicion, and a lead column invites the same reading a second time.
Shuffle seed: 20260728, random.Random(20260728).shuffle(rows) after sorting rows by source-span character offset.
4. Conditions
Four, run as independent stateless calls, three raters each — 12 calls.
- A — control. The two sense definitions, verbatim from the current (post-
D-20260727-09) page, plus the items. - B — treatment. Identical, plus rule R1 verbatim under the banner.
- C — sham. Identical, plus rule R0 verbatim under the banner.
- B2 — repeat of B, verbatim, to bound within-model instability on this item format. S043 measured ≈0.956 self-agreement on described items; a format with no description may be less stable, and the headline cannot be read without that ceiling.
Task, identical in all four. For each item, exactly one of:
style-correspondence · cultural-mediation · both · neither · composite
with both and composite distinguished in the instructions, in all four conditions:
both— the marked element belongs under both senses as a whole and cannot be split between them.composite— the marked element separates into more than one part, and the parts belong under different senses.
Five options, not four, and the difference from S043 is deliberate: composite is the only way clause (c) can be observed being used or not used, and offering it in every condition means the rule is not the only thing that makes decomposition available.
5. Measures
- Raw pairwise agreement — mean over the three rater pairs, per stratum, per condition.
- Chance-corrected agreement — Fleiss' κ over three raters, per stratum, per condition, expected agreement from the observed marginals of that stratum × condition cell. Reported next to raw agreement and not instead of it:
RS-20260727e§3 records that the two answer different questions and disagreed. - Concentration — the proportion of votes on the largest single label, per stratum, per condition. S043's B-condition figure was 0.925 and it is what made κ fall.
compositerate andbothrate, per stratum, per condition.- Self-agreement — B against B2, per rater, over all items.
- Bootstrap 95% CI on the A→B and A→C differences, 10,000 item resamples, seed frozen at
20260728.
6. Predictions, pre-registered
- Q1 (instrument check). In condition A, mean raw pairwise agreement on stratum N is ≥ 0.60. N is where the current page wording is least contested — named things, customs, offices, formal features — and if raters cannot reach 0.60 there, they cannot use this item format and nothing else in the run is interpretable.
- Q2 (the headline — does the gain survive the leak fix?). Mean raw pairwise agreement on stratum P rises by ≥ 0.15 from A to B. This is S043's P2, on items the lead did not write, against a control condition whose definitions have since been narrowed.
- Q3 (the sham must not do it). The A→C rise on stratum P is less than half the A→B rise. S043's pre-committed kill condition, reused verbatim in force.
- Q4 (clause (a) has nothing to point at — the new prediction, and it is the one the lead expects to fail). Under B, the proportion of stratum-P votes going to
style-correspondenceis below 0.80. S043's figure over all seam items was 0.925, produced by clause (a) firing on Japanese honorific morphology. Classical Chinese offers clause (a) no morphological contrast set at any site, so if R1's sorting power is clause (a)'s, the concentration must fall here. If it stays at or above 0.80, R1 is sorting on something other than what its own clause (a) says it sorts on, and that is a finding about the rule rather than about the material. - Q5 (clause (c) is finally testable). In condition B,
compositeis used on at least 3 distinct items by at least 2 of the 3 raters. This is a testability threshold, not a directional prediction: below it, clause (c) remains untested for a second run and this experiment says exactly that instead of reporting a null.
7. Failure criteria, pre-committed
- Q1 fails → the run is uninterpretable. Report that and stop. No claim about the leak, in either direction, and
ARM-sense-boundarycloses with the question open and the reason recorded. - Q2 fails → R1's measured gain does not survive removal of the lead's descriptions. The ratifying vote's objection is sustained on evidence rather than on argument. The seam is recorded as undrawn on
wiki/goodness-senses.md, R1 and both measurements are kept on the record as an attempt that failed, and the arm closes on its completion criterion 2. - Q2 holds and Q3 fails (the sham moves agreement at least half as much) → the rise is not attributable to R1. Same disposition as Q2 failing.
- Q2 and Q3 both hold → the gain survives the leak fix. The arm closes on completion criterion 1, and R1's adoption is opened as a decision for a later session to ratify. It is not applied to the page by this session, on the
D-20260727-09precedent andcontinue-prompt.md§3. - Q4 and Q5 are reported whatever they do and do not gate anything. They are what this run adds beyond replication.
8. What a pass would and would not license
Would: "The agreement gain R1 produces at S043 is not an artefact of the lead having written the item descriptions, and it is not produced by a language whose politeness marking is morphological."
Would not: that R1 cuts in the right place. Charter §4 — panel agreement is not validation, and this instrument is evidence in the failing direction only. Three models agreeing could still reflect one shared conventional account of how people are addressed in classical Chinese.
Would not: anything about scoring translations. This classifies decision sites, not prose quality. Tier D is unpassed and nothing here changes that.
9. Known limitations, recorded before the run
- The English half of every item is still the lead's. The rendering is the artifact under discussion, not a description of the decision, and that is the whole reduction this design achieves — from 100% lead prose about the decision to the source verbatim plus the lead's actual English. It is a reduction and not an elimination: a rendering can telegraph its own classification (rendering 君 as "sir" is itself a legible choice). This is the residual form of
RS-20260727e§6.3 and it is not removable by any design in which the lead is the translator. - The lead chose the text, and the extractor's brief was written by the lead. The brief names no sense and returns no label, but a brief that asks for "places where the rendering is not settled by dictionary meaning alone" shapes which sites exist. Reported verbatim in §3 so it can be attacked.
- One language, one translator, one text. S043 was one language (Japanese) on the held-out side; this is one language (classical Chinese). Two single-language runs pointing the same way is better than one and is not generality.
- Three raters is three raters, with wide κ intervals on strata of this size. Reported with intervals.
- The
compositeoption is new, so the response space is five-way here and four-way at S043. Agreement numbers are therefore not comparable across the two experiments in absolute terms; what is comparable is the within-run A→B difference, which is what Q2 is stated on. - The extraction call could fail or return junk. The mechanical checks in §3 are the guard, and the drop count is reported rather than absorbed. If fewer than 24 items survive the checks, the run does not dispatch and the session reports that instead.
10. Amendments
Dated amendments made before any rating call was dispatched.
A1 (2026-07-28, before dispatch) — stratum P came back empty, and two predictions are therefore unevaluable rather than failed
The extractor returned 40 sites, all of which pass every mechanical check (0 dropped for any reason: no non-verbatim span, no duplicate, no leak). Not one of the 40 contains a character from the frozen L_P list. The extractor identified no site of person reference at all — not 君, not 僕, not 兄/弟, not 六郎, not 客 — in a 1,533-character tale that uses those forms twenty-odd times.
Consequence, stated before the run rather than after it:
- Q2 and Q4 are UNEVALUABLE. Both are stated on stratum P and stratum P has n = 0. The pre-committed disposition for "Q2 fails" does NOT fire: a prediction that cannot be evaluated has not failed, and reading an empty stratum as a refutation of R1 would be exactly the kind of move this design exists to rule out.
ARM-sense-boundarytherefore cannot close on completion criterion 1 or 2 by this run. - Q1, Q3 and Q5 remain evaluable, Q1 and Q3 over the whole 40 (stratum N = the whole set), Q5 over the whole 40.
- The A→B and A→C differences over the whole item set are reported as SECONDARY, NON-PRE-REGISTERED figures, labelled as such wherever they appear. They are still the first measurement of R1's effect on items containing no lead description of any decision, which is the leak fix the ratifying vote asked for; they are simply not the stratified test that was pre-registered.
- The run dispatches. §9.6's abort threshold is fewer than 24 surviving items and 40 survive. Nothing about the item set is defective — the items are verbatim, non-lead-selected, and leak-free. What is defective is the fit between the brief and the prediction.
The empty stratum is not treated as an accident to be worked around; it is reported as a result. Two readings are live and this run cannot separate them:
- The brief selected against the class. §3's brief asks for places where the rendering "is not settled by dictionary meaning alone", and 君 → you is settled by dictionary meaning in the only sense a dictionary records. If so, the fault is the lead's wording and a differently-worded brief would find the sites.
- The class is not visible to a reader who has not been pointed at it. If an independent reader working through a source and its translation does not register lexical honorifics as decisions at all, that bears on how much the seam matters in practice.
No second extraction was run. Re-briefing the extractor after seeing that the first brief produced an inconvenient item set is fitting the materials to the prediction, and it is the precise move the D-20260727-09 ratifying vote refused this rule's evidence for.
A2 (2026-07-28, before dispatch) — critic dispositions
Applied from the independent pre-run critic pass; see critic/dispositions.md for what was accepted, what was declined, and why. Ten findings accepted, three declined in writing.
A3 (2026-07-28, before dispatch) — what this run may and may not say, and A1 corrected
A1 was wrong about Q3 and the critic caught it (finding 1.2). Q3 is stated on stratum P — "the A→C rise on stratum P is less than half the A→B rise" — so moving it to the whole set makes it a different comparison chosen after the materials were seen. Q3 joins Q2 and Q4 as UNEVALUABLE. The sham therefore cannot perform its pre-registered role of attributing any rise to R1, and no causal attribution is available from this run at all.
Prohibited vocabulary, binding on the result page and on every downstream citation of it. This run may not be described as showing that R1's gain "survives the leak fix", that it "replicates S043", or that it "supports R1". The whole-set A→B and A→C differences are secondary and non-pre-registered and may only be stated in the form the critic proposed:
Among 40 extractor-selected sites, none of them in stratum P, condition B differed from condition A in raw agreement by X. This was not a pre-registered test and cannot adjudicate the contested seam or the leak-fix question.
Q5 is relabelled (critic 1.4): it measures use of the composite response category, not application of R1's clause (c), because the response schema collects a label and not a decomposition. Clause (c) remains untested, for the second run running.
Q1's pass is narrowed (critic 1.3). §6 justified the 0.60 threshold on N being the least contested stratum. N is now the whole item set and is a grab bag — ritual realia, offices, idioms, images, and ordinary lexical choices together. A pass therefore licenses only "raters can use this item format and produce non-random labels", and not "raters agree where the page already separates."
Why the run dispatches anyway, stated before the data exists. Not for the headline, which is gone. For one question this item set is unusually well suited to and which RS-20260727e §3 left open: there, under R1, chance-corrected agreement fell while raw agreement rose, because 92.5% of votes landed on one label. A largely out-of-scope item set is a strong test of whether R1 buys agreement by funnelling answers into a default — clause (d)'s residue in particular. The analysis therefore adds, before dispatch: full A→B and A→C label-transition matrices; a collapse diagnostic counting items that go from split under A to unanimous under B, and separately to unanimous neither; per-condition neither and composite rates; a sentence-clustered bootstrap (three source sentences carry two items each, so the item-level resample overstates independence — critic 1.6); a first-half/second-half position check; and all three rater pairs separately, which is also the check on the critic/rater overlap.
If the collapse diagnostic shows the agreement gain is mostly A-split → B-unanimous-neither, that is the finding, and it is a finding against R1's usefulness rather than for it.