Repository path: workshop/experiments/E-20260727c-sense-boundary/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260727c-sense-boundary |
| status | frozen |
| created | 2026-07-27 |
| updated | 2026-07-27 |
| senses | style-correspondence, cultural-mediation |
| links | wiki/arms/ARM-sense-boundary.md, wiki/goodness-senses.md, wiki/findings/results/RS-20260727-log-typology.md, workshop/experiments/E-20260727-log-decision-coding/codes.md |
E-20260727c — does rule R1 bound the style-correspondence / cultural-mediation seam?
FROZEN before the held-out material exists. This file was committed before a word of T-saigo-no-ikku-R04-v1 was translated. That ordering is the design: rule R1 below is drawn from the 26 C4a rows of the S037 corpus, which are therefore its training set, and the decisions generated by the new translation are its test set. A rule written after seeing the new decisions would be fitted to them.
1. Question
wiki/goodness-senses.md lets both style-correspondence and cultural-mediation claim grammatical politeness and honorific marking, and neither defers (RS-20260727-log-typology §2). Can a written rule assign such decisions to one sense or the other in a way that independent readers reproduce?
The question is deliberately not "is R1 the correct boundary?". Nothing available to this project can answer that. What is measurable is whether R1 defines a set — whether readers given only its words land in the same place. That is the operational content of the S042 finding that a criterion declared soft was still unbounded in fact (note (bcl)): boundedness is an inter-rater property, not a property of how carefully the author hedged.
2. Rule R1 — frozen verbatim, not revisable by this experiment
R1. Where does the marker's meaning come from?
(a) Contrastive →
style-correspondence. The marker's meaning comes from the selection it makes among alternatives the source's own grammar or morphology offers at that same site. Swapping in a different member of the set would change what is said about the speaker, the addressee, or the relation between them, while leaving unchanged what is referred to. Test: name the alternative the source did not use. If you can name it, and swapping it in changes the social reading without changing the referent, the decision isstyle-correspondence's.(b) Denotational →
cultural-mediation. The marker's meaning comes from what it picks out in the source's world — an office, institution, rank, title, place, procedure, or role the target reader may not hold. Test: ask what the target reader would have to be told. If what is lost by flattening is knowledge of a thing, the decision iscultural-mediation's.(c) Composite sites are decomposed, never double-scored. Many address forms carry both — an office name (denotational) plus an honorific suffix or predicate form (contrastive). Split the site into two decisions and score each once, under its own sense. The same component is never scored under both senses.
(d) Residue. A marker that fails both tests — no nameable alternative at the site, and nothing the reader would need to be told — belongs to neither sense and is recorded as residue rather than forced into one.
What R1 would change on the page if it survives: cultural-mediation's item list currently reads "realia, honorifics, allusion, politeness deixis, measure words, names" and would narrow "honorifics, politeness deixis" to their denotational aspect (titles and terms of address as social-world referents); style-correspondence's definition, which names sentence shape, repetition, sound play and typographic play but not socially-indexical morphology, would gain it explicitly — which its own S013 grounding note already assumes.
R1's derivation, stated so it can be attacked: it is the two senses' own established mechanisms turned into a partition. style-correspondence's cross-linguistic result (S013) is that grammar-borne meaning is transcoded into lexis at a cost in systematicity — a property only a contrast set has. cultural-mediation's (S017/S018) is that its load is cultural distance × inheritability, and the S018 correction located the variable in reader distance: whether the target's readers still hold the signifier's sense. Those are the two clauses of R1.
3. Materials
Item set, ~40–50 rows, three strata:
- SEAM-C — the 26 rows coded
C4ainworkshop/experiments/E-20260727-log-decision-coding/decisions.tsv. R1's training set. Included so the corpus the seam was found in is measured, not only the new material. - SEAM-H — held-out seam decisions from
T-saigo-no-ikku-R04-v1, a lead translation of Mori Ōgai 「最後の一句」 §4, made after this file was frozen. - CLEAN — control items from the same translation that are unambiguous under the current page wording: named realia and office names with no politeness component (
cultural-mediation), and marked formal features with no social component — sentence shape, repetition (style-correspondence). These are the instrument check. If raters cannot separate items the page already separates, nothing else in the run is interpretable.
Item format, fixed. Each item gives: source language; the source form, quoted; and one neutral sentence stating what the translator had to decide. Items are numbered and presented in one fixed shuffled order, the same order to every rater in every condition.
Leak control, checked mechanically before the run. No item text may contain any of: formal, form, style, stylistic, marked, culture, cultural, culture-bound, realia, honorific, politeness, deixis, register, correspondence, mediation, contrast, contrastive, denote, denotational, refer, referent. The first two strata are about honorifics and politeness, so naming the phenomenon in the item would hand the rater the answer under either condition. A script verifies this and the run does not dispatch if it fails.
4. Conditions
Two, run as independent stateless calls — there is no carryover, because no rater sees both prompts in one context.
- A — control. The rater is given the current
wiki/goodness-senses.mddefinitions of the two senses, verbatim, including their grounding notes, and the item list. - B — treatment. Identical, plus rule R1 verbatim.
Task, identical in both: for each item, answer exactly one of style-correspondence / cultural-mediation / both / neither. both and neither are offered in both conditions, so the rule is not the only thing that makes a partition available.
Raters: three panel models, three labs — P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5. Six calls total. Raters are blind to the lead's own assignments, to each other, and to which stratum any item belongs to.
The lead's own assignments under R1 are frozen in this directory before the run (lead_assignments.tsv) and are reported as a secondary, non-evidential figure. The evidence is between-rater agreement among the three panel models. The lead may not adjudicate its own criterion.
5. Measures
- Raw pairwise agreement — mean over the three rater pairs, per stratum, per condition.
- Chance-corrected agreement — Fleiss' κ over three raters, per stratum, per condition, with expected agreement from the observed marginals of that stratum × condition cell.
bothrate — proportion ofbothresponses, per stratum, per condition.
6. Predictions, pre-registered
- P4 (instrument check). CLEAN items reach ≥ 0.75 mean raw pairwise agreement in condition A.
- P1 (a seam exists). In condition A, mean raw pairwise agreement on SEAM (C and H pooled) is at least 0.15 below that on CLEAN.
- P2 (the rule bounds). Mean raw pairwise agreement on SEAM rises by at least 0.15 from condition A to condition B.
- P3 (secondary). The
bothrate on SEAM falls from A to B. - P5 (secondary). P2's effect holds on SEAM-H alone — the held-out stratum — not only on the training stratum SEAM-C.
7. Failure criteria, pre-committed
- P4 fails → the run is uninterpretable. Report that and stop. No claim about the seam, in either direction.
- P1 fails → there was no seam to draw. The overlap is real on the page and not observable in rater behaviour. The backlog item is discharged as a non-problem; R1 is not adopted.
- P2 fails → the boundary is not drawn. Record the seam as undrawn on
wiki/goodness-senses.md, keep R1 and this measurement on the record as an attempt that failed, and do not amend the sense definitions. - P1, P2 and P4 all hold → R1 is adopted, and the page is amended as §2 describes.
8. What a pass would and would not license
Would: "Independent raters given R1's words apply it to seam decisions more consistently than they apply the current page wording." That is what boundedness means here and it is the whole claim.
Would not: that R1 is the right place to cut. Charter §4 — panel agreement is not validation, and this instrument is evidence in the failing direction. Three models agreeing could reflect a shared prior about honorifics rather than a well-drawn rule; the design cannot separate those. Any result page must say so in the same breath as the number.
Also would not: anything about how a jury would score a translation. This measures classification of decision descriptions, not evaluation of prose. Tier D is unpassed and nothing here changes that.
9. Known limitations, recorded before the run
- R1 was drawn from SEAM-C by the lead, so SEAM-C is a training stratum. P5 exists because of this.
- Three raters is three raters. κ on ~26 and ~15 item strata has wide intervals; the design reports the intervals rather than pretending otherwise.
- SEAM-H is one language. Japanese has the densest and most systematic politeness morphology of any language in the corpus, which likely makes R1's clause (a) easier to apply there than on, say, Spanish
vuestro. SEAM-C's nine languages are the check on that, and they are the training stratum. This confound is not removable within one session and is not removed. - The item descriptions are the lead's, in both strata. A rater classifies the lead's one-sentence account of a decision, not the decision. Leak control (§3) bounds the crudest version of this and not the subtle version.