Repository path: workshop/experiments/E-20260731d-sense-axes/design/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260731d-sense-axes |
| status | frozen |
| created | 2026-07-31 |
| updated | 2026-07-31 |
| track | T2 |
| senses | voice, style-correspondence |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-sense-axes.md, wiki/goodness-senses.md, workshop/translations/levsha/R06-v1/translation.md, wiki/findings/results/RS-20260729b-graded-drift.md, wiki/findings/results/RS-20260729g-graded-senses.md, wiki/method-notes.md |
E-20260731d-sense-axes — are voice and style-correspondence one graded axis?
Frozen before dispatch. Nothing below was written or altered after any rater output existed.
ARM-sense-axes step 1. Track T2.
1. Question
wiki/goodness-senses.md holds nine categorical senses. RS-20260729b showed on the project's other
typology that splitting one unreliable categorical label into two graded axes moved three-rater agreement
from α 0.51 to 0.78/0.89. ARM-sense-axes exists to ask whether the nine senses conflate axes the same
way. This experiment asks it of one pair:
Are
voiceandstyle-correspondencetwo categories, or one graded axis that a categorical label cuts?
Why this pair. It is one of the four watch-pairs recorded at the typology's ratification (S002:
voice/style-correspondence), it has never been tested, and — the reason it is the strongest candidate —
the two definitions differ on a stated quantity: style-correspondence is defined as local and
formal, voice as global and cumulative. If any pair on that page is one axis cut at a threshold,
this is it. It is deliberately not the style-correspondence/cultural-mediation seam, which
ARM-sense-boundary and ARM-graded-typology have already worked and where the graded instrument was
retired on its own numbers (RS-20260729g).
This is not a jury design and must not become one (arm constraint, charter §5, the S015 line). No rater is asked whether any rendering is good, better, or worse than any other. Every question is which label fits this site or how far does this site reach. Tier D has not passed and nothing here depends on its passing.
2. Materials
40 decision sites from T-levsha-R06-v1 — Leskov, «Левша», chapters 6–8, 974 Russian words into
English, R06 lead single pass, translated and its log frozen at commit e2f8d96 before this design
existed, and its contamination gate run after that freeze and before this item set was built (note
(bcd)'s prescribed order): longest common run 11 tokens, 0 shared 12-grams, against the PD comparator
(Gutenberg #61172).
Why this material. «Левша» is skaz. There is no author's language in it, only a Tula townsman's, and every formal property of the text is simultaneously a property of the person the text sounds like. That is the seam under test, at maximum density, in a text nobody chose for that reason. It is also the material's known limitation and it is stated in §7.
Item strata (design/items_source.json, frozen; built by build_items.py):
| stratum | n | what it is |
|---|---|---|
SEAM |
26 | sites where both senses plausibly apply |
CF |
4 | positive control, form pole — a source feature whose handling is a matter of whether English has the construction |
CV |
4 | positive control, passage pole — a site where no single formal marker is at stake |
NF |
6 | foils — sites belonging to cultural-mediation or to accuracy's compelled-specification clause, so that neither is reachable (notes (bdq), (bfq)) |
Every item is three lines: the Russian, the English, and one sentence stating what was chosen. No evaluation, no reason, no quality word.
The wording constraint, checked mechanically and enforced by build failure. No item text may contain
either sense's name or any of narrat*, persona, perspective, stance, speaker, teller,
style, stylistic, or the other seven sense ids, matched as whole words. build_items.py exits
non-zero and writes nothing if any item violates it. (The first build flagged impersonally and
distance under substring matching; the matcher was changed to word boundaries, which is the constraint
as stated and not a relaxation of it. Recorded because it happened before the items were frozen.)
3. Conditions — the same raters, the same items, two question shapes
Within-item, per the arm's constraint that a design changing items and question shape together measures neither.
CAT. For each item, two independent fields:
- label ∈ {A, B, neither}, where A and B are the two senses' verbatim definitional sentences
from wiki/goodness-senses.md, presented unnamed, as A and B, so no rater can answer from
a sense's title. The A/B assignment is counterbalanced across raters — P1 and P3 see A = the
voice sentence, P2 sees A = the style-correspondence sentence — and every answer is mapped back to
a sense id before any analysis. (Amendment made after the pre-run critic pass, on the lead's
initiative and not at the critic's request: a single fixed order cannot detect an A-side bias, and
S065 measured a 33% order-flip rate on a forced choice in this project. It adds no call and does not
move the cost estimate.)
- needs ∈ {ONE, BOTH} — is one of the two descriptions enough for this site, or does it need
both? — asked as its own field, not as a third option in the label slot. This is the note (ber)
test; see prediction P5.
GRAD. For each item, two integers 0–4: - REACH — how far beyond this one site does what is at stake here extend? 0 = this site only; 4 = the whole of the three chapters. - SURFACE — how much of what is at stake is a property of the source's text-surface (its shapes, sounds, marks, forms), as against a property of the sort of person the passage sounds like? 0 = entirely the latter; 4 = entirely the former.
Neither axis word (reach, surface) occurs in either definition.
Order and ids. CAT runs first for all three raters, then GRAD. The two blocks present the same 40
items in different frozen shuffles under different ids (C01…C40, G01…G40; 4 of 40 items land at
the same position in both), so no answer transfers by position. The order is not counterbalanced —
three raters cannot support two orders — and the CAT→GRAD anchoring confound is declared in §7.
Raters. P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5, temperature 0,
max_tokens 6000. Reserve declared before dispatch (note (bfc)): RES = deepseek/deepseek-v4-pro. (The reserve was called P5 at freeze; the pre-run critic's one finding was that this collides with prediction P5, and it is renamed RES everywhere.)
Acceptance requires finish_reason == "stop" and a non-empty body; a non-accepted attempt falls through
to the reserve and is never retried on the same slug (notes (b), (bdl), (bdb)).
4. Registered nulls — computed and committed BEFORE dispatch
The arm requires a no-effort null, on the S067 lesson that a threshold can turn out to be exactly the trivial-satisfier baseline and the only way to know is to compute the baseline first.
N1a — nearest-definition null. Bag-of-words cosine of each item's text against each sense's
definitional sentence, crude-stemmed, stopword-filtered; the nearer wins, zero-on-both goes to neither.
Uses the project's own words, not a lexicon invented for this run. Realised distribution: 26 neither,
8 style-correspondence, 6 voice.
N1b — hand-frozen lexicon null. A frozen list of surface-property words against a frozen list of
text-attitude words; more of the former → style-correspondence, more of the latter → voice, tie →
neither. Realised distribution: 33 style-correspondence, 4 voice, 3 neither.
N1a was degenerate on the first build (30 of 40 neither) and N1b was added for that reason, before
dispatch. Both are frozen in build_items.py and both are reported. That N1b lands 33 of 40 on one side
is itself a fact about the lead's item prose and is reported as one.
Registered interpretation, and it binds. If the mean rater↔null agreement is at or above the mean rater↔rater agreement on CAT, this run cannot distinguish "the raters read the site" from "the raters read the lead's wording", and no conclusion about the two senses may be drawn from it. A low rater↔null agreement does not certify the instrument; it only fails to condemn it.
N2 — length null. Spearman ρ between each item's character length and each graded axis. Registered: |ρ| > 0.5 on an axis means that axis is confounded with the length of the lead's prose and its α is not reportable as a fact about the sense.
N3 — permutation floor for the derived label. The derived four-way label's α is compared against 1,000 within-rater shuffles of the axis scores. The derived α must clear the 95th percentile of that floor to be read as anything.
5. Predictions — frozen, each able to fail
- P1. α(CAT, nominal, three-way) < 0.70. The other seam ran 0.553 categorically and this pair has never been measured.
- P2. The derived four-way label — each axis cut at that rater's own median, per
RS-20260729b— will land within ±0.10 of α(CAT). Rise > 0.10 supports the one-graded-axis reading. Fall > 0.10 reproduces the S054 mechanism here and not at the other seam. This is the arm's central quantity. - P3 — the note (bfl) test. Pooled Spearman(REACH, SURFACE) will have |ρ| < 0.25, against the
−0.358 (bfl) measured on the
form/referentpair. Rationale: those two axes were complements by construction — a fixed quantity divided — and these two are not, because a source feature can be both text-surface and passage-wide, or neither. If |ρ| ≥ 0.358 again, (bfl) is a property of raters splitting any judgement in two, not of the S059 wording, and that is the stronger finding. - P4 — the corner test. Both corners a single dimension forbids will be occupied: ≥ 3 of 120 cells (3 raters × 40 items) at REACH ≥ 3 and SURFACE ≥ 3, and ≥ 3 of 120 at REACH ≤ 1 and SURFACE ≤ 1. Zero occupancy of either corner is the direct signature of one dimension wearing two names.
- P5 — the note (ber) test.
BOTHtake-up ≥ 5 of 120 cells. Two prior designs offered a third option inside the label slot and got 0 of 240 and 0 of 266. This design moves the question to its own field. Zero again generalises (ber) past the response format; substantial take-up narrows (ber) to single-slot option lists. - P6 — the positive control, and it can void the run. On the control items endorsed by the pre-run
critic (see §6), ≥ 75% must receive the intended CAT label from ≥ 2 of 3 raters, and ≥ 75% must
have the intended axis on the intended side of the midpoint (
CF: SURFACE ≥ 3;CV: REACH ≥ 3) from ≥ 2 of 3 raters. Control failure voids every seam conclusion in this run. - P7 — reported, not predicted (note (bfq)). The realised label vocabulary per rater is printed
before any agreement figure is quoted, for CAT labels, for
needs, and for each graded axis. A rater using fewer than 2 of the 3 CAT labels is reported as such and its agreement contribution is flagged.
6. Procedure
build_items.py— constraint check, nulls, two orderings. Done and committed before dispatch.- Pre-run critic, one non-rater, non-reserve seat (
qwen/qwen3.7-max, probed-but-not-selected, so the S053 role-collision fix holds). It is given this design and the item set and asked, among the usual, to endorse or contest each of the 8 control items sight-unseen — the S070 precedent, where the critic's split of the lead'scompositedeclarations was what made the control interpretable. The critic's endorsement, not the lead's declaration, defines the control set P6 is evaluated on. All findings are answered in writing; BLOCKING findings are applied or the run does not go. - CAT block, 3 seats, one call each.
- GRAD block, 3 seats, one call each.
analyse.py— every registered quantity, in the order above, nulls first.verify.py— an independent recomputation of every number that reaches the result page, from the stored.rawbytes, walking the reserve chain and assertingfinish_reason == "stop"on the body it reads (note (bdt)); plus mutation tests that must fail.
7. Declared limitations
- The material is skaz, chosen because the seam is dense in it. If the two senses are separable anywhere, this is the hardest place to show it; if they separate here, that is strong. A null here is weaker evidence than a positive.
- The order is not counterbalanced. CAT anchors GRAD for every rater. The shuffle and the re-id prevent positional transfer, not memory.
- The lead wrote the items, the axes and the control declarations. N1a/N1b bound the first, the critic's endorsement bounds the third, and nothing bounds the second.
- Three raters, all non-Anthropic panel models, none human. Nothing here is evidence about human readers, and panel agreement is not validation (charter §4).
- Tier D has not passed; every self-assessment arising is
provisional.
8. Cost
Worst case built from max_tokens, note (abc): critic 8,000 cap ≈ $0.13; six rater calls at 6,000 cap
≈ $0.30; one reserve firing ≈ $0.055. Declared worst case $0.50. Day headroom at design time
$4.448946523. Key-usage snapshots are taken around each call, not around the session, and the
cross-check is declared void for any inter-call interval that moves (note (bfv)).