Repository path: workshop/experiments/E-20260729g-graded-senses/design/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260729g-graded-senses |
| status | frozen |
| created | 2026-07-29 |
| updated | 2026-07-29 |
| senses | style-correspondence, cultural-mediation |
| provisional | true |
| links | wiki/arms/ARM-graded-typology.md, wiki/findings/results/RS-20260729b-graded-drift.md, wiki/findings/results/RS-20260728f-nonlead-items.md, wiki/goodness-senses.md, wiki/arms/ARM-sense-boundary.md, config/models.md, config/budget.md |
E-20260729g-graded-senses — does the graded/categorical result transfer from the drift typology to the project's own?
Frozen before the held-out material exists and before any call is dispatched. ARM-graded-typology
step 2, the arm's last. The design is committed first, then the translation limb is made, then the
held-out sites are extracted, then the run is dispatched — each stage in a separate commit, in that
order.
1. The question, and why it is this one
RS-20260729b-graded-drift (S054) found, on the drift typology: three readers agree about how far a
word has drifted (α_ord 0.783) and not about which side of a line it falls on (α_nom 0.515); and
binarising the graded scores at each rater's own median drops α to 0.514, indistinguishable from the
categorical label's 0.516. The readers were not disagreeing about the words. They were disagreeing
about where to cut.
The project's central object, wiki/goodness-senses.md, is a categorical typology built the same way —
one reader, one set of considerations — and ARM-sense-boundary closed at S049 with one of its seams
recorded as undrawn, because with the lead's descriptions removed from the items the categorical
instrument stopped being reproducible.
So the question is whether S054's finding is a fact about the drift typology or a fact about
categorical typologies of this kind. ARM-graded-typology's completion criterion commits the arm to
one sentence in wiki/goodness-senses.md about what the S054 result does or does not license about the
project's other typology. This design exists so that sentence is a measurement rather than an
analogy.
2. Materials
Stratum ZH (40 items, pre-existing, unaltered). workshop/experiments/E-20260728f-nonlead-items/items.tsv
— 40 sites in 蒲松齡〈王六郎〉 chosen by a non-lead model from the source, presented as the classical
Chinese sentence with the site bracketed plus the lead's English with the site bracketed, and no lead
prose about any decision. This is the item set on which the categorical instrument's reproducibility is
already measured, by the same three raters.
Stratum KO (held-out, does not yet exist). A passage of Korean prose translated by the lead in this
session, from which a non-lead model will extract sites under a frozen brief, in exactly the ZH
format. Why Korean: S049's pre-registered person-reference stratum came back empty, because
classical Chinese marks social relation entirely lexically and therefore has nothing for R1's clause (a)
— a morphological contrast set — to point at anywhere in it. Korean marks it grammatically: speech
levels, the honorific infix -시-, humble benefactives, and a dense stock of role-and-rank address terms.
Stratum KO is therefore the material on which the two axes below can both fire on one site, which ZH
by construction cannot supply.
Neither text is named to any rater, and the two strata are shuffled into one list, as in
RS-20260729b §7.
3. The instruments
Three conditions, each run on the same single shuffled list.
CAT — the categorical instrument, as the project actually holds it. The prompt is
E-20260728f-nonlead-items/run.py's condition A: the current wiki/goodness-senses.md definitions
of style-correspondence and cultural-mediation quoted verbatim, and for each site exactly one of
style-correspondence / cultural-mediation / both / neither / composite.
GF and GR — the graded instrument, two independent axes, scored in separate calls. Each axis is 0–100 on every site, and each carries a counterfactual clause, which is what stops the two from being the two ends of one scale:
- F (form) — to what extent is what the translator had to handle at this site a marked formal feature of the source's expression (grammar, morphology, word order, sound, repetition, sentence shape, register-marking), such that the difficulty would remain even if the target culture held everything the source refers to?
- R (referent) — to what extent is it a thing in the source's social or material world (referent, custom, institution, title, role, artefact, allusion) that a target reader may not hold, such that the difficulty would remain even if the two languages had identical grammar?
A site may legitimately score high on both, low on both, or high on one. This is the property CAT cannot express and is the whole point of the comparison.
4. Raters, replication, and the defect this design exists to fix
Raters P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5 — the same three
that produced S049's condition A, so every comparison below is rater-matched and item-matched on
stratum ZH.
Every condition is replicated, byte-identically, at temperature 0. ARM-sense-boundary's closing
page returned this to wiki/backlog.md as a named defect of the S043 design that S049 inherited
wholesale: "a design that repeats every condition, not only the treatment … Cheap to fix and it doubles
the call count." It is fixed here. The reason it is not optional: S049's own byte-identical repeat
moved condition-level α_nom from 0.3594 to 0.6022, a swing of 0.243 on this item format. Against
noise of that size, an unreplicated 0.78-versus-0.65 comparison would mean nothing, which is why H1's
threshold below is stated relative to the measured noise rather than as a fixed margin.
3 conditions × 3 raters × 2 replicates = 18 rating calls, plus one extraction call and one independent pre-run critic call.
5. The registered baseline, computed from stored data before dispatch
From S049's stored bodies, three raters, 40 ZH items, Krippendorff α nominal:
| S049 condition | α_nom | unanimous / split |
|---|---|---|
| A — the page's own wording (the CAT instrument) | 0.6506 | 26 / 14 |
| B — rule R1 | 0.3594 | 15 / 25 |
| B2 — byte-identical repeat of B | 0.6022 | 25 / 15 |
| C — sham | 0.3818 | 16 / 24 |
The 14 ZH items on which S049's three raters split under condition A are 5, 12, 13, 15, 16, 18, 19, 21, 22, 24, 25, 30, 31, 38. This partition is frozen here, before any graded score exists, and is H3's secondary population.
6. Predictions, registered
| # | prediction | decides |
|---|---|---|
| H1 | α_ord(F) and α_ord(R) on stratum ZH each exceed α_nom(CAT) on the same items and raters, **by more than the largest replicate-to-replicate | Δα |
| H2 | The like-for-like control. Binarise each rater's (F − R) at that rater's own median → α_nom. Predicted: it falls to within 0.10 of α_nom(CAT), and at least 0.15 below α_ord(F). | whether the gain is bought by not drawing a line — the S054 shape |
| H3 | min(F̄, R̄) — "both-ness" — is higher on CAT-split items than on CAT-unanimous ones. One-sided Mann–Whitney, p < 0.05, on this run's replicate-1 CAT labels; reported again on §5's frozen S049 partition | whether the disagreements sit where both axes fire, the (bdm) shape |
| H4 | The noise floor. Per-rater test–retest Spearman ≥ 0.80 on F and on R | whether H1 and H2 are interpretable at all |
| H5 | α_ord(F) and α_ord(R) on stratum KO each ≥ 0.50, with α_nom(CAT) on KO reported beside them | whether the instrument transfers to material it has never seen, in a language whose politeness marking is grammatical |
Secondary, not registered as a prediction: CAT replicate 1 restricted to stratum ZH is a third measurement of S049's condition A, eleven sessions later. Its agreement with S049's stored labels is reported as a cross-session stability figure and nothing turns on it.
7. Failure criteria, written before the run
- F1 — a defective body. Any call returning fewer than all items, a non-integer score, a score
outside 0–100, or a label outside the five options is re-run once at a higher
max_tokens(note (bdl): on this project's evidence the remedy that works is raising the cap, not falling through). If it fails twice, the declared reserve P4moonshotai/kimi-k3substitutes, and the resulting change in panel composition is declared in the result rather than absorbed (the S057 consequence). - F2 — H4 fails. If test–retest falls below 0.80 on either axis, H1 and H2 are reported as uninterpretable, not as failed, and the arm closes on disposition D4.
- F3 — the extractor returns fewer than 12 KO sites. H5 is UNEVALUABLE, not failed — S049's disposition for an empty stratum, and the reason that disposition exists.
- F4 — the Korean competence screen. Four mechanically checkable Korean gloss questions are appended
to every rater's prompt and answered after the ratings, so they cannot anchor them. A rater scoring
below 3 of 4 has its stratum KO cells dropped, its ZH cells retained, and the drop declared; KO
figures then stand at n = 2. Korean has never been on this panel's competence screen
(
RS-20260725-anchor-verificationscreened Russian, French and Japanese), and this measures rather than assumes it. - F5 — no quality judgment. Nothing in this design asks any model whether any translation is good.
Tier D is not a gate on it, for the same reason it was not a gate on
ARM-sense-boundary: what licenses these raters is the factual-adjudication finding, not calibration.
8. Dispositions, registered — so the sentence is decided by the data
The arm's completion criterion commits to one sentence in wiki/goodness-senses.md. Which sentence is
fixed here, before the run:
- D1 — H1 passes and H2 passes. The senses typology's unreliability is measured to live in the cut. The graded instrument is adopted as a standing instrument in the narrow role of locating seam sites, explicitly not of adjudicating them.
- D2 — H1 fails. The graded instrument is retired with its measurement; the S054 result is recorded as local to the drift typology; the sentence says the transfer was tested and did not hold.
- D3 — H1 passes and H2 fails. The senses case is unlike the drift case: a cut through these axes
does reproduce. The sentence records that, and
ARM-sense-boundary's closure gains a named successor rather than being reopened by this session. - D4 — H4 fails. Double null. Per the arm's own words, "the claim's evidence class falls rather than being repaired."
CL-20260726-drift-window's restatement does not depend on this run — it is determined by evidence
already in hand at S054 — and is written whatever happens here.
9. Budget, worst case built from max_tokens (note (abc))
| stage | calls | max_tokens |
worst case |
|---|---|---|---|
| pre-run critic — P4 | 1 | 16,000 | $0.28 |
| KO site extraction — P5 | 1 | 16,000 | $0.09 |
| GF, GR | 12 | 10,000 | $1.20 |
| CAT | 6 | 6,000 | $0.28 |
| input, all 20 calls | — | ~7k each | $0.25 |
| total | 20 | ≈ $2.10 |
Against $3.369370 headroom on UTC day 2026-07-29. Priced at list out-rates per rater; P5's routing
caution in config/models.md is noted and P5 carries one cheap call only.
10. Verification
verify.py imports nothing from analyse.py, re-parses every body from the stored response JSON,
re-implements Krippendorff's α (nominal and ordinal), Spearman and the Mann–Whitney statistic
independently, recomputes §5's baseline from S049's stored bodies, and checks every number reported.
Mutation-tested per note (bdj): faults are injected and the verifier must catch and name them.
Amendments, from the independent pre-run critic — applied before any material existed and before dispatch
P4 moonshotai/kimi-k3, Fireworks, stop, in 4,276 / out 5,256, $0.137502. Verdict
NEEDS-AMENDMENT, nine findings — three BLOCKING, four MANDATORY, two ADVISORY. All nine accepted.
Note (rr), seventeenth consecutive session. Full body at runs/critic.json.
The critic also stated plainly which of the seven attacks it was asked to press fail: F4's screen is well-designed, F1/F3's dispositions are right, the replication fix is correct and its justification sound, and H2 is not arithmetically guaranteed — median-binarisation preserves agreement when raters share a latent cut, so the prediction is falsifiable in principle. Those are recorded because a critic that only ever finds faults is not measuring anything.
A1 (finding 1, BLOCKING) — H1 is demoted from a disposition gate to a screening statistic. α_ord on
0–100 and α_nom on a 5-class label are different metrics with different chance corrections; a 61-vs-63
disagreement is discounted where a style-vs-cultural disagreement is counted whole, so H1 is close
to guaranteed to pass for metric reasons alone — and D1/D2/D3 were being selected by it. The
dispositions are re-gated on H2, in its amended form below, which is metric-matched. H1 is still
computed and reported; nothing is decided by it.
A2 (finding 2, BLOCKING) — H2 is rebuilt as a genuine like-for-like control, and this is the
substantive change. Binarising (F − R) at a median and comparing a 2-class α_nom to a 5-class
α_nom mismatches the chance correction and, worse, discards exactly the both-ness that CAT encodes and
H3 analyses: a site at F = 80, R = 75 is a clear "both", and the difference-median assigns it to
whichever side noise puts it on. New H2: split each axis at that rater's own median and derive a
four-way label — high-F/low-R → style-correspondence, low-F/high-R → cultural-mediation,
high/high → both, low/low → neither — then compute α_nom of that against α_nom(CAT). Same category
set, same raters, same items; only the location of the cut differs. CAT's composite is reported
as-is in the primary and folded into both in a registered secondary. The old (F − R) binary is kept as
a secondary statistic.
A3 (finding 3, MANDATORY) — the axes are checked for independence rather than assumed to have it. Per-rater Spearman ρ(F, R) is registered as a descriptive, reported with the headline numbers. If pooled ρ < −0.3 the axes are anti-correlated by construction, (F − R) is an artificially stretched variable, min(F̄, R̄) is depressed everywhere, and H2 and H3 are flagged as computed on anti-correlated axes — with the D1 sentence obliged to say so. Zero calls.
A4 (finding 4, MANDATORY) — a check that high α is not bought by an easy bimodal item set. α_ord(F) and α_ord(R) are recomputed restricted to items whose cross-rater mean falls in that axis's interquartile band — the contested middle, which is the only place the seam question lives — and reported beside the full-set figure, with per-item score SD as a descriptive. If mid-range α collapses while full-set α is high, D1's sentence is qualified accordingly.
A5 (finding 5, BLOCKING) — stratum KO's generative procedure is frozen here, because the header claimed it was and the document did not contain it. Three things the critic correctly found missing, now binding:
- Work and span, by rule rather than by discretion. The work is 현진건 「운수 좋은 날」 (Hyŏn Chin'gŏn,
A Lucky Day, 1924) — public domain, freely reachable, and chosen before any result existed on
the stated criterion that Korean marks social relation grammatically (speech levels, honorific
-시-, humble benefactives) where classical Chinese marks it lexically, which is precisely why S049's person-reference stratum was empty. The span is mechanical: from the story's first sentence to the first paragraph break at or past 900 Korean words. No locus is chosen by the lead. Declared deviation: the work was chosen by the lead, not by a non-lead model or an external pointer. The criterion is written above and predates the material; the span rule removes discretion about where inside it to look. That is weaker than a non-lead choice and is recorded as such. - The extraction brief, verbatim, is committed before the extractor is called — as
design/extraction-brief.md, in its own commit, before dispatch. - The extractor sees the Korean source alone. It returns bracketed source spans; the lead then attaches each span's English rendering mechanically, by locating it in the frozen translation. The extractor never sees the lead's English, so sites cannot be selected for being places where the lead's translation made a visible decision. Sites whose source span cannot be mapped to a contiguous English stretch are dropped, and the number dropped is reported.
A6 (finding 6, MANDATORY) — H3's primary population becomes the frozen S049 partition. Splitting items by this run's CAT labels and testing graded scores from the same three models in the same session invites correlated-rater-error confounding — not mechanical circularity, as the critic says plainly, but real. Primary: §5's frozen partition (items 5, 12, 13, 15, 16, 18, 19, 21, 22, 24, 25, 30, 31, 38), labelled eleven sessions earlier. Secondary: this run's CAT partition, as a cross-session robustness check.
A7 (finding 7, MANDATORY) — H1's noise bar is stated within-metric, and the difference gets a CI.
The max of three single-draw |Δα| values is an unstable estimate and mixes metrics — it would have been
set by CAT's nominal swing, which S049 measured at 0.243, stacking the deck on top of A1. Graded gains
are judged against the graded conditions' own replicate swings and CAT against CAT's, and analyse.py
bootstraps a CI for each α difference over items.
A8 (finding 8, ADVISORY) — H5 is demoted to descriptive. §4 argues that fixed margins mean nothing against measured noise and then H5 registered a flat 0.50; with n possibly dropping to two raters under F4, a vacuous pass is the likely outcome. KO figures are reported with their replicate swings and carry no pass/fail.
A9 (finding 9, ADVISORY) — two sentences corrected. (a) §3's "This is the property CAT cannot
express" is false — CAT has both and composite. It should read: CAT can say both but not how
much of each, and not where on the scale the site sits. (b) §4's "every comparison below is
rater-matched and item-matched" holds for the within-run comparisons and for CAT-versus-S049; it does
not hold for H5's KO figures, which are item-matched to nothing, nor for the cross-session stability
figure, which is rater-matched and deliberately not session-matched. The sentence overclaimed and is
qualified here.
Dispositions, as re-gated by A1
- D1 — H2 passes (the derived four-way label's α_nom falls to CAT's level): the senses typology's unreliability is measured to live in the cut; the graded instrument is adopted as a standing instrument for locating seam sites, not for adjudicating them — qualified by A3 and A4 if either fires.
- D2 — H1 fails (the graded axes do not even screen better): the instrument is retired with its measurement and the S054 result is recorded as local to the drift typology.
- D3 — H1 passes and H2 fails (the derived label's α_nom stays well above CAT's): the senses case
is unlike the drift case, a cut through these axes reproduces, and
ARM-sense-boundary's closure gains a named successor rather than being reopened here. - D4 — H4 fails: double null; per the arm's own words, the claim's evidence class falls rather than being repaired.