Repository path: workshop/experiments/E-20260728g-scale-usage/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260728g-scale-usage |
| status | frozen |
| created | 2026-07-28 |
| updated | 2026-07-28 |
| provisional | true |
| internal-judgment-only | true |
| senses | accuracy, naturalness, voice, style-correspondence, literary-quality, cultural-mediation |
| links | wiki/arms/ARM-tierD-repair.md, wiki/findings/results/RS-20260727b-tierD-rules.md, wiki/findings/results/RS-20260726d-tierD-heldout.md, workshop/experiments/E-20260726d-tierD-heldout/design.md, workshop/experiments/E-20260725-tierD-ladder/design.md, workshop/translations/bargamot/R04-v1/translation.md, config/models.md, wiki/backlog.md |
Frozen design (v2) — settling the per-juror scale-usage gate
v1 was rejected by the independent pre-run critic pass (critic/dispositions.md, verdict NEEDS-REDESIGN, nine findings). This is the amended version: all nine accepted, five of them by withdrawing or narrowing a claim v1 had made. Dispositions are tabulated in critic/dispositions.md and dated in §11. The analysis may proceed only against this version.
v1 was frozen 2026-07-28 (S050) at a6100be, a commit that contains this page, T-bargamot-R04-v1 and T-bargamot-R06-v1, and contains no analysis script, no result page and no fault variant. ARM-tierD-repair step 2, condition (ii) of wiki/backlog.md's merged owed entry.
No API call is made by the analysis. Every number below is recomputed from stored responses of two runs that are already paid for. The one API call this design authorises is the independent pre-run critic pass (§9), dispatched after this page is frozen and before any statistic is computed.
0. What this is and what it is not
ARM-tierD closed naming §10's pooled scale-usage gate as defective, and S040 discharged half of that: RS-20260727b-tierD-rules §4 showed that 56.8% of the pooled sum of squares is between jurors, so the pooled statistic counts juror disagreement about the mean as evidence that a juror can move. S040 stated the repair in form — per juror, on within-juror dispersion, every juror passing — and deliberately did not settle the threshold, writing: "0.75 was chosen for a pooled statistic and inherits nothing." That sentence is this page's whole job.
This is not a rescue of S034 and not a Tier D run. The S034 verdict stands (RS-20260726d, §1), no repaired rule is applied retrospectively to it, and Tier D remains NOT PASSED. Nothing here calibrates anything (charter §5.5).
This is not a claim about any translation. No quality claim is asserted about T-bargamot-R04-v1 or about any stored text. The lead never judges its own translation (charter §5).
1. Questions
Q1 — siting. §10 reads the gate off the sham stage, whose two texts differ by eight substitutions constrained to be quality-neutral, word-count-neutral and sense-preserving. Is within-juror dispersion measured there representative of the same juror's dispersion on the arms the gate licenses?
Q2 — statistic. The gate exists so that a 0.75-point margin between two texts is measurable (§6.6). Dispersion of levels and capacity to express differences are different quantities. Do they come apart in the stored data?
Q3 — threshold, and this is the step the arm asks for. Is there a threshold on the per-juror form of §10's statistic that (a) fails a juror who cannot express the criterion margin and (b) does not fail a juror who demonstrably can?
Q4 — replacement. If Q3 has no answer, what should the gate be instead, and what materials does it need?
2. Materials — stored, paid for, and not re-run
Two independent jury runs, the same three jurors (P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P5 deepseek/deepseek-v4-pro), the same six senses, the same 1–7 scale, the same paired-text item format with mandatory order swap.
| run | session | responses | arms present |
|---|---|---|---|
E-20260725-tierD-ladder |
S020 | 60 | sham · targeted (O1, O2, O3, O4) |
E-20260726d-tierD-heldout |
S034 | 60 | sham · targeted (heavy, light) · held-out |
120 responses, 12 scores each = 1,440 scores. Every score is re-parsed from the model's own returned text, not from the runner's cached _parsed, exactly as E-20260726d/analysis/score.py does (§11 of that design).
S020 is the cleaner test of Q1 and is named as such in advance. Its sham and targeted arms sit on the same two passages (A, B), so an arm difference there cannot be a passage difference. S034's sham items (S-GA, S-K) and targeted items (T3/T8-GA/HB/K) overlap in reference but are not the same set, so S034 is a replication with a known confound, not an independent clean test. Any disagreement between the runs is reported and is not resolved in favour of either.
Slot convention, taken verbatim from the S034 analysis and to be re-asserted by the verifier against each run's own items manifest: order 0 → slot A is the reference; order 1 → slot B is the reference.
3. Definitions, fixed here
For a given (run, arm, juror):
- L — every raw 1–7 score that juror gave in that arm: (items × 2 orderings × 2 slots × 6 senses). n = 48 per juror per sham arm in both runs.
- SD(L) — population standard deviation of L. This is §10's statistic, computed per juror rather than on the pool. It is the quantity the repaired gate is stated on.
- SD(L|sense) — (added at v2, critic A3) the same dispersion with each sense's own mean removed within the cell. SD(L) pools numbers that are not exchangeable: a juror who marks
cultural-mediationtwo points aboveaccuracyon every item scores dispersion for doing so. SD(L|sense) is the per-juror analogue of exactly the decompositionRS-20260727b§4 applied to the pooled statistic, and its absence from v1 was an inconsistency. The gap SD(L) − SD(L|sense) is how much of a juror's apparent scale usage is sense-level offset. - I(L) — number of distinct integers in L. R(L) — realised range, max − min.
- D — the paired differences: for each (item, ordering, sense),
d = score(reference) − score(second). n = 24 per juror per sham arm. - SD(D), mean|D|, and z(D) = fraction of d equal to 0.
- maxD — the largest |d| that juror produced in that arm.
- M — (added at v2, critic C2) the realised sense margin: for each sense, the mean of
dover that cell's items and orderings; M is the largest of those six means. M is the quantity §6.6's 0.75 criterion is actually stated on, and it is what the gate exists to certify a juror can produce.
The unit rule, quoted verbatim from E-20260726d-tierD-heldout/design.md §6.3 rather than referred to (added at v2, critic D3): "unit = (juror × item), the two orderings combined. +1 reference preferred in BOTH orderings, −1 second preferred in both, 0 split." Per-juror detection for §4.5 is the count of +1 units for that juror on that run's targeted items, against that juror's total targeted items.
4. Procedure, in this order
- This page is frozen. (Commit contains it; contains no
analyse.py, no result page, no fault variant.) - Independent pre-run critic pass (§9). Dispositions written before step 3.
- Compute the §3 quantities for every (run, arm, juror) cell.
- Test the four predictions of §5.
- Per-juror detection outcomes for both runs, computed the same way (§6.3 units of the S034 design; the analogous rule in S020), because prediction 3 turns on them.
- Verify with a second script that recomputes every reported number from
runs/alone and shares no code with step 3. - Build the fault variant
variant-F1.mdonT-bargamot-R04-v1Unit B, to the repaired specification ofRS-20260727b§6.2–6.3, and characterise it on the two axes that result named. This step happens last and is not conditional on any result above — it is materials, not a test.
(Added at v2, critic E2.) The variant is used by no statistic on this page. It is frozen in its own commit, and any future design that uses it as a control must be frozen after that commit — the constraint travels with the artifact rather than with this page. Every site must instantiate one of O4's four documented failure types (wrong referent, wrong word sense, invented detail, dropped negation), whose catalogue basis is external to the lead by charter §5.2 — A-shaw-spider-thread, Swann 1974 on Turney, A-chekhov-pari on Koteliansky & Murry 1915 — and the type is named per site. A site instantiating none of them is not admissible.
5. Predictions, written before anything is computed
(P2, P3 and P4 restated at v2 on critic findings A2, A4 and C2. The v1 wording is preserved in §11.)
P1 — arm effect on dispersion. For every juror in both runs, SD(L) on the targeted arm exceeds SD(L) on the sham arm. Falsified by any juror in either run where it does not.
Predicted magnitude, so that "large" is not decided after the fact. P5's sham SD(L) in S034 is already published at 0.143 with two integers used (RS-20260727b §4). The prediction that goes with P1 is that P5's targeted SD(L) in the same run is at least twice that.
P2 — levels and differences come apart (restated, critic A4). On the sham arm of at least one run there is a juror with z(D) ≥ 0.50 — half or more of that juror's paired differences are exactly zero — who is not the juror with the lowest SD(L) in that run. Falsified if in both runs every juror with z(D) ≥ 0.50 is also the run's lowest-SD(L) juror, or if no juror reaches z(D) ≥ 0.50. v1 conjoined SD(L) > 0.40, a bar any spread clears; the rank contrast replaces it because it can fail.
P3 — the gate's operating characteristic (restated, critic A2 — v1's "invalid" reading is withdrawn). Applying SD(L) ≥ 0.75 per juror, all jurors passing to each run's sham arm fails both runs, while in each run every juror individually went on to produce that run's targeted detection outcome. Falsified if either run passes the gate, or if any juror's targeted detection is not recoverable per juror from stored data (§7.4). What this licenses is a false-negative count on the only two cases that exist — a cost of the threshold, not a proof that the statistic is the wrong one. The wrong-statistic claim rests on P2 and P4.
P4a — separability (split from v1's P4, critic C2). Every (run × juror) cell reaches M ≥ 0.75 on its targeted arm. Falsified by any cell below it. If P4a holds, then no juror the project has ever run is incapable of the criterion margin, so no positive threshold on any prior statistic is licensed by this data — the gate would have to pass everybody.
P4b — monotonicity (split from v1's P4). The ordering of the (run × juror) cells by sham SD(L) does not match their ordering by targeted M. Falsified if the two orderings agree within each run. v1 stated a non-existence claim over all thresholds and then tested a rank ordering, which is necessary and not sufficient; the two halves are now separate claims with separate falsifiers.
6. What each outcome licenses — pre-registered readings
(P1 and P3 rows narrowed at v2 on critic findings A1 and A2; split-outcome and UNEVALUABLE rows added on C3.)
| outcome | licensed statement |
|---|---|
| P1 holds | The gate is read off the arm on which these jurors' dispersion is lowest. §8 concedes that heterogeneity and damage-presence are the same thing in this data, so no claim is made about why. Under either mechanism the consequence is the same and is the whole point: a dispersion figure read on the sham arm does not describe the juror's dispersion on the arms the gate licenses. v1's "sited on the arm where the jurors move least" claimed a general property of neutral material and is withdrawn. |
| P1 fails | Dispersion does not vary with arm for these jurors; the gate is well sited and only the number is open. |
| P2 holds | SD(L) and difference capacity are dissociated within a run: a juror can rank above another on dispersion of levels while expressing fewer differences. A threshold on SD(L) cannot be derived from a criterion stated on differences. |
| P2 fails | The two orderings agree; SD(L) is a usable proxy for difference capacity on this data, and the threshold question is only about the number. |
| P3 holds | This gate, at this inherited threshold, would have blocked both runs — and in both, every juror separately went on to produce the targeted detection outcome. That is a false-negative count of 2 of 2 on the only cases available, i.e. a cost of the threshold. It is not a demonstration that the statistic is the wrong one; a prior gate that blocks a run which then succeeds may simply be conservative. |
| P3 fails | The inherited threshold survives its first check; step 2 may close by keeping 0.75 per juror, with the check on record. |
| P4a holds | Every juror the project has run reaches the criterion margin when there is damage to find. No positive threshold on a prior statistic is licensed by this data, because there is no incapable juror for it to exclude. The honest settlement of Q3 is then "no discriminating threshold exists on this panel", which is an answer and not a deferral. |
| P4a fails | At least one juror cannot produce the margin. A gate has something to exclude, and P4b decides whether SD(L) can find it. |
| P4b holds | SD(L) on the sham arm is not monotone in the realised margin. No cut on it can order jurors the way the criterion does. |
| P4b fails | SD(L) is monotone in the realised margin on this data; a cut exists, and the result page must report it together with the fact that it rests on six cells. |
| split outcome (added v2) | A prediction holding in one run and failing in the other is reported as a split and is not resolved in favour of either run. S020 is the cleaner test of Q1 (§2) and that is stated when it applies; it does not make S020 the arbiter of P2, P3, P4a or P4b, none of which turn on the passage confound. |
| UNEVALUABLE (added v2) | Any prediction whose §7 failure criterion fires is recorded UNEVALUABLE and no statement of its licensed form is made in either direction — the S049 precedent, binding here in advance. |
The prohibition, rewritten at v2 to bind by form of inference rather than by word list (critic E1). v1 banned calibrating, validating and passing, and the critic correctly observed that the conclusions being reached — "the gate is sited…", "invalid at its inherited threshold" — are validation-shaped sentences in other words. The binding form is this: every statement this analysis makes must be expressible as "on these 1,440 stored scores, statistic S at threshold T would have decided D", with S, T and D named. Any sentence that cannot be rewritten in that form is not licensed, whatever words it uses. Nothing here calibrates, validates or passes anything, and Tier D remains NOT PASSED.
7. Failure criteria — what voids or qualifies this analysis
- Re-parse mismatch. If re-parsing any stored response's own text yields a different score-set from the runner's cached
_parsed, the affected cell is reported as a discrepancy and the analysis states its count before any other number. A discrepancy rate above 5% voids the run. - Slot-convention failure (rewritten at v2, critic D1). If the order→reference mapping cannot be re-derived from a run's own items manifest, that run is excluded entirely and the exclusion is the headline of the result. v1 let such a run "contribute to L-statistics only", which silently changes the composition of the cells being compared; a loud exclusion is the safe failure.
- Cell size. Any (run, arm, juror) cell with fewer than 24 scores in L is reported and excluded from threshold arithmetic.
- P3 is not evaluable if the two runs' detection outcomes are not per-juror recoverable from stored data. In that case P3 is recorded
UNEVALUABLEand no statement of its licensed form is made — the S049 precedent, and it binds here in advance.
8. Known threats, stated in advance
- Two runs, three jurors, one model generation. Nothing here generalises to other jurors. Every statement is about these three models on these materials.
- The arms differ in more than heterogeneity. A targeted arm's second text is damaged; a sham arm's is not. "Material heterogeneity" and "presence of damage" are the same thing in this data and cannot be separated by it. The claim P1 licenses is therefore about siting, not about a general psychophysics of these jurors. This is the sharpest limitation on the page and it is not repairable from stored data.
- The held-out arm sits in only one run and is reported for completeness, not compared across runs.
RS-20260726dreports the S034 sham arm as out of band (0 of 6 units preferred the unedited text), which is a fact about that arm's materials. It does not affect a dispersion statistic and is not used to discount one.- The lead is measuring an instrument the lead's own project designed. No anchor licenses any of this; the page carries
internal-judgment-onlyandprovisional: true.
9. The one API call
Independent pre-run critic, one call, on this frozen page. Role: a non-Anthropic panel model that is not one of the three jurors whose data is analysed here — P1, P2 and P5 are all subjects, so the critic must come from outside them. Resolved from config/models.md: P3 x-ai/grok-4.5, with P4 moonshotai/kimi-k3 explicitly off the call list (note (b); it has burned two full budgets returning nothing) and first reserve qwen/qwen3.7-max on an empty or length return, per S044/S045/S046's own lesson: fall through, do not retry.
Pre-flight estimate, built from max_tokens and not from an assumed output length (note (abc)): max_tokens 8000 at P3's listed $6.00/M out = $0.048, plus ~5,000 prompt tokens at $2.00/M in = $0.010. Worst case $0.058. Today's headroom before this call: $3.513016 of the $5.00 UTC-day cap.
reasoning: {"effort": "low"} is sent, and the prompt instructs brevity — both are standing lessons and neither is a guarantee (S044: a provider silently ignored the parameter).
10. Verification
A second script, sharing no code with the analysis, that recomputes from runs/ and the items manifests alone:
- the count of responses loaded, per run, per arm, per juror;
- every SD(L), I(L), R(L), SD(D), mean|D|, z(D), maxD reported;
- the re-parse discrepancy count of §7.1;
- the slot convention, re-derived from each run's items manifest — and, independently of the manifest, the sign check the critic's finding D2 produced:
mean d > 0on every targeted cell. The reference is the undamaged text, so an inverted convention flips everydand fails this check loudly without reading the manifest at all; - the per-juror detection outcomes of §4.5, by an independent path;
- a second regression on the D-pairing itself (added at v2, critic D2): the count of paired differences in one named cell (
S034 / sham / P1), derived independently as items × orderings × senses, checked against the length of D as the analysis built it; - the between/within variance decomposition of
RS-20260727b§4 on the S034 sham arm — 43.2% within, 56.8% between, pooled SD 0.783 → 0.515 — as a regression test that this analysis is reading the same data S040 read. A mismatch here is a defect in this analysis, not a correction to that one, until shown otherwise.
Every check prints PASS/FAIL and the script exits non-zero on any failure.
11. Amendment record (v1 → v2, 2026-07-28, S050)
v1 frozen at a6100be. Critic pass dispatched against it; verdict NEEDS-REDESIGN, nine findings, all nine accepted. Full dispositions: critic/dispositions.md. Changes:
| § | change | finding |
|---|---|---|
| 3 | SD(L|sense) added — dispersion with sense means removed within the cell | A3 |
| 3 | M added — the realised sense margin, the quantity the 0.75 criterion is actually stated on | C2 |
| 3 | the S034 unit rule quoted verbatim instead of referred to | D3 |
| 5 | P2 restated: the inert SD(L) > 0.40 conjunct replaced by a rank dissociation that can fail |
A4 |
| 5 | P3 restated: the "it is invalid" reading withdrawn; P3 now licenses a false-negative count and nothing more | A2 |
| 5 | P4 split into P4a (separability) and P4b (monotonicity), each with a falsifier that matches its claim | C2 |
| 6 | P1's row narrowed: no claim about why sham dispersion is lowest; v1's neutrality claim withdrawn | A1 |
| 6 | P3's row narrowed to an operating characteristic | A2 |
| 6 | split-outcome and UNEVALUABLE rows added | C3 |
| 6 | the vocabulary prohibition rewritten to bind by form of inference, not by word list | E1 |
| 7.2 | the "L-statistics only" fallback deleted; an unrecoverable slot convention now excludes the run entirely | D1 |
| 4.7 | the fault variant's external-catalogue constraint and its freeze-order constraint written onto the step | B, E2 |
| 10 | the manifest-independent sign check (mean d > 0 on targeted cells) and a D-pairing regression added |
D2 |
v1 wording preserved for the two withdrawn claims, because a withdrawal that deletes what was withdrawn is not a record:
- v1 P3: "the repaired gate is not merely stricter; it is invalid at its inherited threshold… A gate that stops a run whose detection subsequently fired in every juror separately is not conservative; it is measuring the wrong thing."
- v1 §6, P1 row: "The gate is sited on the arm where the jurors move least. §10 reads scale usage off material constructed to be indistinguishable, then uses it to certify that a juror can distinguish. The siting is the defect, independently of the threshold."