Repository path: workshop/experiments/README.md · rendered 2026-09-09
Page metadata (front matter)
| type | note |
|---|---|
| id | experiments-readme |
| status | active |
| created | 2026-07-23 |
| updated | 2026-07-31 |
Experiments
Regime comparisons and evaluation studies. The discipline (charter §3) is non-negotiable for anything citable:
- Frozen design written before the run: question, materials, procedure, predictions, and what would count as failure. File:
design.md, committed before execution. - Independent pre-run critic pass — a critic agent (or non-Anthropic panel model) reviews the design; its critique and any design change are recorded. Changes after the critic pass re-freeze the design.
- The run, with all raw outputs preserved under
runs/. - Post-run verification — every reported number recomputed from the raw outputs by a fresh pass; discrepancies reported, not smoothed.
Honest nulls are first-class results (write them into wiki/findings/results/).
Layout: workshop/experiments/E-YYYYMMDD-<slug>/ with design.md, critic.md, runs/, analysis.md, verification.md.
Pilots (charter §10.7) are exempt from the full discipline but must be labeled status: pilot, evaluated only informally, and never cited as evidence.
Standing critic dispositions — check these before freezing a design
Added 2026-07-25 (S021) because NEXT.md note (n) named the gap: the critic passes work, the memory of them does not. Three of S020's twenty-one blockers were regressions — defects an earlier critic raised, that were accepted and fixed, and that a later design reintroduced. This list is what has already been established. Re-deriving one of these in a fresh critic pass is a process failure, not a discovery. Add to it whenever a pass establishes something general; keep each entry to the rule and its origin.
- No pseudo-replication. Repeated dispatches of the same payload to the same model are not independent units. Count units, not calls. (calibration-v1 critic; reintroduced S020 D-design.)
- The instrument is frozen before the design that uses it, not authored alongside it. An instrument written after the design can be shaped to the design's predictions. (calibration-v1; reintroduced S020.)
- Every prediction states its failure condition. A prediction with no way to fail is not pre-registered. (S010; reintroduced S020.)
- Order swap is mandatory on any paired-item judgment. Slot preference has been measured at 0.50–0.85 in this project. (S010.)
- A threshold must be reachable by the thing it is applied to, in both directions. Check that the rule can fire and can fail before pre-registering it. Three violations in S020 alone. (S020 note (o).)
- A control has to be built and measured, not asserted. Any page claiming a control is neutral, matched, or undamaged shows the measurement. (S020 note (l).)
- A control needs a case it should pass as well as a case it should reject. A negative control alone proves the instrument is not vacuous and nothing about what a pass means. (S021
critic.md6.2.) - The measured axes must be the axes the control is for. S021 v1 measured form, period, length, architecture and register; the axis that disqualified its candidate was an outright error. (S021.)
- Matchers, tokenizers and sentence splitters are frozen in the design page, verbatim. A gate whose acceptance region is coded after the freeze can only fail by implementer choice. (S021
critic.md1.2, 1.5, 4.6.) - Cost estimates are built from per-call maxima, and the run is sized so the worst case fits. (S020 note (m).)
max_tokensis sized for reasoning + output. Reasoning models here spend 1.7k–11k tokens before emitting anything; S021's critic spent 9,742. (S010 note (b).)- Never relax a charter-mandated control inside the experiment page that benefits from the relaxation. Open a decision page. (S020 critic F31; exercised correctly by S021 at
D-20260725-06.) - Verify an agent's factual claims about stored files by exact match before acting on them — and verify the tallies over them too. (S015 note (g); S021 verified 38/38.)
The cross-day drift gate — binding on every rating design from S072
Two wiki/backlog.md rows opened at S062 reached the review-or-retire rule at S072 and are discharged
here, by conversion into a condition at the point of use rather than by staying in a table where their
age was a number nobody had to act on (the S061 precedent, framework/control-arm-spec.md).
The finding they carry. Two instruments of different kinds — RS-20260730-grain-clause (categorical,
23 items, 9 labels) and RS-20260730c-revision-close §3 (graded, 216 items, 0–100) — both moved on
byte-identical requests one day apart, the second in 4 of 5 clean cells, with mean |Δ| up to 8.83
and α falling 0.82 → 0.72. Every κ, α and mean this project has published from a panel was measured
once. And the cheapest tripwire for it is already in every stored usage block and has never been
read: on the same byte-identical prompt a day apart, RS-20260730b-c16-redraw §4 saw one seat's
completion go 3,439 → 423 tokens and another's 1,903 → 2,879 while both returned the same 23
site lines.
The condition. A design that will publish an agreement statistic, a mean or a rate from panel output must do one of these, and say in its own text which:
- include a byte-identical repeat of one condition in the same session, and report the repeat's delta beside the headline figure; or
- state plainly that it has no same-day anchor, and that its figure is therefore a single measurement of an instrument known to move across days.
Either way, a design that reports a panel figure should also report completion length per seat, which costs nothing and is the drift signal already sitting in the stored bodies.
First application, and it is against this session. E-20260731d-sense-axes took option 2 by default
rather than by choice — it has no same-day repeat, and its α figures (0.521 CAT, 0.398 derived, 0.700 and
0.415 per axis) are single measurements. Recorded here because a gate whose first instance is an
exemption nobody wrote down is a gate that will not fire again.
What the drift gate does NOT measure, added 2026-07-31 (S073)
The gate above offers a same-day repeat, and the failures that motivated it were CROSS-day. That
was written into wiki/backlog.md as an owed cross-day repeat of E-20260730d's four decision
cells; the row reached the review-or-retire rule at S073 and is absorbed here, because the honest
place for it is beside the gate it qualifies rather than in a table.
So, stated plainly: option 1 bounds within-session instability and says nothing about the
day-to-day instability that RS-20260730-grain-clause (κ +0.285 and +0.356 on byte-identical
requests a day apart) and RS-20260730c §3 measured. A design that needs a cross-day figure must
dispatch the same condition on two different days and say so; no design in this project has, and the
quantity remains unmeasured.
S073 is the gate's first affirmative application — E-20260731e took option 1 by choice — and
what it bought is on the record: a floor of 0.375 on one seat and 0.000 on another, which
was larger than the effect the design was built to detect and which turned a reportable-looking
result into a null.
Standing materials condition, absorbed from wiki/backlog.md at S074 (row opened S064, review-or-retire fired)
A study whose items are derived from a translator's log may not treat lead-written glosses as a samplable nuisance. RS-20260730e §6 registered a gloss control, withdrew it on the critic's Finding 3, and then found the pool held 2 glossed items in 37 — too few for any statistic, and the structural reason is worse than the sample: a gloss exists exactly where the translator judged the source needed one, so glossing is confounded with difficulty by construction.
Either the item set is authored by someone other than the translator, or the design states in advance that the leak is uncontrolled and reports the affected count. There is no third option, and sampling is not one.
The general form of the repair is the one E-20260731f stage 2b used against its own pre-run critic's BLOCKING finding: put the classification in hands that are not the author's, and report the disagreement as a number. It cost two calls, and it moved a reported figure from 18 of 41 to 1 of 41.