Repository path: workshop/experiments/E-20260727-log-decision-coding/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260727-log-decision-coding |
| status | frozen |
| created | 2026-07-27 |
| updated | 2026-07-27 |
| senses | accuracy, naturalness, voice, style-correspondence, affect, literary-quality, cultural-mediation, purpose-fit, consistency |
| links | wiki/arms/ARM-typology-logs.md, workshop/experiments/E-20260727-log-decision-coding/codes.md, workshop/experiments/E-20260727-log-decision-coding/decisions.tsv, wiki/goodness-senses.md, config/models.md, config/budget.md |
| internal-judgment-only | true |
| provisional | true |
E-20260727 — is the decision-class scheme shared, or is it one coder's?
Frozen before any API call. Written after codes.md and decisions.tsv existed and before either was shown to anything outside this session.
What is already done, and what is not an experiment
Steps 1–3 of ARM-typology-logs are a reading, not an experiment. Twenty frozen logs were read end to end; 361 decisions were extracted into decisions.tsv; fourteen classes were derived from the material and mapped onto the nine senses in codes.md. There was no pre-registered design and there could not have been one: open coding derives its categories from the corpus, and a category set fixed in advance would be the thing the arm exists to avoid.
Two things were nonetheless pre-committed, and they are what keeps the reading from being unfalsifiable:
- A held-out log.
T-rayo-de-luna-R04-v1(Spanish → English, a pair the project had never worked) was translated, logged and frozen at commitb4cf674before any of the twenty prior logs was reopened. Its 19 decisions were coded only after the fourteen classes were fixed. A class appearing in the held-out log and absent from the derivation set would be a saturation failure. - The mapping is at class level and written down. The residue claim is a claim about four named classes, not about individual rows, so it can be attacked by disputing fourteen judgments rather than by re-litigating 380.
The question this experiment does ask
Is the fourteen-class scheme shared, or is it idiosyncratic to the lead?
The scheme was derived by one coder, applied by the same coder, over a corpus in which one of the twenty-one logs is that coder's own work from this session and the other twenty are that coder's own work from previous sessions. That is the weakest joint in the unit. If an independent coder given only the class definitions cannot reproduce the assignments, the class-level percentages in RS-20260727-log-typology are a report of one agent's intuitions and must be labelled as such.
Materials
- The instrument:
codes.md§"The classes" — the fourteen definitions with their examples, verbatim. The mapping table is withheld: a coder who knows which classes are the residue could route items toward or away from them. - The items: 60 decisions drawn from the 361-row derivation set by systematic sampling, every 6th row from row 1, so the draw is reproducible and neither the lead nor the coder selects it. The held-out log is excluded. Item text is the
decisioncolumn verbatim;log,lang,contamandidare withheld. - The coders: panel roles P1 and P2 (
config/models.md), both non-Anthropic, one call each,temperature0, blind to each other and to the lead's codes.
Procedure
- Sample the 60 items with
sample.py(fixed rule, no randomness), writingsample.json. - One call per coder: the fourteen definitions, the 60 items, and the instruction to return exactly one class id per item and nothing else, in a JSON object keyed by item number.
- Preserve raw request/response JSON under
run/. score.pyrecomputes every number reported, importing nothing from the analysis script.
Predictions and failure criteria — pre-committed
Primary measure: raw agreement with the lead's code, per coder, over 60 items. Chance is 1/15 ≈ 6.7% (fourteen classes, C4 split in two).
- ≥ 70% on both coders → the scheme is reproducible from its written definitions; the percentages may be reported as properties of the corpus.
- 50–69% → reported as weak; percentages carry a stated caveat.
- < 50% on either coder → the scheme is not shared. The class-level percentages are withdrawn as corpus properties and reported only as the lead's coding, and the residue argument must stand on the four named classes' definitions rather than on their counts.
Secondary measure, and the one the finding actually depends on: residue-class retention. For items the lead coded C4b, C11, C12 or C13, what fraction do the independent coders also place in a residue class?
- Pre-committed: if independent coders route residue-coded items into the eleven non-residue classes at a rate materially above the overall disagreement rate, the residue claim is damaged and
RS-20260727-log-typology§3 must say so. "Materially above" is fixed here as residue retention below (1 − overall disagreement), i.e. residue items disagreeing more than average items. - The reverse is also pre-committed and is the more likely failure: coders may put residue items into
C1, becauseC1is large and inviting. That would be a real result and is not to be explained away.
What this experiment cannot do. It cannot show the classes are right, only that they are legible. Two models agreeing with the lead is not validation (charter §4 forbids treating panel agreement as validation); it is evidence in the failing direction only — disagreement would be informative, agreement is weak.
Pre-flight cost estimate
Built from the max_tokens cap actually sent, per method note (abc) — not from an expected output length.
| in (tok) | out cap | list rate | list cost | ×4 routing worst case | |
|---|---|---|---|---|---|
P1 openai/gpt-5.6-terra |
~2,600 | 2,000 | $2.50 / $15.00 | $0.037 | $0.148 |
P2 google/gemini-3.6-flash |
~2,600 | 2,000 (+~1,700 hidden reasoning, billed as output) | $1.50 / $7.50 | $0.032 | $0.128 |
Worst case total $0.28, against $5.00 of untouched headroom on 2026-07-27. The ×4 factor is the measured routing caution in config/models.md. Actuals recorded from usage.include and cross-checked against the key-usage delta.