Repository path: workshop/experiments/E-20260730d-register-centre/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260730d-register-centre |
| status | frozen |
| created | 2026-07-30 |
| updated | 2026-07-30 |
| senses | naturalness |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-evidence-audit.md, wiki/goodness-senses.md, wiki/base/anchors/A-mansfield-garden-party/A-mansfield-garden-party.md, wiki/base/anchors/A-doctorow-little-brother/A-doctorow-little-brother.md, wiki/base/anchors/A-mchugh-presence/A-mchugh-presence.md, workshop/translations/postmaster/R07-v1/translation.md, workshop/regimes/R07-fluency.md, wiki/base/sources/S-venuti-invisibility.md, wiki/findings/open-questions/OQ-20260723-target-register.md |
E-20260730d — is the unmarked centre of naturalness a register, or the absence of two?
Frozen before any call. ARM-evidence-audit step 3, the arm's last. Nothing below may be reworded after this commit; the pre-run critic's findings are applied as numbered amendments appended to this file, never as edits to the text above them.
The question the arm asks
wiki/goodness-senses.md's naturalness entry says the sense is register-relative and anchors it on two marked poles — A-mansfield-garden-party (period-idiomatic, 1922) and A-doctorow-little-brother (contemporary-vernacular, 2008) — and then says: "the vernacular end of contemporary and a neutral/literary-contemporary register are still un-anchored." ARM-evidence-audit's completion criterion item 3 is that the unmarked-centre anchor exists, or the arm records in writing why the project should not build one.
The armchair argument for not building one is available and this experiment exists because it is an argument and not a measurement. It runs: an anchor's job is to supply a catalogue of positive features read off Tier 1 target-language writing; the defining property of the unmarked centre is that it has no markers; so a centre catalogue would consist of the features the two poles share, minus the poles' markers, and would be derivable from the two anchors already built rather than adding independent evidence.
This experiment decides between that and its negation, on the arm's licensed side of the S015 line. The panel is used only for is this property present in this passage — a factual-adjudication task, where the four panel members used scored 0.886–0.917 discrimination against planted false claims — and never for is this passage good. A step that drifts across that line is outside the arm's licence.
Materials — seven passages, roles declared before anything runs
Built by build_passages.py, which declares each source file's header convention and paragraph convention per file rather than sniffing them, asserts the first words of each body, and cuts by one mechanical rule for all seven: whole paragraphs while under 1,500 words; if the next paragraph would pass 1,560, that paragraph is cut at the last sentence boundary keeping the total ≤ 1,560. Hard line wrapping is removed where a file has it, so line-wrapping cannot be read as a property of any passage. SHA-256 of every passage in materials/passages/manifest.json.
| id | text | year | role |
|---|---|---|---|
| A | Mansfield, "Miss Brill" | 1922 | marked pole — period-idiomatic; the existing anchor |
| B | Doctorow, Little Brother ch. 1 | 2008 | marked pole — contemporary-vernacular; the existing anchor |
| C | McHugh, "Presence" | 2005 | candidate centre 1 |
| D | Kessel, "The Snake Girl" | 2008 | candidate centre 2 |
| E | Kessel, "The Last American" | 2008 | marked control, same author as D |
| F | McHugh, "Frankenstein's Daughter" | 2005 | marked control, same author as C |
| G | lead, T-postmaster-R07-v1 |
2026 | R07 fluency-rule translation — Venuti's own description of the unmarked centre, executed |
Word counts 1,505–1,548, spread 2.9%. C, D, E, F are CC BY-NC-SA 3.0 (Small Beer Press), storable under the grant on the same footing as B; A is public domain.
Why two candidate centres, each with a same-author marked sibling. One centre text cannot distinguish this register has no profile of its own from this author has no profile of her own. E and F are by the same two authors, from the same two books, cut by the same rule, and are strongly marked in different directions — E is a pseudo-documentary register (a hagiography reviewed by a future archivist), F a first-person teenage vernacular. If C and D come out featureless while E and F do not, the featurelessness is a property of the register.
Why G. R07 is the project's execution of Venuti's review-corpus checklist — current, standard, idiomatic, continuous, precise, resolved — which wiki/backlog.md's row on this anchor calls "a characterisation of exactly the unmarked published English that Anglo-American reviewing rewards." G lets the census ask, factually, whether English written to that description carries the features of published unmarked-centre writing or the features of the poles. G is not judged. Feature presence is a textual fact; the lead never judges its own translation (charter §5) and no seat here is asked to.
Procedure
Order of operations, stated because the critic is not first. This design is frozen by its commit; then the blind enumeration runs, because it is materials construction and its output is an input the critic needs to review; then the pre-run critic reviews this design and the 56 predicates; then amendments are applied in writing; only then does the census, which is the measurement, run. The critic therefore precedes every measurement and follows one materials-construction stage — the same shape as S062, where the blind control builder ran before the critic saw the items it had built. What the critic cannot block is the enumeration brief, which is frozen here, in this file, before it runs.
Stage 1 — blind feature enumeration. qwen/qwen3.7-max, a materials-construction seat: not a rater, not the raters' declared reserve, not the critic (the S053 role-collision fix). Seven separate calls, one passage per call. Identical brief. The model is told nothing about the other passages, nothing about register, nothing about poles or a centre, nothing about a hypothesis, and not that any comparison exists. It is asked for exactly 8 predicates about the English of the passage, each of which
- names a concrete, inspectable property of the language (diction, syntax, tense, person, figuration, punctuation, morphology, register markers);
- could be checked as present or absent in a passage about a completely different subject;
- names no character, place, plot element, subject matter or proper noun.
This is the S062 blind-build discipline: the lead does not write the items, does not know them before they exist, and cannot have written the translator's log to target them — the log was frozen at 001405b and the predicates do not exist until this stage runs.
Stage 2 — the pre-run critic's screen (part of the critic call, before any census). The critic is given the 56 predicates with their sources stripped and asked to flag any that are content-specific in violation of the brief. Flagged predicates are dropped, and the count dropped is reported.
Stage 3 — cross-census. The surviving predicates are pooled, shuffled with fixed seed 20260730, and their sources stripped. For each of the seven passages, two raters — P1 openai/gpt-5.6-terra and P3 x-ai/grok-4.5 — answer PRESENT / ABSENT / UNCLEAR for every predicate. 14 calls, temperature 0, max_tokens 6,000. Judgment is not parallelised across raters within a cell; each cell is one independent call.
Each rater is asked, after its census answers, whether it recognises the passage and to name work and author if so. Recognition is measured, not assumed — ARM-evidence-audit's standing constraint says authorship-stripping is not blinding on famous public-domain texts and that every step must measure rather than assume.
Stage 4 — same-day byte-identical repeat. Cell (C, P1) is re-sent byte-identical, same day, same parameters. Registered here, before the run, as a standing condition — S061 and S062 each found a rating instrument that did not reproduce across a day, on two different instrument kinds, and every κ, α and mean this project has published from a panel was measured once.
Statistics, frozen
Consensus rule: a predicate is PRESENT in a passage iff both raters say PRESENT; ABSENT iff both say ABSENT; otherwise SPLIT. All three counts are reported for every cell.
Let own(X) be the predicates enumerated from X that are PRESENT in X by consensus.
H(X)— own-text hit rate:|own(X)|out of the 8 (or fewer, after the screen) enumerated from X. A validity check: a predicate read off a text must hold of that text.Dist(X)— distinctiveness: the fraction ofown(X)that is ABSENT by consensus in both A and B.Excl(X)— pole-exclusion: the fraction ofown(A) ∪ own(B)that is ABSENT by consensus in X.Asym(X)=Excl(X) − Dist(X).
Predictions, registered, with a decision rule that can return three answers
H1 — residual. The centre has no positive profile; it is the poles minus their markers. Then Dist(C) and Dist(D) are low, Asym(C) and Asym(D) are strongly positive, and both exceed Asym(E) and Asym(F).
H2 — third register. The centre is a register with features of its own. Then Dist(C) and Dist(D) are comparable to the marked texts' and are not distinguished from Dist(E), Dist(F).
Decision rule. Let m = min(Dist(E), Dist(F)).
- H1 supported iff
Dist(C) < mandDist(D) < m. - H2 supported iff
Dist(C) ≥ mandDist(D) ≥ m. - Undecided otherwise — one centre above and one below. This is the outcome the two-centre design exists in order to be able to see, and it will be reported as undecided rather than resolved by picking the centre that agrees with the lead.
Secondary, and reported whichever way the primary goes. Where does G sit? Of own(A) ∪ own(B), how many predicates are PRESENT in G, against how many in C and in D; and Dist(G) against Dist(C), Dist(D).
Failure criteria — any one fires and no headline is reported, only estimates
- Inter-rater raw agreement over all cells < 0.70. Cohen's κ is computed and reported alongside but does not gate, because the base rate is expected to be skewed and κ and raw agreement have already been observed disagreeing in this project (
RS-20260727e§3). - The repeat control flips the decision. The entire decision rule is recomputed with the repeat cell's answers substituted for the original (C, P1) answers. If the decision changes, the headline is withheld. The raw per-predicate flip count is reported descriptively. This is the tolerance-sizing fix that S062's critic had to impose after the fact: the gate is not a tolerance in different units from the effect — it is the effect, recomputed.
- Mean
Hacross the seven passages < 5 of 8 (62.5%). The enumerator's predicates do not hold of the texts they were read off, so nothing downstream means anything. - The critic's screen drops more than 14 of 56 predicates as content-specific.
What this experiment cannot establish, stated before it runs
- It is not a calibration. Tier D has not passed; nothing here licenses any quality claim about any of the seven passages. Every number is a presence count.
- n = 1 per role per author. Two centre texts and two marked siblings is enough to separate register from author only if these four texts are representative of their registers, which is not established and cannot be by four texts.
- Recognition is a live confound and is measured, not removed. A and B are famous; C, D, E, F are CC-licensed and openly downloadable and so may also be in training data; G is two hours old.
- The predicate set is a sample of one enumerator's attention. A different enumerator, or a different brief, would produce a different 56. What the design protects against is the lead choosing them, not enumerator variance.
- "Unmarked" is not "good." Venuti's argument is that this register is a norm Anglo-American reviewing rewards, not a virtue; the backlog row that scheduled this work says so and the anchor page must repeat it.
Pre-flight cost estimate — worst case from max_tokens, note (abc)
Prices read from GET /api/v1/models this session (2026-07-30) rather than from the table: P1 $1.25/$7.50, P3 $2.00/$6.00, P4 $3.00/$15.00, P2 $1.50/$7.50, qwen/qwen3.7-max $1.475/$4.425 per M in/out.
| stage | calls | max_tokens |
worst case |
|---|---|---|---|
| pre-run critic (P4) | 1 | 16,000 | $0.261 |
| blind enumeration (qwen) | 7 | 3,000 | $0.303 — see note |
| census P1 | 7 | 6,000 | $0.352 |
| census P3 | 7 | 6,000 | $0.311 |
| repeat control (P1) | 1 | 6,000 | $0.050 |
| total, no reserve | 23 | $1.277 | |
| declared reserve P2, if a census seat fails | ≤3 | 6,000 | $0.154 |
| total with reserve | ≤26 | $1.431 |
Note on the enumeration row, because the honest figure is not the cap-literal one. At max_tokens 3,000 the cap-literal worst case is $0.0168 per call, $0.118 for seven. But at S062 this same slug billed 8,902 output tokens against a cap of 4,000, the excess being hidden reasoning tokens. The figure declared above uses 3× the cap, the S062-observed ratio. Both numbers are stated; the larger is the one being declared, which is what note (abc) requires.
Today's ledger before this run: $0.9710133477 of $5.00, headroom $4.0289866523. Opening key snapshot 29.516047048, persisted at runs/snap/key-open.json.
Amendments (appended after the pre-run critic; nothing above is edited)
Pre-run critic: moonshotai/kimi-k3 (P4), one call, accepted first time, Together, in 5,351 / out 6,866, 191 s, $0.1188702. Verdict NEEDS-REDESIGN — two BLOCKING, four MANDATORY. All six findings are accepted; one is narrowed with a stated reason. Raw at runs/critic.txt, prompt at runs/critic.prompt.txt. This is the second consecutive NEEDS-REDESIGN in this project and the third ever (S021, S062, S063).
The two BLOCKING findings are answered by one change, and it replaces the primary statistic. Finding 2 is correct and is the finding of the session so far: because the predicates were enumerated from each passage, Dist(X) is low for an unmarked text by construction — the most salient features of an unmarked text are the generic ones, which the marked poles also have — so H1's entire registered signature (low Dist(C), low Dist(D), positive Asym) would have been produced with no register fact involved. Dist was not a test of H1; it was a restatement of which passage had been called unmarked. Finding 1 is also correct: an 8-item denominator with a strict inequality and no band turns one predicate flip into a headline.
A1 — the materials seat keeps no reserve, and one transport failure is on the record
The first enumeration run reached A, B and C and then qwen/qwen3.7-max at Alibaba returned 1,892 bytes of SSE keep-alive padding and no JSON payload at all on passage D. This is not note (b)'s failure mode (a model returning length or an empty completion); nothing was served. The seat was not replaced, because a heterogeneous enumerator would put enumerator identity inside the design's own statistic; the request was re-issued to the same slug and accepted first time. call.py gained a PayloadFree exception so the case is recorded rather than crashing, runs/enum-attempts.json logs every attempt, and the failed call shows no billing in the key-usage delta at the time of check (delta below the per-request sum, i.e. note (bco)'s settling lag, so this is evidence and not proof). New method note (beq).
A2 — the screen ran; two predicates dropped
The critic's SCREEN block dropped 2 of 56: predicate 1 (two properties in one, violating constraint (d)) and predicate 22 ("specialized technical jargon", which is defined by the passage's subject matter, violating (b) and (c)). Failure criterion 4 allowed up to 14; 2 fired, so the criterion does not fire. The drops are adopted verbatim.
A3 — the primary statistic is replaced (Findings 1 and 2)
Dist(X) and Asym(X) are withdrawn from the decision rule. Excl(X) is retained as descriptive only. The new primary statistic, and the whole of the decision:
Disc(X)— the discriminating power of X's catalogue. Over every (class, other-passage) pair with the class inown(X)and the other passage one of the six that is not X, the fraction ABSENT by consensus. Denominator|own(X)| × 6= 42 or 48. Range 0 — X's salient features hold of every other passage, so a catalogue of them discriminates nothing — to 1, X's features are unique to X.
What this makes the experiment about, and it is narrower than the frozen text above claims. Disc measures whether a catalogue read off a candidate text can discriminate texts at all. A low Disc(C) therefore supports "an anchor built on this material grounds nothing the two existing anchors do not already imply" — a claim about anchorability, which is what ARM-evidence-audit item 3 actually asks. It does not support "the unmarked centre is not a register of English", and this result may not be written that way. Finding 2 is the reason the claim is narrowed, and the narrowing is the critic's, not the lead's.
Granularity: steps of 1/48 = 0.021 rather than 1/8 = 0.125, which is Finding 1's fix arriving through the same change.
A4 — the item pool is deduplicated using the critic's own classes (Finding 6)
The critic named nine duplicate groups across the 56. They are adopted verbatim, so no lead judgment enters the pool, with one correction, declared: the critic's single tense class merged predicates 15 and 25 (past tense) with 17 and 41 (present tense), which are opposites rather than duplicates. It is split into tense-past and tense-present. Nothing else is altered.
Result (build_pool.py, materials/pool.json): 54 survivors → 38 classes. Canonical wording of a class is its lowest-numbered surviving member, used verbatim. own(X) is every class with at least one member enumerated from X: own(A) 7, own(B) 8, own(C) 7, own(D) 8, own(E) 8, own(F) 8, own(G) 8. Ten classes are owned by more than one passage and that is reported rather than resolved.
A5 — per-pair comparisons replace min() (Finding 3)
m = min(Dist(E), Dist(F)) is withdrawn: taking the minimum stacks the threshold at whichever control happens to be less distinctive. Two registered per-pair comparisons instead, each within an author, each with the control's marked direction declared before the census:
- C (McHugh, unmarked) against F (McHugh, first-person teenage vernacular)
- D (Kessel, unmarked) against E (Kessel, pseudo-documentary)
Finding 3's other half is accepted as a stated limitation and not repaired: F differs from C in grammatical person and tense as well as in markedness, and E differs from D likewise, so each pair holds author, book and cut rule fixed and does not hold person and tense fixed. Four texts cannot. This is recorded in the result as a confound, not as a control.
A6 — the decision rule, restated, with a band (Findings 1, 3, 5)
Let Δ_C = Disc(F) − Disc(C) and Δ_D = Disc(E) − Disc(D). Band 0.10 (≈5 of 48 pairs, ~5× the smallest step).
- H1 (the centre's catalogue does not discriminate) supported iff
Δ_C ≥ 0.10andΔ_D ≥ 0.10. - H2 (it discriminates as well as a marked sibling's) supported iff
Δ_C ≤ −0.10andΔ_D ≤ −0.10. - Undecided in every other case, including any
|Δ| < 0.10and any split between the two pairs. A tie is undecided, not H2 — Finding 1 noted that the frozen rule silently counted ties as H2 and it was right.
Finding 5 is answered by deletion, which the critic's own FIX offered as one of two options: Asym is removed from H1's registered prediction rather than added to the rule, because Asym was built out of the withdrawn Dist. Disc(A), Disc(B), Disc(G) and all Excl values are reported descriptively for every passage, which is Finding 2's requested base-rate reporting.
A7 — the repeat control is widened to every decision-relevant cell (Finding 4)
Failure criterion 2 is replaced. All four decision-relevant cells for one rater — (C, P1), (D, P1), (E, P1), (F, P1) — are re-sent byte-identical, same day, same parameters, and the decision is recomputed under each substitution separately and under all four together. If the decision changes under any substitution the headline is withheld. Per-class flip counts are reported descriptively.
Finding 4's other half cannot be repaired by this session and is owed: the instrument failures that motivated the control (S061, S062) were cross-day, and a session inside one UTC day cannot supply a second day. This is the same limitation NEXT.md records for the third-day C15/C16 repeat. The result states that the control run is same-day and that a cross-day repeat of these cells is outstanding.
A8 — revised pre-flight, worst case from max_tokens
Spent to this point: enumeration $0.14302 (7 accepted calls, one payload-free at no visible cost) + critic $0.1188702 = $0.2618902.
| stage | calls | max_tokens |
worst case |
|---|---|---|---|
| census P1, 7 passages × 38 classes | 7 | 6,000 | $0.346 |
| census P3, 7 passages × 38 classes | 7 | 6,000 | $0.301 |
| repeat control, P1, cells C/D/E/F | 4 | 6,000 | $0.198 |
| declared reserve P2 if a census seat fails | ≤3 | 6,000 | $0.154 |
| remaining worst case, with reserve | ≤21 | $0.999 |
Experiment worst case $1.261, against $4.0289866523 headroom at session start. The pool shrank from 56 items to 38, so the census prompts are smaller than the frozen estimate assumed; the per-call figures above are rebuilt from the actual prompt sizes and the caps actually sent, and both the frozen $1.431 and this $1.261 are on the record rather than only the one that flatters the estimate (note (abc)).