Repository path: workshop/experiments/E-20260810c-register-reach/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260810c-register-reach |
| status | frozen |
| created | 2026-08-10 |
| updated | 2026-08-10 |
| senses | style-correspondence |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-register-reach.md, workshop/regimes/R23-device-crossed-low.md, workshop/regimes/R22-placeless-low.md, workshop/translations/botchan/R22-v1/translation.md, workshop/experiments/E-20260809g-device-cross/design.md, wiki/findings/results/RS-20260809g-device-cross.md, wiki/findings/results/RS-20260808f-placeless.md, workshop/experiments/E-20260808e-low-pole/site-criterion.md, framework/v0.2/README.md, config/models.md, config/budget.md, wiki/method-notes.md |
E-20260810c — the powered cell: how much of English's way down is bought by the place, and how much by the spelling?
Frozen 2026-08-10 (S150) before dispatch. ARM-register-reach step 1. The translation limb
(T-botchan-R22-v1, its log, and its contamination measurement) was committed at 915bf89
before this file existed.
1. Question
framework/v0.2 §7.1 refuses a register-carriage recommendation and names, in writing, the one
thing a successor needs:
the register comparison needs more sites per language cell than fourteen across two […] A cell of ~20 marked narration sites in one pair, with sites chosen for device availability rather than only for source markedness, is what would make
P1areadable.
This is that cell.
At narration sites where a source drops below its own neutral written register, how much of English's downward movement is bought by a LOCATED IDIOM and how much by a NONSTANDARD SPELLING — asked where there are enough sites for either answer to be legible?
2. What is carried from E-20260809g and is not re-asked
P3is settled and is not re-tested. Nonstandard spelling is not placeless:loc(B)= 0.5714 / 0.5357 against a registered 0.25 on two independent judges, respelling placed more often than located idiom in all four judge × cell blocks. This run measuresloconly as a manipulation check (G2), not as a hypothesis.A21— where an arm's English at a site is byte-identical to that hand's∅,dis 0 by construction and the seats' codes are not differenced.A19's conclusion matrix is carried in §11 with the cell count changed and nothing else.- Note (bkr) — no content-parity check gates anything here. The
errandaddquestions are asked inside the reproducedlocprompt and are reported descriptively only. REGis a direction, not a magnitude, and is batch-sensitive by a factor of three (RS-20260806g). Every difference is read within one call: the rating is split by site, with all eight arms of a site kept together in one call, and absolute per-block means are never pooled across blocks.
3. Five registered changes from E-20260809g, each with its reason
| # | change | reason, from the record |
|---|---|---|
| C1 | One cell, Japanese, ~20 sites instead of two cells and fourteen | The IT cell died of ties (5 non-tied of 8); the JA cell died of n (5 non-tied of 5 available, against a bar of 6). Japanese narration sites are clause-shaped and their tie rate at S145 was zero, so n is the whole repair (framework/v0.2 §7.1, third measurement) |
| C2 | The site list is censused by three INDEPENDENT annotators and the lead is not the spine | E-20260808e's A4 spined the vote on List A. Here the lead has just translated this exact span under R22 and published a 21-row table of its low sites knowing the arm's question. A lead-spined list would be the hypothesis selecting its own evidence. The lead's table endorses nothing and is compared afterwards (P5) |
| C3 | H2 changes from mistralai/mistral-medium-3-5 to qwen/qwen3.7-max |
S145's mechanical reach count: H2-JA-I = 1 of 6. A hand that does not apply the manipulation measures nothing, and config/models.md records mistral as "weakest on Japanese texture" |
| C4 | No PUB arm |
A14 replaced G6 with G6′ (mean(d_AB) > 0), which needs no published comparator. Aligning Morri 1918 to freshly censused sites would put lead judgement inside the run for a figure nothing rests on |
| C5 | G5 raised from ≥ 6 to ≥ 12 non-tied sites |
The bar a larger n is bought for. At 12 the exact two-sided sign-flip test attains P = 0.00049; at 6 it attains 0.03125 and nothing else |
No bar carried from S145 is loosened. G2, G3, G6′ and every prediction threshold are the
S145 values verbatim.
4. Materials
| cell | JA → EN, and it is the only cell |
| source span | 夏目漱石「坊っちゃん」(1906) 一, paragraphs 2, 3, 4, 5, 6, 11, 13, 15, 16, 17, 18, 19 — 3,482 non-space characters, the whole of chapter 1 that is not the Kiyo thread. workshop/translations/botchan/source-ja-span3.txt |
| copy-text | Aozora Bunko 000148/752_14964, stored at wiki/base/anchors/A-morri-botchan/botchan-soseki-1906-ch1-ja.txt |
| site criterion | workshop/experiments/E-20260808e-low-pole/site-criterion.md, verbatim and unmodified, frozen 2026-08-08 before any published English of any census work was opened |
| type | N (narration) only, by that criterion's mechanical framing rule |
| translation limb | T-botchan-R22-v1, frozen at 915bf89. It is rated by nobody and enters no figure; its log supplies P5's registered comparison and nothing else |
Why this span is the right material and not a convenience. It is 3.66× the Japanese the S145 JA
cell was cut from, it is continuous, it is the part of chapter 1 with no deferential speech in it,
and it is clause-shaped throughout — Sōseki's 俺-narration is finite predicates end to end, which is
the syntactic property framework/v0.2 §7.1 identifies as what gives a respelling anything to
attach to.
5. Stages
Stage 0 — critic. One adversarial pre-run pass over this frozen file plus R23,
openai/gpt-5.6-terra, reasoning off (note (bkw)), cap 16,000. Findings accepted or overruled in
writing, before dispatch.
Stage 1 — the site census, three independent annotators. Each is given the source span, the
frozen criterion verbatim, and nothing else: no English, no mention of translation, no hypothesis,
no indication that register direction or any device is at issue. Each returns a JSON array of
stretches with src and type.
B1mistralai/mistral-medium-3-5·B2z-ai/glm-5.2·B3nvidia/nemotron-3-ultra-550b-a55b- The vote, frozen in
code.pybefore dispatch.B1is the spine by dispatch order and not by merit. AB1stretch is endorsed byB2orB3if that census lists a stretch overlapping it undervote_sites.py'soverlap()copied verbatim fromE-20260808e. A stretch is admitted at ≥ 2 of 3.typeis the majority of the votes cast on it; a tie takes the spine's. - Only
type: Nsites go forward. If more than 24 are admitted, the first 24 in source order are taken — registered here, before the census exists. If fewer than 14 are admitted,G5cannot be met and §10F7fires. |B2 \ B1|and|B3 \ B1|are reported as the coverage the spine missed.
Stage 2 — generation, R23's four cells, two hands. H1 x-ai/grok-4.5 · H2
qwen/qwen3.7-max. Each hand writes its own ∅ and then three minimal revisions of its own ∅,
one call per cell, using E-20260809g's BASE/CLAUSE/REVISE/TAIL prompt strings word for
word. A hand is given the source stretches, the rule text for its cell, and (for a revision) its
own ∅. It is not told the hypothesis, that other cells exist, or that any comparison will be
made. 8 calls, cap 6,000.
Stage 3 — loc. Two judges over every ⟨site, arm⟩ pair (8 arms). The prompt is
E-20260809g's LOC_HEAD word for word, including its two companion questions. L1
nvidia/nemotron-3-ultra-550b-a55b (S137's and S145's checker) · L2 z-ai/glm-5.2 (S145's).
Split by site into two blocks; cap 16,000.
Stage 4 — REG. The signed −3…+3 register-direction scale, reproduced word for word from
E-20260806b via E-20260806g, E-20260808e, E-20260808f and E-20260809g. Three seats: J1
openai/gpt-5.6-terra · J2 google/gemini-3.6-flash · J3 deepseek/deepseek-v4-pro — S145's
seats, same instrument, no new competence claim is made. Split by site into three blocks
with all eight arms of a site together; independent shuffle per seat under a fixed seed; cap 12,000.
Role double duty, declared. The panel and reserves reachable here are nine models and this
design needs eleven roles. B2/L2 and B3/L1 are the same models in two stages. A census
annotator saw the Japanese only and was told nothing about translation; a loc judge is asked
about English placedness. No generating hand, and no REG seat, is a census annotator.
openai/gpt-5.6-terra is both critic and J1, as at S145.
Not dispatched, deliberately: moonshotai/kimi-k3 — note (bkh-corr).
6. Quantities
For arm x and site s, REG(x,s) is the mean over the returned seats of the register-direction
code. Positive means the English is higher than its source. For a generating hand h:
d_I(h,s)=REG(h∅,s) − REG(hA,s)— what the located idiom buys.d_S(h,s)=REG(h∅,s) − REG(hB,s)— what the respelling buys.d_AB(h,s)=REG(h∅,s) − REG(hAB,s)— what both buy.
Positive d = the arm went lower than its own ∅. Site-level values pool the two hands.
7. Gates, registered before dispatch
| gate | what it protects | bar |
|---|---|---|
G1 manipulation S |
that +S cells respell and −S cells do not |
mechanical, frozen in code.py: each hand's B and AB contain ≥ 1 respelling; each hand's ∅ and A contain 0 |
G2 manipulation I |
that the located arm is actually located | loc(A) − loc(∅) ≥ +0.20, pooled over hands, on both judges. E-20260808f G8 bar, reproduced |
G3 scale one-sidedness |
that the scale is not pinned | ≥ 10% of all returned REG codes negative |
G5 power |
that a null is a null and not a tie structure | ≥ 12 non-tied sites on d_I − d_S |
G6′ positive control |
that the batch is measuring register direction at all | mean(d_AB) > 0 |
G7 returns |
that the analysis is not reading a truncation | ≥ 90% of expected rating cells returned after one re-dispatch |
8. Predictions, registered
P1a — which device, policy contrast (reach included). Over all admitted sites, hands pooled
within site: mean(d_I) − mean(d_S), tested by an exact two-sided sign-flip permutation
blocked at the site, α = 0.05.
Registered in both directions, and the direction the record points is stated so it cannot be claimed afterwards. S145's unread JA margin was −0.5833 (
mean d_I0.2083,mean d_S0.7917) — the spelling, in all four hand × cell blocks and under every leave-one-seat-out.P1a′(spelling buys more) is therefore the expected direction and is the weaker claim to make;P1a(idiom buys more) would overturn a consistent margin. Either is licensed only at |difference| ≥ 0.50 and P < 0.05 and the same sign in both hands.
P1b — conditional contrast. The same test restricted to sites where, for that hand, both
A ≠ ∅ and B ≠ ∅. P1a is about a policy; P1b is about the devices where both were applied,
and the two answer different questions (S145 A1+A2+A6).
P2 — the bundle. mean(d_AB) ≥ +0.50, pooled over hands and sites, de-linked from S137's
magnitudes.
P4 — reach, mechanical, no jury. For each hand and device, the share of admitted sites at
which that device's cell differs from that hand's own ∅ at all.
P5 — the translator's claim, tested against hands that never read it.
T-botchan-R22-v1's frozen log states that at 19 of its 21 low sites (0.905) the readiest
English was locatable and had to be refused. Registered: reach(I) pooled over the two hands
≥ 0.75. If the independent hands find a located option far less often, the log's claim is a fact
about the lead's English and not about English. This is the wire between the two limbs and it costs
nothing to run.
Also reported, descriptively and with no test: the share of independently censused N sites that fall inside the lead's 21 rows, in both directions.
9. Failure criteria
| # | condition | consequence |
|---|---|---|
F1 |
G3 fails |
every REG-based figure withheld (P1a, P1b, P2) |
F2 |
G2 fails on either judge |
P1a and P1b withheld — an unlocated A prices nothing |
F3 |
G1 fails for a hand |
that hand's affected contrast is void; if both hands fail for a device, that device is unmeasured |
F4 |
G5 fails |
P1a is reported UNDERPOWERED and is not read as a null; no verbal claim that a device "does not work" |
F5 |
G6′ fails |
the cell contributes to no REG figure at all |
F6 |
G7 fails |
affected items dropped, the loss reported, and any figure resting on < 14 sites withheld |
F7 |
fewer than 14 type: N sites admitted by the census |
G5 is unreachable; P1a and P1b are not dispatched at all and the census result is the session's finding |
No bar in this table moves after it fires. If a gate fires, the margin is reported and not exploited.
10. Limits, declared in advance
- Site-level generation and site-level rating. Arms render and are rated on short stretches out
of context. A device's effect over a whole paragraph is not measured, and
T-botchan-R22-v1's own log argues the low register of this narration lives in its speed, which is a paragraph property. This design cannot see that, and a null here does not reach it. - One work, one author, one pair. Sōseki's 俺-narration is a particular kind of low, and nothing here generalises past Japanese narration of this shape.
- Two hands. Both are LLMs.
- No human rater exists in this project. The standing BLOCKING from
E-20260808e's critic is unrepaired and is a property of the project, not of this run. - The seats are three LLMs and the
locjudges are two more, two of which also censused the Japanese (§5, role double duty). T-botchan-R22-v1is the lead's and the lead knew the arm's question while writing it. It is rated by nobody here.P5uses only the counts in its frozen log, not its English.- Tier D is NOT PASSED. Nothing here is a quality judgment; the seats code a source–target register relation and no goodness sense is scored.
11. Conclusion matrix — binding on the result page
Read down the first row that applies. Nothing below the line a gate draws may be claimed.
| condition | what may be claimed |
|---|---|
F7 fires |
the census result only; no device is compared |
G3 fails |
nothing on the REG scale. P4, P5 and the loc figures stand; they need no scale |
G6′ fails |
the cell contributes to no REG figure — not P1a, not P1b, not P2 |
G2 fails on either judge |
P1a and P1b withheld. P2, P4, P5 and every descriptive figure stand |
G1 fails for a hand |
that hand's affected contrast is void; if both hands fail for a device, that device is unmeasured and no comparison involving it may be stated |
G5 fails |
P1a is UNDERPOWERED: its mean is printed, no null is claimed, and no verbal claim that a device "does not work" may be made from it |
| the sign differs by hand, or flips when a seat is dropped | P1a is descriptive only; the words supported, shows and establishes are forbidden of it |
G1, G2, G3, G5, G6′ all pass and the sign holds by hand and under leave-one-seat-out and P < 0.05 |
P1a may be stated as a finding, in the form "at marked narration sites in one Japanese work, permitting a located idiom buys N scale points of register direction and permitting a respelling buys M" — and in no other form |
| any outcome | no additive model of d_AB may be fitted or asserted; the devices are not orthogonal in reach |
| any outcome | T-botchan-R22-v1 enters no claim, no gate and no pooled figure, and the word replication is not used of it |
| any outcome | no claim that spelling is placeless. That is RS-20260809g P3, already FAILED, and is not re-opened |
| any outcome | no framework recommendation may be written if G5 failed |
12. Pre-flight cost
Built from the caps the requests permit, not from expected output (note (abc)).
| stage | calls | cap each | worst case |
|---|---|---|---|
| 0 critic | 1 | 16,000 | $0.40 (P1 at $6/M out, ×4 routing caution) |
| 1 census | 3 | 8,000 | $0.20 |
| 2 generation | 8 | 6,000 | $0.35 |
3 loc |
4 | 16,000 | $0.40 |
4 REG |
9 | 12,000 | $0.75 |
| input tokens, all stages | — | — | $0.15 |
| re-dispatch contingency | — | — | $0.15 |
Declared ceiling: $2.40. UTC day 2026-08-10 stands at $0.071580820991 of $5.00 before this
session, so the ceiling fits the day's headroom of $4.928419 with $2.53 to spare. Actuals recorded
from usage.cost per response and reconciled against the key-usage delta.
13. Amendments from the pre-run critic (2026-08-10, before dispatch)
critic.md carries the twenty findings, the adjudication, and the six overruled remedies in full.
openai/gpt-5.6-terra, NEEDS-REDESIGN, 20 findings — 4 BLOCKING, 11 SERIOUS, 5 MINOR — of
which 17 accepted, 3 accepted-in-part with the overruled half written out, and 6 individual
remedies overruled with the reason. Amendments A1–A20 are in force and supersede §§4–11
above wherever they conflict. The originals are left standing rather than rewritten, so that what
was frozen and what was changed are both visible.
The three that changed the experiment:
A1(BLOCKING 1). The estimand is a permission-policy effect, not a device effect.Aopens a broad lexical/idiomatic/grammatical class andBopens one narrow orthographic class; they are not matched doses. Every sentence reads permitting X, never device X buys.A3(BLOCKING 3). The word exact is struck from the test. Nothing is randomized; the sign-flip reference distribution rests on an assumed within-site exchangeability of the two permission labels under the sharp null, that assumption is printed wherever the p-value is, and inference is conditional on these sites, these two hands and these three seats.A16(MINOR 16). The census vote is symmetric: candidates are the union of all three annotators, clustered under the verbatimoverlap()rule, admitted at ≥ 2 of 3 distinct annotators. There is no spine, so a stretch two annotators found and the third missed is admitted.
Also in force, in brief. A2 — P1b demoted to exploratory and barred from every claim, being
post-treatment conditioning. A4 — the licensed sentence replaced with procedure-level wording.
A5 — cap 24 → 30, and above it a paragraph-stratified fixed-seed sample replaces first-n.
A6 — non-tie defined in code at the analysis level, G5 evaluated once. A7 — the whole analysis
algorithm frozen in analyse.py before dispatch. A8 — G6′ renamed a minimal internal
manipulation check. A9 — G3 strengthened: ≥ 10% negative overall and ≥ 10% among ∅/A
items alone. A10 — every REG figure is an LLM-panel-perceived relation. A11 —
naturalness struck from senses:, nothing here collects it. A12 — site-level compliance
coded; P1a declared intention-to-treat. A13 — G1 split into per-site purity and aggregate
application. A14 — content parity reported per arm plus an add-excluded sensitivity, and no
parity check gates anything (note (bkr), overruled in writing). A15 — the sample is labelled an
enriched feasibility sample. A17 — the outcome is a comparative panel judgment within a call.
A18 — P5 loses its threshold and is a coarse cross-procedure comparison. A19 — the request map
fixed: every seat rates every block. A20 — the word powered struck; this is an enlarged
measurement run.
14. Two amendments registered during the run, before any English existed
A21.A16's union-and-cluster rule, implemented as single-link clustering, chained 312 stretches into 26 blobs. The unit of agreement becomes the sentence — character-level coverage per annotator, 111 sentences, admitted at ≥ 2 of 3. Symmetric, order-independent, cannot chain, and it is the coarsest unit the criterion's own exclusion 3 permits. Registered with no English and no rating in existence.A22. The census admits 97 of 111 sentences (0.874). The criterion was built to find islands of low register and this narration is low throughout. The phrase "marked narration sites" is struck: the sample is 30 of the 82 criterion-positive narration sentences of this span, drawn paragraph-stratified at seed 5501, and every figure says so.
Both are stated in full in critic.md.