Repository path: workshop/experiments/E-20260809g-device-cross/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260809g-device-cross |
| status | frozen |
| created | 2026-08-09 |
| updated | 2026-08-09 |
| senses | style-correspondence, naturalness |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-register-devices.md, workshop/regimes/R23-device-crossed-low.md, workshop/translations/malavoglia-i/R23-v1/translation.md, wiki/findings/results/RS-20260808f-placeless.md, wiki/findings/results/RS-20260808e-low-pole.md, framework/v0.2/README.md, config/models.md, config/budget.md, wiki/method-notes.md |
E-20260809g — which device buys the low pole in English narration: the place, or the spelling?
Frozen 2026-08-09 (S145) before dispatch. ARM-register-devices step 1. The translation limb
(T-malavoglia-i-R23-v1, its log, and its dependence measurement) was committed at 18ab5ff
before this file existed.
1. Question
framework/v0.2 §7 states a problem and refuses a recommendation for one named reason: the
constraint that produced RS-20260808f's effect bundles regional/class idiom with
eye-dialect, and that run could not say which half the register was paid to. §7 then names the
design that would settle it. This is that design.
At narration sites where a source drops below its own neutral written register, does English reach the low pole by a LOCATED IDIOM, by a NONSTANDARD SPELLING, or only by both — and is the spelling really the placeless device the framework assumes it is?
2. Materials — nothing is re-cut
Two cells, two language pairs, the site lists frozen at S136 and not touched here. Sites were
voted by three annotators against E-20260808e/site-criterion.md, which was written before any
published English was opened; a site is admitted at ≥ 2 of 3 votes. Only type N (narration)
sites are used, by the mechanical typing rule A5 of that run.
| cell | source span | admitted N sites | published hand |
|---|---|---|---|
| IT | Verga, «I Malavoglia» (1881), Ch. I opening, 595 tokens | 8 — IT-03, 05, 06, 08, 09, 10, 11, 12 | Mary A. Craig, 1890 |
| JA | Sōseki, 「坊っちゃん」 (1906), Ch. 1, paragraphs 3, 5, 6 — 952 characters | 6 — JA-05, 06, 08, 09, 10, 12 | Yasotarō Morri, 1918 |
14 sites. The IT sites are the same eight on which RS-20260808f measured the bundle at
Q1 = +0.7917 and Q2 = +0.8000, which is why they are reused: this run decomposes a specific
measured quantity at the sites it was measured on.
3. Arms
R23's four cells, ∅ (−I −S) · A (+I −S) · B (−I +S) · AB (+I +S).
Primary arms — two independent generating hands, blind. Each hand writes its own ∅ and then
three minimal revisions of its own ∅, one call per cell, in the R23 §Procedure shape. A hand
is given the source stretches, the rule text for the cell it is writing, and (for a revision) its
own ∅. It is not told the hypothesis, that other cells exist, what is being measured, or that
any comparison will be made.
| role | slug | why |
|---|---|---|
| H1 | x-ai/grok-4.5 |
wrote LOW-B at S136; is not a rating seat here |
| H2 | mistralai/mistral-medium-3-5 |
wrote LOW-A (S136) and LOW-P (S137); non-panel; is not a rating seat |
Secondary arm — the lead's quadruple, T-malavoglia-i-R23-v1, IT only. Rated in the same call
as the model arms so it is on the same scale, and read as a human replication and never as a
primary (that consequence is fixed on the artifact, before any measurement). Its ∅ is
T-malavoglia-i-R22-v1, written at S137 before this hypothesis existed; A, B and AB were
written this session by a party who knows the hypothesis.
Reference arm — PUB, the published hand's English at each site, from the frozen site list. It
generates nothing and exists as the positive control of §6 G6.
Arms rated: IT 13 (8 model + 4 lead + PUB) × 8 sites = 104 items · JA 9 (8 model + PUB)
× 6 sites = 54 items.
4. Procedure
Site-level generation and site-level rating, as E-20260808f. Every raw body is written to runs/
before anything is computed from it; a dead body is rotated into runs/discarded/ and never
overwritten (note (bhd)).
- Stage 0 — critic. One adversarial pre-run pass over this frozen file,
openai/gpt-5.6-terra, reasoning off (note (bkw)). Findings accepted or overruled in writing, before dispatch. - Stage 1 —
∅. Two calls (H1×2 cells… four calls: H1-IT, H1-JA, H2-IT, H2-JA). - Stage 2 — the three revisions. Twelve calls: each hand × each cell × {
A,B,AB}, each given that hand's own∅for the same sites. - Stage 3 —
loc. Two independent judges, neither a generating hand nor a rating seat, over every ⟨site, arm⟩ pair. The question is reproduced word for word fromE-20260808f's stage A, including its two companion questions, so thatlocfigures are comparable with that run's. LOC-1nvidia/nemotron-3-ultra-550b-a55b(S137's checker). LOC-2z-ai/glm-5.2, dispatched second and smallest-call-first as a live slug probe; if it returns a zero-contentfinish_reason: length, the registered fallback is single-judgelocwith S137's own single-judge limit carried, and no bar moves. - Stage 4 —
REG. The signed −3…+3 register-direction scale, reproduced word for word fromE-20260806bviaE-20260806g,E-20260808eandE-20260808f. Three seats, one call per cell per seat, arms interleaved and shuffled under a fixed seed. Seats J1openai/gpt-5.6-terra, J2google/gemini-3.6-flash, J3deepseek/deepseek-v4-pro— S137's seats, same instrument, same languages, no new competence claim is made.
Registered dispatch fallback. If an IT rating body returns finish_reason: length with no
usable JSON, the call is split by site into sites 1–4 and 5–8, all thirteen arms kept together
within each half, so that every within-site contrast this design reads remains inside one call.
Absolute per-cell means are then reported per half and never pooled across the split.
Not dispatched, deliberately: moonshotai/kimi-k3 — note (bkh-corr), two zero-content
finish_reason: length bodies in this project at $0.125778 and $0.179433.
5. Quantities
For arm x and site s, REG(x,s) is the mean over the three seats of the register-direction
code. Positive means the English is higher than its source; going low is negative. For a
generating hand h:
d_I(h,s)=REG(h∅,s) − REG(hA,s)— what the located idiom buys.d_S(h,s)=REG(h∅,s) − REG(hB,s)— what the respelling buys.d_AB(h,s)=REG(h∅,s) − REG(hAB,s)— what both buy.
Positive d = the arm went lower than its own ∅.
6. Gates, registered before dispatch
| gate | what it protects | bar |
|---|---|---|
G1 manipulation S |
that +S cells respell and −S cells do not |
mechanical: each hand's B and AB contain ≥ 1 respelling; each hand's ∅ and A contain 0 |
G2 manipulation I |
that the located arm is actually located | loc(A) − loc(∅) ≥ +0.20, pooled over the model hands, on LOC-1. E-20260808f's G8 bar, reproduced |
G3 scale one-sidedness |
that the scale is not pinned | ≥ 10% of all returned REG codes negative. E-20260808e's G1 bar, reproduced |
G4 minimal-pair integrity |
that a revision is a revision | per hand, ≥ 70% token identity between each revision cell and its own ∅, at ≥ 10 of 14 sites |
G5 power |
that a null is a null and not a tie structure | ≥ 10 of the 28 ⟨hand, site⟩ pairs non-tied on d_I − d_S |
G6 positive control |
that the batch is measuring register direction at all | REG(PUB) − REG(model ∅) > 0 in each cell |
G5 is the registered repair of RS-20260808f §3.1.1, where four of eight sites were exact
ties and P = 0.125 was the smallest value the test could return — the design could not have reached
significance whatever the data did, and nobody knew until afterwards. Here the question is asked
before dispatch and its answer withholds an interpretation rather than a number.
What is NOT a gate, by rule
Content parity gates nothing in this design. Note (bkr): a content-parity check cannot
license a register claim about register-marked sites, because at those sites nothing passes it —
RS-20260808f measured the published hand as omitting or misstating at 0.6809 and supplying at
0.7660, the highest supply rate in that run. The err and add questions are asked (they are part
of the reproduced loc prompt) and are reported descriptively. They withhold nothing, and no
figure in this design is conditioned on them.
7. Predictions, registered
P1 — which device. Pooled over 2 hands × 14 sites, mean(d_I) − mean(d_S) ≥ +0.50, with
the same sign in both hands. Exact two-sided paired permutation over the 28 within-⟨hand, site⟩
differences.
P1is registered in both directions and neither reading is post-hoc.T-malavoglia-i-R22-v1's log says "the live ones are the located ones" and predictsd_I > d_S.RS-20260808f§3.1.2's post-hoc decomposition points the other way: the two largest movers in that whole run were its only two eye-dialect sites, and removing them took the effect from +0.7917 to +0.2778. Ifmean(d_S) − mean(d_I) ≥ +0.50on the same terms,P1′holds and the framework's recommendation is the opposite one. The record contains two readings pointing in opposite directions; both are written down here before the data exist.
P2 — the bundle reproduces. mean(d_AB) ≥ +0.50 pooled over hands and sites.
RS-20260808f measured +0.7917 and +0.8000 for the same bundle at the IT sites, on a different
instrument and with four exact ties; a value near zero here means the S137 figures do not survive a
minimal-pair design and that is a result about them.
P3 — is the spelling placeless? loc(B) ≤ 0.25, pooled over model hands and sites, on
LOC-1. This is the bar E-20260808f's G6 set for LOW-P (which met it at 0.0588). P3 is
framework/v0.2 §7's own assumption stated as a testable proposition — that a translator "can
adopt phonetic spelling without relocating a book." The lead's expectation, recorded so that it
cannot be claimed afterwards, is that P3 FAILS.
P4 — reach, mechanical, no jury. For each hand and device, the share of the 14 sites at which
that device's cell differs from the hand's own ∅ at all. Registered: reach(S) < reach(I) and
reach(S) ≤ 0.50. The lead's own quadruple gives reach(I) 8/8 and reach(S) 3/8 at the IT
sites, which is where the prediction comes from; the model arms are where a hand that has not read
that log gets asked.
8. Failure criteria
| # | condition | consequence |
|---|---|---|
F1 |
G3 fails |
every REG-based primary withheld (P1, P1′, P2) |
F2 |
G2 fails |
P1 and P1′ withheld — an unlocated A prices nothing |
F3 |
G1 fails for a hand |
that hand's affected contrast is void; if both hands fail, the device is unmeasured |
F4 |
G5 fails |
P1 is reported UNDERPOWERED and is not read as a null |
F5 |
G6 fails in a cell |
that cell contributes to no REG primary |
F6 |
< 90% of expected rating cells returned after one re-dispatch | affected items dropped, the loss reported, and any primary resting on < 20 pairs is withheld |
No bar in this table moves after it fires. If a gate fires, the margin is reported and not exploited.
9. Limits, declared in advance
- Site-level generation and site-level rating. Arms render and are rated on short stretches out of context, as at S136 and S137. A device's effect on a whole paragraph is not measured.
REGis a direction, not a magnitude, and is batch-sensitive by a factor of three (RS-20260806g). Every figure is comparative within its own call; no figure here is comparable with any S136 or S137 figure, including theQ1/Q2values this run decomposes.- Two hands, two pairs, fourteen sites. Small.
- The seats are three LLMs and the
locjudges are two more. No human rater exists in this project; the standing BLOCKING fromE-20260808e's critic is unrepaired and is a property of the project, not of this run. - The lead's quadruple knows the hypothesis at
A,BandAB, and its∅carriesR22's declared Craig exposure. It is a replication and never a primary. +Swas applied by the lead under a stated policy that declined h-dropping and thet'article (T-malavoglia-i-R23-v1§3). The model hands are under no such policy and may use them; if they do,loc(B)is measuring a differentBfor the models than for the lead, and the two are reported separately.- Tier D is NOT PASSED. Nothing here is a quality judgment; the seats code a source–target register relation and no goodness sense is scored.
10. Pre-flight cost
Built from the caps the requests permit, not from expected output (note (abc)).
| stage | calls | cap each | worst case |
|---|---|---|---|
| 0 critic | 1 | 16,000 | $0.40 (P1 at $6/M out, ×4 routing caution) |
| 1–2 generation | 16 | 3,000 | $0.20 |
3 loc |
4 | 16,000 | $0.30 |
4 REG |
6 | 16,000 / 12,000 | $0.45 |
| input tokens, all stages | — | — | $0.10 |
| re-dispatch contingency | — | — | $0.35 |
Declared ceiling: $1.80. UTC day 2026-08-09 stands at $1.965544 of $5.00 before this session, so
the ceiling fits the day's headroom of $3.034456 with $1.23 to spare. Actuals recorded from
usage.cost per response and reconciled against the key-usage delta.
11. Amendments from the pre-run critic (2026-08-09, before dispatch)
critic.md carries the findings, the adjudication and the two overrulings in full.
openai/gpt-5.6-terra, NEEDS-REDESIGN, 24 findings — 10 BLOCKING, 13 SERIOUS, 1 MINOR — of
which 22 accepted and 2 overruled in writing. Twenty amendments A1–A20 are in force and
supersede §§5–8 above wherever they conflict. The originals are left standing rather than
rewritten, so that what was frozen and what was changed are both visible.
The one that changed the experiment: P1 is split into P1a (the policy contrast, reach
included) and P1b (the conditional contrast among applied sites), because d_I and d_S were
never measured on the same units — each device revises only where its own opportunities are, and
the design was about to read a reach difference as a strength difference.
Also in force, in brief: generator blindness withdrawn (A2); every compliance judgement mechanical
and frozen in code.py before dispatch, with the lead adjudicating nothing (A3); both loc
judges required and the weakening fallback deleted (A4); G4 replaced by a site-level purity test
(A5); a content check on the A/AB cells that withholds nothing at run level (A6); the
estimand narrowed to within a domesticating regime (A7); permutation blocked at the site,
run within each language cell, pooling demoted to a labelled secondary (A11); G5 re-derived
from the exact attainable P and set at ≥ 6 non-tied sites per cell (A12); G6 replaced by
G6′, mean(d_AB) > 0 (A14); P2 renamed and de-linked from the S137 magnitudes (A17); the
lead's quadruple quarantined out of every gate, test and pooled figure (A18).
12. Conclusion matrix (A19) — binding on the result page
Read down the first column that applies. Nothing below the line a gate draws may be claimed.
| condition | what may be claimed |
|---|---|
G3 (one-sidedness) fails |
Nothing on the REG scale. P4 (reach, mechanical) and the loc figures stand; they need no scale |
G6′ fails in a cell |
that cell contributes to no REG figure at all — not P1a, not P1b, not P2 |
G2 fails on either loc judge |
P1a and P1b withheld. P2, P3, P4 and every descriptive figure stand |
G1 fails for a hand |
that hand's affected contrast is void; if both hands fail for a device, that device is unmeasured and no comparison involving it may be stated |
G5 fails in a cell |
that cell's P1a is UNDERPOWERED: its mean is printed, no null is claimed, and no verbal claim that a device "does not work" may be made from it |
A13's robustness fails (sign differs by hand, or flips when a seat is dropped) |
P1a is descriptive only; the words supported, shows and establishes are forbidden of it |
all of G1, G2, G3, G5, G6′ pass and A13 holds and the direction is compatible in both cells |
P1a may be stated as a finding, in the form "across two pairs, permitting device X buys N scale points of register direction across a book's marked narration sites and permitting device Y buys M" — and in no other form |
| any outcome | no additive model of d_AB may be fitted or asserted; the devices are not orthogonal in reach (critic.md §What the critic did not catch) |
| any outcome | the lead's quadruple enters no claim, no gate and no pooled figure, and the word replication is not used of it |
| any outcome | no claim that spelling is placeless unless P3 is met on both loc judges |
| any outcome | no framework recommendation may be written from a cell whose G5 failed |
13. A21 — registered after generation, before any rating
Where an arm's English at a site is byte-identical to that hand's ∅, d is 0 by
construction and the seats' codes are not differenced. Identical English cannot differ in register
from itself. The rule was registered before the first loc or REG call and before any rating data
existed; it can only remove apparent differences, and it makes G5 harder rather than easier. Rate
of seat disagreement on byte-identical items is reported descriptively as a property of the
instrument. Full statement and the mechanical compliance results in critic.md §A21.