Repository path: workshop/experiments/E-20260815-register-room/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260815-register-room |
| status | frozen |
| created | 2026-08-15 |
| updated | 2026-08-15 |
| senses | style-correspondence, voice, naturalness |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-elevation-resolution.md, wiki/findings/results/RS-20260814d-elevation-resolution.md, workshop/regimes/R26-light-ennoblement.md, workshop/regimes/R25-ennoblement.md, workshop/regimes/R06-lead-single-pass.md, workshop/translations/zloumyshlennik/R06-v1/translation.md, workshop/translations/zloumyshlennik/R26-v1/translation.md, workshop/translations/zloumyshlennik/R25-v1/translation.md, config/models.md, config/budget.md, framework/v0.2/README.md |
E-20260815-register-room — does a one-step register policy act at a source's low places when they have room?
ARM-elevation-resolution step 2 (T4). Frozen before any API call of this run. The three
translations this design measures were frozen at commit 3fa489e, before this file existed.
1. The question, and why it is not settled
RS-20260814d (S183) measured R26 (light ennoblement) against R06 (no register rule) on
Andersen's «Flipperne» and found them +0.0833 apart at the sites where the Danish drops below its
own level and +0.6667 apart where it sits at it. It deposited this reading:
A one-step-up register policy is not experienced as raising the source. It is experienced as not letting the translation drop below it — and the place it does that work is the ordinary run of the prose, not the marked places.
Its own §6 offers a rival explanation of the same numbers, and the arm cannot close honestly without separating them:
- (A) The policy does not raise. A one-step rule has no upward effect at low places in any
source; what it does is prevent a fall elsewhere. If this is right, the sentence above goes into
framework/v0.2§8 Q-e as written. - (B) There was no room. Six of seven low sites in «Flipperne» are shorter than four Danish words — «Las!», «Snærpe!», «Forlovet!» — and "a policy that raises by one step has nowhere to go inside a one-word shout." If this is right, the sentence is conditional on span length, and the framework must say so, because a translator working on sustained low-register speech would otherwise draw the wrong conclusion from it.
The question: on a source whose low register comes in long turns as well as in one-word items, does a one-step-up register policy show at the low places?
Chekhov's «Злоумышленник» (1885) is the material that separates (A) from (B) inside one text, which is why it is worth a session: the peasant Denis Grigoryev's 26 turns run from 1 word to 79 (pool: 1, 1, 2, 2, 6, 7, 7, 9, 10, 10, 13, 14, 19, 21, 21, 25, 26, 30, 30, 35, 38, 41, 46, 55, 60, 79). The «Flipperne» condition and its opposite are present in the same story, the same author, the same translator and the same three regimes, so span length is tested within the design rather than across two runs.
Subject rule check (wiki/tracks.md, continue-prompt.md §4.5), in one sentence: this measures
whether a moderate register policy does any visible work where a source actually drops, and so tells
a translator whether such a policy is worth adopting on colloquial dialogue or only on ordinary
narrative prose. That is a claim about translating, not about the project's apparatus.
2. Materials — frozen, and all of them lead-produced at $0
| arm | regime | words | ratio to source | role |
|---|---|---|---|---|
R06 |
lead single pass, no register rule | 1,396 | 1.300 | floor |
R26 |
light ennoblement, N1–N7 | 1,429 | 1.331 | the arm under test |
R25 |
ennoblement, E1–E8 | 1,886 | 1.756 | calibration control |
Source: 1,074 Russian words, 57 paragraphs, workshop/translations/zloumyshlennik/source.txt
(stored at S165, unchanged). All three arms preserve the source's 57 paragraphs, asserted by
build_pool.py and by the verifier, so span alignment is 1:1 and mechanical.
2.1 Contamination, measured before the design was written (dependence.json)
| pair | 7-grams | 12-grams | 15-grams | longest run |
|---|---|---|---|---|
R06 ~ R26 |
644 | 475 | 407 | 115 |
R06 ~ GARNETT (1921) |
130 | 37 | 10 | 19 |
R26 ~ GARNETT |
88 | 17 | 3 | 17 |
R25 ~ GARNETT |
21 | 2 | 0 | 13 |
R06 ~ R25 |
18 | 0 | 0 | 10 — clean |
R25 ~ R26 |
45 | 2 | 0 | 12 |
(reference) GARNETT ~ PANEL_P1 |
89 | 13 | 3 | 16 |
(reference) PANEL_P1 ~ PANEL_P3 |
190 | 68 | 42 | 30 |
Garnett's 1921 "A Malefactor" and the two 2026 panel renderings are stored in this repository from
E-20260812i (S165). None of the three was opened by the translator; the table was produced by
tools/dependence_check.py, which reads files without displaying them, after all three arms were
frozen and committed.
Three things this table settles before the run, and one it does not:
R06~R26is the most dependent pair the project has measured — 475 shared twelve-grams and a 115-token run, far above note (bhb)'s previous densest firing (27 tokens,RS-20260808g; 113 twelve-grams,RS-20260814d). This is by construction:R26's N1 leaves the neutral alone and 30 of 57 spans are character-identical. ~~It is conservative for the primary — two arms that share this much are harder to tell apart, not easier~~ STRUCK by amendment A4 (critic F-4). That defence is a property of pairwise preference and detection tasks, and this design never asks a seat to compare two arms: every English span is coded against the Russian alone under labels scrambled per (site, seat), so the overlap does not enter the statistic at all — except at character-identical spans, where it is biasing rather than conservative, and where amendment A1 now excludes it from the LOW strata. The overlap remains fatal to any use of these two arms as independent renderings, which this design does not make.RS-20260814d§3's finding reproduces on a second language and a second author. The high rule set walks out of the plain register entirely:R06~R25is clean at 0 twelve-grams between two renderings by the same translator, on the same day, from the same copy-text, whileR06~GARNETTsits at 37. Overlap is a property of the register a policy targets, not of having a policy. Second measurement, second pair, same direction.R06is dependence-flagged against Garnett (37 twelve-grams, a 19-token run). Under the standing rule (CLAUDE.md, contamination) the lead may not serve as the independent third translator where a design's validity turns on independence from a published rendering. This design's validity does not turn on it: no published hand is an arm, and every comparison is between two register policies of the same translator. Stated here so the exclusion is visible rather than assumed.- What it does not settle: whether the lead's
R06is itself pulled toward Garnett's register by that overlap. If it is,R06is not a neutral floor. Nothing here measures that, and §9 carries it as a limit.
3. Seats
Per config/models.md. P5 (deepseek/deepseek-v4-pro) is not used on any task shape, note
(bne).
- Coding seats:
P1openai/gpt-5.6-terra,P2google/gemini-3.6-flash,P3x-ai/grok-4.5. - Pre-run adversarial critic:
P4moonshotai/kimi-k3— a seat that codes nothing in this run.RS-20260814h(S187) had to report its primary twice because the critic seat was also a judging seat; that is not repeated here.
4. Procedure
Stage 0 — token-cap probe
tools/panel_probe.py-style probe per note (bnl): one throwaway call per coding seat on a
two-span version of the stage-2 task, to read actual reasoning-token use, with caps set at the
observed maximum + 400. Probe calls are ledgered.
Stage 1 — the site list, from the Russian alone
Each coding seat receives the whole Russian story and the 57 mechanically built spans (the source's paragraphs, in order, no selection whatever) and codes each span:
below— below the ordinary written register of this storyat— at itabove— above it
No English appears anywhere in the stage-1 prompt. The seat also answers three frozen comprehension questions about the Russian, which gate its retention:
- What object is the accused charged with having removed? (a nut / bolt from the railway)
- What does he say he and his fellow villagers make out of them? (fishing sinkers / weights)
- What does the official tell him at the end will happen to him? (he is to be taken into custody / sent to prison)
G0 — source-comprehension gate. A seat answering fewer than 2 of 3 correctly is dropped from
the run entirely, and its stage-1 codes are discarded. This is RS-20260811g's lesson (a seat that
could not read the source still produced confident codes) applied before the money is spent on
stage 2.
The site list is the three-seat majority, exactly as RS-20260808e's and RS-20260814d's.
A span with no majority is excluded.
Stage 2 — register height, blind, one site at a time
Sites are selected mechanically from stage 1's majority classification crossed with the mechanical word count of the source span:
| stratum | rule | n |
|---|---|---|
| LONG-LOW | majority below and source span ≥ 25 words and R06 ≠ R26 |
4 |
| SHORT-LOW | majority below and source span ≤ 7 words and R06 ≠ R26 |
4 |
| NEUTRAL | majority at |
4 |
Within each stratum, spans are ordered by the SHA-256 of the source span text (already in
pool.json) and the first four are taken. No span is chosen by the lead, and the ordering was
fixed before stage 1 was run.
The identical-text exclusion on the LOW strata is amendment A1 (critic F-1, BLOCKING). Where
the two arms are the same text they cannot differ, and such a site contributes a structural zero
that attenuates Q1 and makes Q2 true by construction. identical_R06_R26 was computed by
build_pool.py before stage 1 existed. Its effect is small and is itself reported: of Denis's 26
turns exactly two are identical (spans 10 and 34), so the exclusion removes at most one
SHORT-LOW candidate and none at all from LONG-LOW, leaving 11 long and 6 short candidates
against a G2 floor of 3. R26's rule fired at every one of Denis's eleven long turns and at
six of his seven short ones — the translator-side fact a reader needs in order to read a null.
The cutoffs 25 and 7 are author-chosen, and amendment A5 (critic F-5) requires that to be said
here rather than implied away. Their basis: Denis's 26 turns have a natural gap in the length
distribution between 21 and 25 words (…13, 14, 19, 21, 21, 25, 26, 30…), and ≤ 7 is the band
that reaches the «Flipperne» condition, whose low sites were under four words. That basis was
chosen with knowledge of the materials by the person who wrote the arms, which is why A5 also
adds a cutoff-free statistic (Q1b) and a required per-site table.
Each body is one site × one seat. It contains the Russian span, and the three English renderings of that span under labels scrambled independently for every (site, seat) pair. The seat codes each English span against the Russian span:
−1the English sits below the Russian's register0at it+1above it
12 sites × 3 seats = 36 bodies. Judgment is not parallelised across sites within a seat.
Blinding. No prompt in either stage contains: the regime ids R06/R25/R26/R08/R21, the
words ennoblement, ennoble, register policy, regime, Berman, Chekhov, Garnett,
translator's log, lit-trans, experiment, hypothesis, the arm word counts, or any statement
that the arms differ in policy. Asserted by the verifier over the stored prompts, matched on word
boundaries. The seats can still recognise the story from the Russian; that is unavoidable in a
source-relative design and is a limit (§9), not a leak — recognition does not reveal which label is
which arm.
Runner discipline, note (bnx). Every body is appended to runs/bodies.jsonl as it returns;
re-invocation resumes from what is already on disk; the dispatch runs in the background with no
foreground wall clock. S187 lost $0.312883 to exactly this and the fix is three lines.
5. Gates — every one of them registered here, before dispatch
| gate | rule | what fails if it fires |
|---|---|---|
G0 |
each seat answers ≥ 2 of 3 source-comprehension questions | that seat is dropped |
G1 |
R25 − R06 at LOW sites (LONG-LOW ∪ SHORT-LOW) ≥ +0.50 |
Q1, Q2 withheld |
G2 |
each stratum has ≥ 3 qualifying spans after the majority rule | that stratum's primary is withheld |
G3 |
a seat is retained only if it returns usable codes at ≥ 2/3 of its sites | seat dropped; if fewer than 2 seats remain the run is void |
F1 |
if G1 fails, Q1 and Q2 are withheld and no override is available to this session |
— |
F2 |
if more than 1/3 of bodies are unusable after the permitted re-dispatches, the run is void | — |
G1 is the gate the whole run rests on and it is built to be failable. RS-20260814d's
analogue realised +0.8333 at marked sites on Danish. If the seats cannot see full ennoblement at
Denis's low places on this story, then a null at R26 says nothing, and the arm closes on that.
Re-dispatch cap: 6. A body that returns malformed or truncated may be re-dispatched; beyond six
across the run, remaining failures count as unusable and F2 applies.
6. Predictions, registered
Q1 — the primary, and it is the discriminator between (A) and (B).
Q1holds → explanation (B): the «Flipperne» near-zero was a span-length artefact, a one-step policy does raise at low places when they have room, andframework/v0.2§8 Q-e must carry the statement conditionally.Q1fails,G1passing → explanation (A): the policy does not raise at low places even with 79 words of room, on a second language pair, andRS-20260814d's sentence goes into the framework as written.
The lead's expectation, recorded so the result cannot be read as either way confirming it: I
expect Q1 to hold. Writing the arm, the rule fired at nearly every one of Denis's long turns
and the changes were substantial (49 contractions removed to 0, inversions, tense and case changes,
whole fragments made into sentences). If the seats cannot see that, the finding is much stronger than
if they can.
Q2 — the «Flipperne» condition, replicated. R26 − R06 at SHORT-LOW sites < +0.25.
Danish realised +0.0833. Predicted to hold.
Q3 — a SANITY CHECK, not a prediction (amendment A9, critic F-9): does the unruled arm fall on
ordinary prose? R06's coded height at NEUTRAL sites. It carries no licensing weight either way.
Danish realised −0.6000, and that fall is where the whole «Flipperne» effect came from.
Predicted here: |R06 at NEUTRAL| < 0.25 — i.e. the Danish neutral fall does NOT reproduce,
because on this source R26's N1 found nothing to hold up: the translator's log records the
magistrate's turns and the narration carried over from R06 verbatim. Registered as a prediction
that the neutral effect was a property of that rendering rather than of unruled passes in general.
C1 — the same-text control, and it is the run's noise floor. build_pool.py finds R06 and
R26 character-identical at 30 of 57 spans, including 25 of 26 of the magistrate's. Wherever
a selected site is one of those, the two arms are the same text under two labels, and the coded
difference must be 0.
Report mean |
R26−R06| over identical-text sites, with its denominator. If it exceeds 0.25, label noise is at the size of the primary, and by amendment A2 (critic F-2, BLOCKING) every licensing cell of §7 is void: the run reports noise only, makes no framework edit, and the arm closes on the second branch of its completion criterion. If fewer than 2 selected sites are identical-text,C1is reported unmeasurable, not silently dropped (amendment A6).
After A1, C1 lives on the NEUTRAL stratum, where 25 of the magistrate's 26 spans are
character-identical between the two arms.
This is RS-20260814h §4's control, which fired: a seat scored the same text 4 and then 1.
Q4 — exploratory, reported and licensed for nothing. R25 − R06 at LONG-LOW against
NEUTRAL: is full ennoblement's visible effect also concentrated at the low places, or spread?
Q1b — cutoff-free, added by amendment A5. Spearman rank correlation between the source
span's word count and (R26 − R06) across all eight selected LOW sites, per seat and
pooled. It uses no cutoff and so cannot be tuned by one. Exploratory; n = 8 sites. A per-site
table of word count against R26 − R06 is a required output, so the relation can be read
continuously rather than only through the two strata.
R25 − R26 at LOW — a required reported quantity, added by amendment A7. The three arm
means satisfy (R25 − R26) + (R26 − R06) = R25 − R06, so if G1 passes at +0.50 and
Q1 fails at +0.25, then R25 − R26 exceeds +0.25 necessarily. A Q1 null therefore arrives
with a demonstration, on the same seats, sites and scale, that an interval of exactly the disputed
size is resolvable in the adjacent gap. This does not make the null powered — n is 4 sites —
but it distinguishes "the instrument saw nothing" from "the instrument resolved something of this
size here, and not between these two arms."
7. What each outcome licenses, written before the numbers exist
G1 passes |
G1 fails |
|
|---|---|---|
Q1 holds |
Framework §8 Q-e carries the elevation statement conditioned on span length; ARM-elevation-resolution closes resolved on its completion criterion's first branch. |
withheld |
Q1 fails |
§8 Q-e records that the effect was not detected at 4 sites per stratum on a second language pair and at both span lengths, that (A) is not excluded and (B) is not supported, and that the statement RS-20260814d proposed remains untested rather than confirmed — reported beside R25 − R26 per A7; arm closes resolved. |
Arm closes resolved on the criterion's second branch: the instrument cannot resolve the interior, and Q-e carries that with the number. |
Every cell above is void if C1 > 0.25 (amendment A2). The Q1-fails cell was rewritten by
amendment A3: a null over 4 sites and 12 cells may not promote a sentence to a settled framework
claim, and the original wording — "carries RS-20260814d's statement as written" — did exactly
that.
In every cell the arm closes and the framework gains a measured statement, which is what the completion criterion demands. Nothing in this run can produce a fourth session on this arm.
8. Budget
Declared ceiling $1.20, against a $5.00 UTC-day cap with $0.00 spent on 2026-08-15 before this run. Pre-flight estimate, worst case computed from the caps the requests actually permit (note (abc)):
| stage | calls | worst case |
|---|---|---|
| stage 0 probe | 3 | $0.05 |
pre-run critic (P4) |
1 | $0.25 |
| stage 1 | 3 | $0.30 |
| stage 2 | 36 + up to 6 re-dispatches | $0.60 |
| total | $1.20 |
Lead translation of all three arms is $0 and is not ledgered (charter §3, A4), as are
counts.py, build_pool.py, the dependence table and the verifier.
9. Known limits, written before the run
- One story, one author, one language pair, one translator's idea of "one step".
R26andR25are this lead's assemblies, not Berman's protocol and not any published translator's declared policy. R06is dependence-flagged against Garnett at 37 twelve-grams. If that overlap has pulled the floor arm's register toward a 1921 hand, every difference measured from it is measured from a floor that is not neutral. Unmeasured here.R06~R26share 475 twelve-grams and a 115-token run. Conservative for the primary, disqualifying for any other use of the pair.- Nothing is calibrated. Tier D is NOT PASSED; every figure is
internal-judgment-onlyandprovisional. Three seats are not a jury. - n is small: 4 sites per stratum, 3 seats, 36 bodies. A stratum mean rests on 12 cells.
- The seats may recognise the story. Blinding covers provenance strings, not the Russian.
- The manipulation is transparent at LOW sites (amendment A8, critic F-8): three renderings of one paragraph at visibly different register heights, so a seat may infer that register elevation is the subject and that the scale invites a spread. Source-relative coding bounds this and does not remove it; demand characteristics are unmeasured here.
- The stratum cutoffs were chosen by the author with knowledge of the materials (A5).
Q1band the per-site table exist because of it, and neither removes the objection. - A source-relative scale cannot see a difference the two arms do not have. Where
R06andR26are identical,C1is a noise measurement and not evidence about register policy.