Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260815-register-room/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260815-register-room
statusfrozen
created2026-08-15
updated2026-08-15
sensesstyle-correspondence, voice, naturalness
internal-judgment-onlytrue
provisionaltrue
linkswiki/arms/ARM-elevation-resolution.md, wiki/findings/results/RS-20260814d-elevation-resolution.md, workshop/regimes/R26-light-ennoblement.md, workshop/regimes/R25-ennoblement.md, workshop/regimes/R06-lead-single-pass.md, workshop/translations/zloumyshlennik/R06-v1/translation.md, workshop/translations/zloumyshlennik/R26-v1/translation.md, workshop/translations/zloumyshlennik/R25-v1/translation.md, config/models.md, config/budget.md, framework/v0.2/README.md

E-20260815-register-room — does a one-step register policy act at a source's low places when they have room?

ARM-elevation-resolution step 2 (T4). Frozen before any API call of this run. The three translations this design measures were frozen at commit 3fa489e, before this file existed.

1. The question, and why it is not settled

RS-20260814d (S183) measured R26 (light ennoblement) against R06 (no register rule) on Andersen's «Flipperne» and found them +0.0833 apart at the sites where the Danish drops below its own level and +0.6667 apart where it sits at it. It deposited this reading:

A one-step-up register policy is not experienced as raising the source. It is experienced as not letting the translation drop below it — and the place it does that work is the ordinary run of the prose, not the marked places.

Its own §6 offers a rival explanation of the same numbers, and the arm cannot close honestly without separating them:

The question: on a source whose low register comes in long turns as well as in one-word items, does a one-step-up register policy show at the low places?

Chekhov's «Злоумышленник» (1885) is the material that separates (A) from (B) inside one text, which is why it is worth a session: the peasant Denis Grigoryev's 26 turns run from 1 word to 79 (pool: 1, 1, 2, 2, 6, 7, 7, 9, 10, 10, 13, 14, 19, 21, 21, 25, 26, 30, 30, 35, 38, 41, 46, 55, 60, 79). The «Flipperne» condition and its opposite are present in the same story, the same author, the same translator and the same three regimes, so span length is tested within the design rather than across two runs.

Subject rule check (wiki/tracks.md, continue-prompt.md §4.5), in one sentence: this measures whether a moderate register policy does any visible work where a source actually drops, and so tells a translator whether such a policy is worth adopting on colloquial dialogue or only on ordinary narrative prose. That is a claim about translating, not about the project's apparatus.

2. Materials — frozen, and all of them lead-produced at $0

arm regime words ratio to source role
R06 lead single pass, no register rule 1,396 1.300 floor
R26 light ennoblement, N1–N7 1,429 1.331 the arm under test
R25 ennoblement, E1–E8 1,886 1.756 calibration control

Source: 1,074 Russian words, 57 paragraphs, workshop/translations/zloumyshlennik/source.txt (stored at S165, unchanged). All three arms preserve the source's 57 paragraphs, asserted by build_pool.py and by the verifier, so span alignment is 1:1 and mechanical.

2.1 Contamination, measured before the design was written (dependence.json)

pair 7-grams 12-grams 15-grams longest run
R06 ~ R26 644 475 407 115
R06 ~ GARNETT (1921) 130 37 10 19
R26 ~ GARNETT 88 17 3 17
R25 ~ GARNETT 21 2 0 13
R06 ~ R25 18 0 0 10 — clean
R25 ~ R26 45 2 0 12
(reference) GARNETT ~ PANEL_P1 89 13 3 16
(reference) PANEL_P1 ~ PANEL_P3 190 68 42 30

Garnett's 1921 "A Malefactor" and the two 2026 panel renderings are stored in this repository from E-20260812i (S165). None of the three was opened by the translator; the table was produced by tools/dependence_check.py, which reads files without displaying them, after all three arms were frozen and committed.

Three things this table settles before the run, and one it does not:

  1. R06~R26 is the most dependent pair the project has measured — 475 shared twelve-grams and a 115-token run, far above note (bhb)'s previous densest firing (27 tokens, RS-20260808g; 113 twelve-grams, RS-20260814d). This is by construction: R26's N1 leaves the neutral alone and 30 of 57 spans are character-identical. ~~It is conservative for the primary — two arms that share this much are harder to tell apart, not easier~~ STRUCK by amendment A4 (critic F-4). That defence is a property of pairwise preference and detection tasks, and this design never asks a seat to compare two arms: every English span is coded against the Russian alone under labels scrambled per (site, seat), so the overlap does not enter the statistic at all — except at character-identical spans, where it is biasing rather than conservative, and where amendment A1 now excludes it from the LOW strata. The overlap remains fatal to any use of these two arms as independent renderings, which this design does not make.
  2. RS-20260814d §3's finding reproduces on a second language and a second author. The high rule set walks out of the plain register entirely: R06~R25 is clean at 0 twelve-grams between two renderings by the same translator, on the same day, from the same copy-text, while R06~GARNETT sits at 37. Overlap is a property of the register a policy targets, not of having a policy. Second measurement, second pair, same direction.
  3. R06 is dependence-flagged against Garnett (37 twelve-grams, a 19-token run). Under the standing rule (CLAUDE.md, contamination) the lead may not serve as the independent third translator where a design's validity turns on independence from a published rendering. This design's validity does not turn on it: no published hand is an arm, and every comparison is between two register policies of the same translator. Stated here so the exclusion is visible rather than assumed.
  4. What it does not settle: whether the lead's R06 is itself pulled toward Garnett's register by that overlap. If it is, R06 is not a neutral floor. Nothing here measures that, and §9 carries it as a limit.

3. Seats

Per config/models.md. P5 (deepseek/deepseek-v4-pro) is not used on any task shape, note (bne).

4. Procedure

Stage 0 — token-cap probe

tools/panel_probe.py-style probe per note (bnl): one throwaway call per coding seat on a two-span version of the stage-2 task, to read actual reasoning-token use, with caps set at the observed maximum + 400. Probe calls are ledgered.

Stage 1 — the site list, from the Russian alone

Each coding seat receives the whole Russian story and the 57 mechanically built spans (the source's paragraphs, in order, no selection whatever) and codes each span:

No English appears anywhere in the stage-1 prompt. The seat also answers three frozen comprehension questions about the Russian, which gate its retention:

  1. What object is the accused charged with having removed? (a nut / bolt from the railway)
  2. What does he say he and his fellow villagers make out of them? (fishing sinkers / weights)
  3. What does the official tell him at the end will happen to him? (he is to be taken into custody / sent to prison)

G0 — source-comprehension gate. A seat answering fewer than 2 of 3 correctly is dropped from the run entirely, and its stage-1 codes are discarded. This is RS-20260811g's lesson (a seat that could not read the source still produced confident codes) applied before the money is spent on stage 2.

The site list is the three-seat majority, exactly as RS-20260808e's and RS-20260814d's. A span with no majority is excluded.

Stage 2 — register height, blind, one site at a time

Sites are selected mechanically from stage 1's majority classification crossed with the mechanical word count of the source span:

stratum rule n
LONG-LOW majority below and source span ≥ 25 words and R06 ≠ R26 4
SHORT-LOW majority below and source span ≤ 7 words and R06 ≠ R26 4
NEUTRAL majority at 4

Within each stratum, spans are ordered by the SHA-256 of the source span text (already in pool.json) and the first four are taken. No span is chosen by the lead, and the ordering was fixed before stage 1 was run.

The identical-text exclusion on the LOW strata is amendment A1 (critic F-1, BLOCKING). Where the two arms are the same text they cannot differ, and such a site contributes a structural zero that attenuates Q1 and makes Q2 true by construction. identical_R06_R26 was computed by build_pool.py before stage 1 existed. Its effect is small and is itself reported: of Denis's 26 turns exactly two are identical (spans 10 and 34), so the exclusion removes at most one SHORT-LOW candidate and none at all from LONG-LOW, leaving 11 long and 6 short candidates against a G2 floor of 3. R26's rule fired at every one of Denis's eleven long turns and at six of his seven short ones — the translator-side fact a reader needs in order to read a null.

The cutoffs 25 and 7 are author-chosen, and amendment A5 (critic F-5) requires that to be said here rather than implied away. Their basis: Denis's 26 turns have a natural gap in the length distribution between 21 and 25 words (…13, 14, 19, 21, 21, 25, 26, 30…), and ≤ 7 is the band that reaches the «Flipperne» condition, whose low sites were under four words. That basis was chosen with knowledge of the materials by the person who wrote the arms, which is why A5 also adds a cutoff-free statistic (Q1b) and a required per-site table.

Each body is one site × one seat. It contains the Russian span, and the three English renderings of that span under labels scrambled independently for every (site, seat) pair. The seat codes each English span against the Russian span:

12 sites × 3 seats = 36 bodies. Judgment is not parallelised across sites within a seat.

Blinding. No prompt in either stage contains: the regime ids R06/R25/R26/R08/R21, the words ennoblement, ennoble, register policy, regime, Berman, Chekhov, Garnett, translator's log, lit-trans, experiment, hypothesis, the arm word counts, or any statement that the arms differ in policy. Asserted by the verifier over the stored prompts, matched on word boundaries. The seats can still recognise the story from the Russian; that is unavoidable in a source-relative design and is a limit (§9), not a leak — recognition does not reveal which label is which arm.

Runner discipline, note (bnx). Every body is appended to runs/bodies.jsonl as it returns; re-invocation resumes from what is already on disk; the dispatch runs in the background with no foreground wall clock. S187 lost $0.312883 to exactly this and the fix is three lines.

5. Gates — every one of them registered here, before dispatch

gate rule what fails if it fires
G0 each seat answers ≥ 2 of 3 source-comprehension questions that seat is dropped
G1 R25 − R06 at LOW sites (LONG-LOW ∪ SHORT-LOW) ≥ +0.50 Q1, Q2 withheld
G2 each stratum has ≥ 3 qualifying spans after the majority rule that stratum's primary is withheld
G3 a seat is retained only if it returns usable codes at ≥ 2/3 of its sites seat dropped; if fewer than 2 seats remain the run is void
F1 if G1 fails, Q1 and Q2 are withheld and no override is available to this session —
F2 if more than 1/3 of bodies are unusable after the permitted re-dispatches, the run is void —

G1 is the gate the whole run rests on and it is built to be failable. RS-20260814d's analogue realised +0.8333 at marked sites on Danish. If the seats cannot see full ennoblement at Denis's low places on this story, then a null at R26 says nothing, and the arm closes on that.

Re-dispatch cap: 6. A body that returns malformed or truncated may be re-dispatched; beyond six across the run, remaining failures count as unusable and F2 applies.

6. Predictions, registered

Q1 — the primary, and it is the discriminator between (A) and (B).

R26 − R06 at LONG-LOW sites ≥ +0.25.

The lead's expectation, recorded so the result cannot be read as either way confirming it: I expect Q1 to hold. Writing the arm, the rule fired at nearly every one of Denis's long turns and the changes were substantial (49 contractions removed to 0, inversions, tense and case changes, whole fragments made into sentences). If the seats cannot see that, the finding is much stronger than if they can.

Q2 — the «Flipperne» condition, replicated. R26 − R06 at SHORT-LOW sites < +0.25. Danish realised +0.0833. Predicted to hold.

Q3 — a SANITY CHECK, not a prediction (amendment A9, critic F-9): does the unruled arm fall on ordinary prose? R06's coded height at NEUTRAL sites. It carries no licensing weight either way. Danish realised −0.6000, and that fall is where the whole «Flipperne» effect came from. Predicted here: |R06 at NEUTRAL| < 0.25 — i.e. the Danish neutral fall does NOT reproduce, because on this source R26's N1 found nothing to hold up: the translator's log records the magistrate's turns and the narration carried over from R06 verbatim. Registered as a prediction that the neutral effect was a property of that rendering rather than of unruled passes in general.

C1 — the same-text control, and it is the run's noise floor. build_pool.py finds R06 and R26 character-identical at 30 of 57 spans, including 25 of 26 of the magistrate's. Wherever a selected site is one of those, the two arms are the same text under two labels, and the coded difference must be 0.

Report mean |R26 − R06| over identical-text sites, with its denominator. If it exceeds 0.25, label noise is at the size of the primary, and by amendment A2 (critic F-2, BLOCKING) every licensing cell of §7 is void: the run reports noise only, makes no framework edit, and the arm closes on the second branch of its completion criterion. If fewer than 2 selected sites are identical-text, C1 is reported unmeasurable, not silently dropped (amendment A6).

After A1, C1 lives on the NEUTRAL stratum, where 25 of the magistrate's 26 spans are character-identical between the two arms.

This is RS-20260814h §4's control, which fired: a seat scored the same text 4 and then 1.

Q4 — exploratory, reported and licensed for nothing. R25 − R06 at LONG-LOW against NEUTRAL: is full ennoblement's visible effect also concentrated at the low places, or spread?

Q1b — cutoff-free, added by amendment A5. Spearman rank correlation between the source span's word count and (R26 − R06) across all eight selected LOW sites, per seat and pooled. It uses no cutoff and so cannot be tuned by one. Exploratory; n = 8 sites. A per-site table of word count against R26 − R06 is a required output, so the relation can be read continuously rather than only through the two strata.

R25 − R26 at LOW — a required reported quantity, added by amendment A7. The three arm means satisfy (R25 − R26) + (R26 − R06) = R25 − R06, so if G1 passes at +0.50 and Q1 fails at +0.25, then R25 − R26 exceeds +0.25 necessarily. A Q1 null therefore arrives with a demonstration, on the same seats, sites and scale, that an interval of exactly the disputed size is resolvable in the adjacent gap. This does not make the null powered — n is 4 sites — but it distinguishes "the instrument saw nothing" from "the instrument resolved something of this size here, and not between these two arms."

7. What each outcome licenses, written before the numbers exist

G1 passes G1 fails
Q1 holds Framework §8 Q-e carries the elevation statement conditioned on span length; ARM-elevation-resolution closes resolved on its completion criterion's first branch. withheld
Q1 fails §8 Q-e records that the effect was not detected at 4 sites per stratum on a second language pair and at both span lengths, that (A) is not excluded and (B) is not supported, and that the statement RS-20260814d proposed remains untested rather than confirmed — reported beside R25 − R26 per A7; arm closes resolved. Arm closes resolved on the criterion's second branch: the instrument cannot resolve the interior, and Q-e carries that with the number.

Every cell above is void if C1 > 0.25 (amendment A2). The Q1-fails cell was rewritten by amendment A3: a null over 4 sites and 12 cells may not promote a sentence to a settled framework claim, and the original wording — "carries RS-20260814d's statement as written" — did exactly that.

In every cell the arm closes and the framework gains a measured statement, which is what the completion criterion demands. Nothing in this run can produce a fourth session on this arm.

8. Budget

Declared ceiling $1.20, against a $5.00 UTC-day cap with $0.00 spent on 2026-08-15 before this run. Pre-flight estimate, worst case computed from the caps the requests actually permit (note (abc)):

stage calls worst case
stage 0 probe 3 $0.05
pre-run critic (P4) 1 $0.25
stage 1 3 $0.30
stage 2 36 + up to 6 re-dispatches $0.60
total $1.20

Lead translation of all three arms is $0 and is not ledgered (charter §3, A4), as are counts.py, build_pool.py, the dependence table and the verifier.

9. Known limits, written before the run