Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260814d-elevation-resolution/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260814d-elevation-resolution
statusfrozen
created2026-08-14
updated2026-08-14
sensesstyle-correspondence, voice, cultural-mediation
internal-judgment-onlytrue
provisionaltrue
linkswiki/arms/ARM-elevation-resolution.md, workshop/regimes/R26-light-ennoblement.md, workshop/regimes/R25-ennoblement.md, workshop/translations/flipperne/R26-v1/translation.md, wiki/findings/results/RS-20260813c-ennoblement-direction.md, wiki/findings/results/RS-20260813g-register-quadrants.md, framework/v0.2/README.md, config/models.md, config/budget.md

E-20260814d — a register axis with three rungs, and a gate that can fail short of ceiling

Frozen 2026-08-14 (S183) before any call was dispatched. ARM-elevation-resolution step 1.

1. The question, and why the third rung is the whole design

framework/v0.2 §7.4: "It says nothing about elevation: the run's Axis E voided on a dead body and where published hands sit above plain English is still unmeasured." That is Q-e's open half.

The predecessor run (RS-20260813c) built the upper pole and could not read the interior. Its gate G1 required the deliberately ennobled arm to read above the plain arm by ≥ +0.50 on a −1 / 0 / +1 scale, and got +1.0000 in 56 of 56 cells, zero variance, while the published hands it existed to measure sat at +0.125 to +0.70. Method note (bnb): a calibration gate that passes at ceiling has not calibrated anything in the range the primaries occupy.

On a source whose own register is uniformly low, how far above it do published translators pitch their English — and can this instrument resolve the interval at all?

The answer to the second half gates the first. A third rung (R26, light ennoblement) is built so that the calibration gate is a one-step movement rather than a maximal one, and can therefore fail.

2. Materials

Source. H. C. Andersen, «Flipperne» (1848), 756 Danish words in 26 paragraphs, copy-text workshop/experiments/E-20260813g-register-quadrants/corpus/DA.txt, re-fetched from Danish Wikisource this session and identical to the stored file apart from one paragraph break. Single-witness and uncollated; two transcription slips recorded on the translation artifacts.

Seven English arms, all of the whole tale, all held in workshop/experiments/E-20260813g-register-quadrants/corpus/:

arm hand role
A Gutenberg #1597, translator unattributed in the edition published
PAULL Mrs. H. B. Paull, 1888 published, dependence-flagged
BRAK H. L. Brækstad, 1900 published
R06 lead, single pass, no register rule lower pole
R26 lead, light ennoblement (N1–N7), written this session middle rung
R25 lead, ennoblement (E1–E8) upper pole
R08 lead, resistancy source-ward arm, exploratory only

2.1 The materials gate, run before this design was written, and it constrains it

tools/dependence_check.py over all 21 pairs. The ten pairs not involving R26 reproduce RS-20260813c §2.1 exactly, which is a check on the tool as well as on the texts.

pair 7-grams 12-grams 15-grams longest run verdict
A ~ BRAK 28 0 0 11 clean
A ~ PAULL 24 5 0 14 DEPENDENT?
PAULL ~ BRAK 41 7 2 16 DEPENDENT?
A ~ R06 49 8 5 19 DEPENDENT?
BRAK ~ R06 38 5 1 15 DEPENDENT?
PAULL ~ R06 18 1 0 12 DEPENDENT?
A ~ R25 18 0 0 11 clean
PAULL ~ R25 9 0 0 10 clean
BRAK ~ R25 9 0 0 11 clean
R06 ~ R25 7 0 0 11 clean
A ~ R08 27 0 0 11 clean
BRAK ~ R08 10 0 0 10 clean
PAULL ~ R08 7 0 0 9 clean
R06 ~ R08 48 8 2 16 DEPENDENT?
R08 ~ R25 9 1 0 12 DEPENDENT?
R26 ~ R06 263 113 72 27 DEPENDENT?
R26 ~ R08 67 17 5 19 DEPENDENT?
R26 ~ BRAK 72 20 3 15 DEPENDENT?
R26 ~ PAULL 31 5 2 16 DEPENDENT?
R26 ~ A 63 2 0 12 DEPENDENT?
R26 ~ R25 9 0 0 11 clean

Five consequences, all binding, all written before any code was seen.

  1. R26 is not an independent hand and no claim about published practice may rest on it. It is a constructed pole. Q3 is the only prediction that puts R26 beside a published hand and it is secondary and dependence-flagged, with R26~BRAK at 20 shared twelve-grams named at every use.
  2. R06 is contaminated with A (run 19) and BRAK (run 15) and is used only as the lower pole, never as an independent hand. RS-20260813c §6's argument carries over: that contamination pulls R06 up, which shrinks every X − R06 difference, so G1 and G2 are conservative gates — the contamination can make them fail and cannot make them pass spuriously. But a second bias runs the other way and the two are stated together, per critic finding 7(b): the crib is a lead-written word-for-word English rendering, and so is R06. A seat anchoring on the crib would push whichever arm most resembles it toward 0, inflating every X − R06 difference — so G1/G2 are conservative with respect to contamination and not with respect to crib anchoring, and the design does not claim otherwise. One handle on it is free and is registered here: the arm that most resembles a word-for-word crib is R08, which calques, not R06, so crib anchoring predicts R08 sitting nearer 0 than R06 does. That comparison is reported.
  3. A ~ BRAK is still the only clean published pair. PAULL is dependence-flagged against both others, so the three published arms are not three independent observations. Q1 and Q2 are stated per hand on the clean pair, with PAULL reported alongside as a dependent sensitivity.
  4. R26 ~ R06 is the closest pair in the table by a wide margin, so a design using their contrast as its manipulation must show they differ at the grain the gate runs at — critic finding 6, and the paragraph-level check the design first carried does not entail it. Checked per candidate span before dispatch and recorded in pool.json: token-level similarity over the whole texts is 0.720 against 0.388 for R06~R25; of the 27 candidate spans, 1 is identical (S23a, «Forlovet!»), 3 differ only in terminal punctuation (S10a, S18a, S20a), and the remaining 23 differ lexically. So 4 of 27 spans carry no lexical manipulation at all, and if they are sampled they contribute cells where G1 measures nothing. G1 is therefore computed twice and both figures are reported: over all marked sites, and over the marked sites where R06 and R26 differ lexically. The first is the registered gate; the second is what says whether a failure is the instrument's or the materials'.
  5. R26 ~ R25 is clean at 0 shared twelve-grams, so the two rungs above the baseline are not each other's copies.

3. Stage 1 — the site list, three annotators, source only

The repair this stage exists for: RS-20260813c §8 measured its own site list against two independent readers and six of ten lead-marked sites were read as unmarked. The list was one annotator's, and that annotator was the lead. RS-20260808e's three-annotator majority vote is the procedure being copied.

3.1 The candidate pool is built mechanically, with no lead judgement in it. Frozen rule, applied to the copy-text before any annotation:

Realised pool: 27 candidates, 20 speech and 7 narration, frozen in pool.json before dispatch.

3.2 The annotation. Three seats — P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5 — each receive the whole Danish tale and the 27 numbered spans, and return one code per span:

Relative to the ordinary written Danish of an 1848 storybook, does the Danish at this span read BELOW that level, AT it, or ABOVE it? → below / at / above.

No English appears anywhere in a stage-1 prompt. Output is a closed form — one code per numbered span, 27 codes — per note (bng).

3.3 The majority rule and the sample, both frozen here.

3.4 The dependence this creates, declared rather than discovered. The three annotators are the same three seats that judge in stage 2. Calls are stateless and no stage-2 prompt contains a stage-1 answer, so the dependence is at the level of a model's dispositions, not of memory — but a seat that reads a span as low may be likelier to read an English rendering as above it, and that is stated as a limit rather than argued away. The alternative — annotating with the reserve seats — was rejected because their Danish is entirely unprobed and P5 is excluded outright by note (bne).

4. Stage 2 — the height run

One call per (site × seat), 3 seats, ≤ 15 sites: at most 45 calls. P5 is not seated, on note (bne). Judgment is not parallelised: each call is stateless and sees one site.

Each site payload carries, for the site the annotators selected:

Each call asks two things, both closed-form:

Two instructions carried over from the predecessor's amended prompt, because they were accepted findings there: length is not level — judge the words chosen, not the number of them (critic finding 7), and no rendering is described as "independent" (finding 9).

Blinding assertions, split per note (bnm). The scaffolding is fixed text and carries the ban absolutely: no prompt contains any of the 19 provenance strings Andersen, Paull, Brækstad, Braekstad, Berman, ennobl, regime, foreigniz, resistancy, Venuti, vulgaris, Gutenberg, R06, R08, R25, R26, 1888, 1900, 1597, matched as substrings (several are stems). 1848 is deliberately NOT on this list — critic finding 1, and it was a BLOCKING error: the date is the period anchor the question itself needs ("the ordinary written Danish of an 1848 storybook"), so the frozen assertion was unsatisfiable the moment it was written, exactly as note (bnm) describes. 1848 identifies no translator and no arm; every English arm in the run postdates it. The crib is lead-written data and carries the 19 register words ban — colloquial, homely, blunt, everyday, literary, elevated, decorous, insult, shout, spoken, plain, formal, register, low, high, ornate, dignified, slang, vulgar — matched on word boundaries, not as substrings, because low is inside below and plain is inside explain (critic finding 10). The verifier counts any breach rather than assuming none.

Caps, per note (bnl). Reasoning is never disabled on a judging seat. One probe call per seat is dispatched on a real stage-2 prompt, reasoning_tokens is read off the response, and the cap for the remaining dispatches is set at measured reasoning + 400 for the answer, floored at 1,200 and ceilinged at 3,000. The probe's own cap is 2,500. Stage 1's cap is 4,000 for all three seats and is not probed, because the stage is only three calls.

Re-dispatch, per note (bnn). A body is usable only if it parses as JSON and carries its stage's key (levels for stage 1, source_level and heights for stage 2) with the right arity. Anything else is re-dispatched once at 2× the cap but never above 4,000 tokens, and the stored record is marked usable: false so the analysis and the F3 count key off the same fact. Stage 2's re-dispatches are capped at 6 in total (critic finding 4): beyond the sixth, an unusable body is recorded usable: false without a retry, and the number not retried is reported. Without that cap the design's own rules permit 42 extra calls and the declared ceiling is not a ceiling.

5. Gates, run before any prediction is read

Aggregation, fixed here for every gate and prediction: a figure is the mean of seat-means over seats surviving G4, computed over cells that returned a usable body. Minimum 2 surviving seats.

6. Predictions, registered before any call

7. Failure criteria, registered

8. What this run cannot establish

9. Money

Pre-flight priced from config/models.md as corrected 2026-08-14 (P2 $0.75/$3.75), from max_tokens and not from an assumed answer length (note (abc)), with re-dispatch in the arithmetic (note (abc)'s S079 firing).

stage calls cap worst case
stage 1 — annotation 3 (+3 re-dispatch) 4,000 $0.16
stage 2 — reasoning probe 3 2,500 $0.06
stage 2 — main 42 ≤ 3,000 $0.95
stage 2 — re-dispatch, capped at 6 (critic finding 4) ≤ 6 ≤ 4,000 $0.16
pre-run critic (P4, effort low, note (bnk)) 1, actual $0.112383 6,000 $0.25
declared ceiling for this run $1.50

Session ceiling $1.50. UTC-day headroom at design time: $2.763776 of $5.00 remaining (2026-08-14, three prior sessions at $2.236224). A run that does not fit is split or deferred; F4 and F5 are both routes to spending materially less than the ceiling and that is a normal outcome.

Lead translation of «Flipperne» under R26 is $0 and is not ledgered (charter §3, A4), as are the copy-text re-fetch, the dependence measurement, the mechanical pool and the verifier.

10. Verification

analysis/verify.py, importing nothing from tools/, recomputes every reported number from the stored bodies and asserts:

  1. every stage-1 and stage-2 prompt is free of all 19 banned provenance strings;
  2. every crib is free of all 19 banned register words, and the count of breaches is reported rather than assumed zero (note (bnm));
  3. every arm span in every payload is a contiguous substring of that arm's stored corpus file, normalising whitespace only — the guard against a rendering being paraphrased into the payload;
  4. the label scramble is a bijection at every (site, seat) and agrees with the runner's record;
  5. every individual cell code matches the stored body, not merely the aggregate — RS-20260813b §8's repair, where a flipped cell survived a majority-level check;
  6. every reported mean recomputes from the per-cell codes;
  7. the site list recomputes from the stage-1 bodies under §3.3's majority rule and hash order;
  8. mutation sensitivity: flipping one stored code, and separately corrupting one payload span, each change at least one asserted figure.

11. Amendment before dispatch: the pre-run critic's findings, all ten accepted

Critic: moonshotai/kimi-k3 (P4), no role in this run's annotation or its jury (P1/P2/P3), and not a seat in it at all. Ten findings, six BLOCKING, four NON-BLOCKING. All ten accepted; two of them in an amended form with the reason written. The design above is the amended one, and nothing had been dispatched. Cost $0.112383000, provider Chutes, 4,034 reasoning tokens of a 6,000 cap. Raw body: run/critic.json.

The verdict line did not arrive: the body hit finish_reason: length mid-finding-10, so the critic's own summary verdict is missing and finding 10 is truncated after its first sentence. The sentence is complete enough to act on and was acted on. Note (bnk) fired again, on the seat it names, with its own remedy already applied — effort pinned low in the first dispatch and the cap sized from the observed 6,000-token pass at S181 — and it still truncated, because ten findings is a longer answer than eight. Note (bng) is the more exact diagnosis: an open-ended enumeration is the one output shape whose cap cannot be guessed, and a critique is exactly that shape. The finding is recorded rather than repaired: a second dispatch at a larger cap would have bought a verdict word this session does not need, at the price of a second $0.11.

# finding disposition
1 BLOCKING — the blinding ban listed 1848, and the design's own question text contains 1848. The verifier's first assertion would have failed on every prompt, before any model was called. Accepted, and it is the most serious finding. 1848 struck from the provenance list, which is now 19 strings; the reason is written into §4. The date is a period anchor the question needs and identifies no translator and no arm. This is note (bnm) exactly — a blindness assertion the design could not hold, unsatisfiable the day it was frozen — caught this time by a critic instead of by a verifier after the money was spent.
2 BLOCKING — G4's "majority of the other seats" is undefined with three seats. A majority of two requires both, and the statistic does not exist when they split — on the path to every prediction through F2. Accepted. G4 rebuilt as agrees with at least one other seat. §5.
3 BLOCKING — G4's fixed bar of 8 sits over a denominator that can be smaller than 8, so on the F4-minimum path (9 sites) a single dead body would void the run by arithmetic. Accepted, and it is note (bna) reintroduced one clause after the note was cited by name. The bar is now a proportion (≥ 2/3 of eligible sites) with a frozen minimum eligible count of 6, below which G4 is indeterminate and reported as not having run rather than failing everybody.
4 BLOCKING — the stage-2 re-dispatch tail is unbounded, so $1.42 was not a worst case. The design's own rules permit 42 extra calls at up to 6,000 tokens against $0.03 of slack. Accepted. Re-dispatches capped at 6 and at 4,000 tokens, the number not retried is reported, and the declared ceiling is raised to $1.50 with the row priced. Note (abc)'s S079 firing said a worst case is max_tokens × attempts × slugs; the attempt count was missing from one row and the critic found it.
5 BLOCKING — the provenance list has 20 strings and §10 asserts 19 twice, so either one string goes unverified or the frozen text is wrong. Accepted. With 1848 struck (finding 1) the provenance list is 19 and the register list is 19; §4 and §10 now name which count belongs to which list.
6 BLOCKING — the manipulation check is at paragraph grain and the gate runs at span grain. If the sampled marked sites are spans where R26 = R06, G1 measures nothing and fails for a materials reason the design would report as an instrument reason. Accepted in full and it is the finding that most changed the run. Every one of the 27 candidates was diffed before dispatch and the result is in pool.json: 1 identical, 3 punctuation-only, 23 lexically different. §2.1(4) now carries the numbers, and G1 is computed and reported twice — over all marked sites, and over those where the two arms differ lexically.
7 NON-BLOCKING — §2.1(2)'s "contamination is conservative" argument does not cover Q3, and the crib introduces a bias in the opposite direction that no section mentions. A literal crib resembles R06 more than any other arm; a seat anchoring on it pushes R06 toward 0 and inflates every X − R06 difference, so G1/G2 can pass spuriously after all. Accepted in full, both halves. §2.1(2) now states the crib-anchor bias and its direction and withdraws the unqualified conservatism claim; Q3 now says its threshold can be passed by contamination and not only failed by it. One amendment beyond what the critic asked: the arm that most resembles a word-for-word crib is R08, which calques, not R06 — so crib anchoring makes a prediction about R08, and that comparison is now reported as a handle on the bias rather than left as an unfalsifiable worry.
8 NON-BLOCKING — the ≥ 4-word floor applies only to narration, so up to 6 of 15 sites could be single interjections, and the "mechanical, judgement-free" rule embeds an unjustified asymmetry. Accepted as a cap rather than as the critic's floor, with the reason written. Applying the floor to speech would drop «Snærpe!», «Las!» and «Forlovet!» — the sites where this source's register most obviously lives, and the ones the predecessor run coded. The asymmetry now carries its justification (a short unquoted remnant is a speech tag with no register content; a short speech turn is a complete utterance), and at most 3 of the 15 sampled sites may be shorter than 4 Danish words.
9 NON-BLOCKING — seats and predictions both use the labels P1–P5, so P5 names an excluded seat in §4 and a live prediction in §6. Accepted, in the form that costs no cross-references. The predictions are renamed Q1–Q5; the seats keep the names config/models.md gives them, since renaming a seat here would desynchronise this design from every other page in the project.
10 NON-BLOCKING (truncated) — the banned-word screen substring-matches, so low fires inside "below" and plain inside "explain". Accepted. The register screen now matches on word boundaries; the provenance screen stays substring, because several of its entries (ennobl, foreigniz, vulgaris) are deliberately stems.

One thing that is not the critic's and is recorded so the amendment list is honest: the site hash keys on span_id only, so none of these amendments changes which spans the sample will draw or which label any arm sits behind.

12. Pre-dispatch amendment A1 — after stage 1, before any stage-2 call

Written and committed 2026-08-14 after stage 1 returned and before a single stage-2 call was dispatched. Stage 1's three annotators returned 3 of 3 usable bodies and produced a majority site list far thinner and far less agreed than the design anticipated: 7 marked, 13 neutral, 0 above, and 7 spans with no majority at all (every one of those seven split three ways, above/below/at). F4 does not fire — its bar is 6 marked and 3 neutral by majority, and both are met.

But two frozen rules then interact in a way nobody foresaw at freeze time. Six of the seven marked spans are shorter than four Danish words, and §3.3's short-span cap (critic finding 8) allows at most three such spans in the sample. The frozen sample is therefore 4 marked sites, of which — by §2.1(4)'s own per-span table — S23a is lexically identical between R06 and R26 and S20a differs only in terminal punctuation. Half the cells of the registered resolution gate would carry no manipulation at all.

The amendment, and what it deliberately does not do.

Two stage-1 facts recorded here rather than in the result, because they bear on the dispatch. P3 returned 6,516 reasoning tokens against a max_tokens of 4,000 and still delivered a complete body — so on that provider the cap did not bound the reasoning at all, and the call cost $0.045616 against a per-call worst case of $0.0272 built from the cap. Note (abc)'s worst case is not a bound on a reasoning seat, which is note (bnk)'s point arriving from the other side. And the three annotators' base rates for below over the same 27 spans are 7, 20 and 3, which is the measurement §5's G4 was written to catch and the reason the stage-2 probe is worth its three calls.