Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260813c-ennoblement-direction/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260813c-ennoblement-direction
statusfrozen
created2026-08-13
updated2026-08-13
sensesstyle-correspondence, voice, cultural-mediation
internal-judgment-onlytrue
provisionaltrue
linkswiki/arms/ARM-ennoblement.md, workshop/regimes/R25-ennoblement.md, workshop/regimes/R21-vulgarisation.md, workshop/translations/flipperne/R06-v1/translation.md, workshop/translations/flipperne/R25-v1/translation.md, wiki/base/sources/S-berman-tendances.md, framework/v0.2/README.md, config/models.md, config/budget.md

E-20260813c — where a published translator sits on a register axis that now has both ends

Frozen 2026-08-13 (S174) before any call was dispatched. ARM-ennoblement step 1.

1. The question, and why it needs a new pole

Berman's fifth and sixth deforming tendencies are one axis: ennoblissement upward, vulgarisation downward. The project built the downward end (R21, S124) because a measurement found it empty. The upward end has never existed, so every register figure this project has published — including RS-20260808e's ten published hands across five language pairs, pooled REG +0.9716, no hand lowering where the source lowers — was measured on a scale with a low pole, a plain pole, and nothing above the line. framework/v0.2 §7 records that it cannot make a register-carriage recommendation.

R25 (frozen today, before either translation was written) supplies the upper pole. The question this run asks is the one that becomes askable only once it exists:

On a source whose own register is uniformly low, where do published translators sit on a register axis that has both ends — at the plain pole, or up beside a rule set built to ennoble on purpose?

A subsidiary question the materials force and the design does not dodge: is Berman's ennoblement, executed deliberately, doing something more extreme than what Victorian translators actually did, or is it merely a description of them?

2. Materials

Source. H. C. Andersen, «Flipperne» (1848), 767 Danish words — the project's first text in Danish, and its sixth language pair for this instrument. Chosen because Andersen's register here is spoken and low almost throughout, which is the condition under which ennoblement and any source-ward markedness point in opposite measurable directions.

Five English arms, all of the whole tale:

arm hand status
A Gutenberg #1597 "The False Collar", translator unattributed in the edition published
PAULL Mrs. H. B. Paull, 1888 (verified: Gutenberg #27200 is the Paull text) published
BRAK H. L. Brækstad, 1900 published
R06 lead single pass, no register rule lead, primed and contaminated — see §2.1
R25 lead, ennoblement rule set E1–E8 lead, clean

2.1 The materials gate, run before this design was written, and it constrains the design. tools/dependence_check.py, all pairs:

pair shared 7-grams 12-grams 15-grams longest run verdict
A ~ PAULL 24 5 0 14 DEPENDENT?
A ~ BRAK 28 0 0 11 clean
PAULL ~ BRAK 41 7 2 16 DEPENDENT?
A ~ R06 49 8 5 19 DEPENDENT?
BRAK ~ R06 38 5 1 15 DEPENDENT?
PAULL ~ R06 18 1 0 12 DEPENDENT?
A ~ R25 18 0 0 11 clean
PAULL ~ R25 9 0 0 10 clean
BRAK ~ R25 9 0 0 11 clean
R06 ~ R25 7 0 0 11 clean

Three consequences, all binding:

  1. R06 is contaminated and may not be read as an independent hand. Its own artifact already restricted it to the lower calibration pole on the ground of declared session priming; the measurement confirms the restriction rather than discovering it. No claim in this design rests on R06's position relative to a published hand — it is used only against R25.
  2. The only clean published pair is A ~ BRAK. PAULL is dependence-flagged against both other hands, so the three published arms are not three independent observations, and no pooled published figure here may be read as a three-hand agreement. P1 is therefore stated per hand.
  3. R25 is clean against every hand. The ennobled arm, written after the plain one by the same primed translator on the same day, shares no 12-gram with anything.

3. Sites

14 sites, frozen in sites.json before dispatch: 10 marked (the Danish sits below its own neutral written level — bare colloquial exclamations, homely compounds, the particle cluster «jo/vel/nok», the third-person address-drop) and 4 neutral (plain narration at the story's ordinary level). Each site carries the Danish, a literal English gloss of the Danish written by the lead, and the five arms' renderings of that site.

4. Procedure

One call per (site × seat). 42 calls: 14 sites × 3 seats.

Seats: P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P5 deepseek/deepseek-v4-pro (the standing non-Anthropic jury; config/models.md). Judgment is not parallelised across sites within a seat's reasoning — each call is stateless and sees one site only.

Blinding. Arms are presented under labels V1…V5, scrambled per (site, seat) by sha256(site_id | seat | "E-20260813c") mod 120, so no arm sits behind a fixed label and order is controlled without paying for two orderings. No provenance, no date, no translator name, no mention of a regime, of Berman, of ennoblement, or of this project appears in any prompt.

Each call asks two things:

Answers returned as strict JSON. Caps sized per note (b)/(bmb): max_tokens 1500 with reasoning_effort pinned low, and the re-dispatch guard triggers on a body that fails to parse, irrespective of finish_reason — note (bmz), whose remedy is applied here for the first time.

5. Gates, run before any primary is read

Aggregation, fixed here for every gate and prediction: a figure is the mean of seat-means over seats surviving G3, computed over cells that returned a usable body. Minimum 2 surviving seats. (Critic finding 8.)

6. Predictions, registered before any call

§2.1's dependence table governs which of these can be primary. Critic finding 5 is accepted: PAULL is dependence-flagged against both other published hands, so it cannot supply an independent observation, and A~BRAK is the only clean published pair.

Why G1/G2 may rest on the contaminated R06 arm, stated rather than assumed (critic finding 5's last clause). R06's contamination is with A (run 19) and BRAK (run 15), both of which sit above the Danish on the hypothesis under test. Contamination therefore pulls R06 up, which shrinks R25 − R06. G1 and G2 are consequently conservative gates: the contamination can cause them to fail, and cannot cause them to pass spuriously. No shared n-gram between R06 and a published hand raises R25, which shares none with anything.

7. Failure criteria, registered

8. What this run cannot establish

9. Money

Pre-flight, priced from config/models.md at max_tokens 1500 (note (abc): the worst case is built from the cap, not from an expected length), with a 4× routing surcharge on every seat:

seat per call worst × 14
P1 $0.0099 $0.139
P2 $0.0126 $0.177
P5 $0.0017 × 4 $0.095
subtotal $0.411
× 4 routing surcharge applied to P1/P2 as well ≈ $1.10
pre-run critic, one call $0.10
declared ceiling for this run $1.20

Session ceiling $1.40 including the $0.160873565 already spent on the D-20260813-17 gate. Day headroom at design time: $2.04 of $5.00 remaining.

10. Verification

analysis/verify.py, importing nothing from tools/, recomputes every reported number from the stored bodies and asserts: every prompt is free of arm names, provenance, dates and the words Berman, ennoble, regime, foreignizing; the scramble is a bijection at every (site, seat); each reported mean recomputes from the stored per-cell codes; and every individual cell code matches the record, not merely the aggregate — RS-20260813b §8's repair, where a flipped cell survived a majority-level check.

11. Amendment before dispatch: the pre-run critic's findings, accepted

Critic: x-ai/grok-4.5 (P3), no role in this run's jury (P1/P2/P5). Verdict NEEDS-AMENDMENT, 5 BLOCKING and 4 NON-BLOCKING findings. All 9 accepted; the design above is the amended one, and nothing had been dispatched. Cost $0.0177364. Raw body: run/critic.json.

# finding disposition
1 BLOCKING — the gloss was a register leak. The frozen glosses described the Danish's register ("a bare noun-shout", "a sharp everyday insult", "homely"), which hands the seats the answer to probe (a) and biases (b) toward whichever rendering resembles the gloss's own English. Accepted in full and it is the most serious finding. Every gloss rewritten to word-for-word crib plus dictionary senses, with a banned-vocabulary check run over all 14 (no colloquial, homely, blunt, everyday, literary, elevated, decorous, insult, shout, spoken, plain, formal, register, low, high). The prompt now names the crib as "not English prose" and instructs that it is not a yardstick of level. This is the same confound D-20260813-17 was ratified about this session — a plainly-written description making the plainer arm win — and the critic caught it in a second design hours later.
2 BLOCKING — P2 largely true by construction. Accepted. P2 restated as is PAULL distinguishable from R25, with the null as the outcome of interest; demoted to secondary and made conditional on G1.
3 BLOCKING — G2 had no threshold; n = 4 neutrals too thin for P4. Accepted. G2 given a numeric bar (+0.25) and an aggregation rule; P4 demoted to exploratory.
4 BLOCKING — G3 would drop a competent seat for disagreeing with the lead's own site labels. Accepted. G3 rebuilt as inter-seat coherence; agreement with the lead's map is now reported as a measurement rather than used as a gate, which is strictly more informative.
5 BLOCKING — the contamination consequences were not carried through into P1/P3, and G1 leans on a contaminated pole. Accepted. P1 narrowed to the clean pair A~BRAK, PAULL demoted to a dependent sensitivity, P3 demoted to exploratory, and the written argument added that R06's contamination makes G1/G2 conservative rather than permissive.
6 NON-BLOCKING — the marked/neutral inventory is lead-theoretic. Accepted in the form finding 4's remedy takes: no independent Danish annotator is reachable, so the seat-majority-vs-lead-map agreement is reported and the limits section carries it.
7 NON-BLOCKING — probe (b) invites scoring expansion as elevation. Accepted. The prompt now says in terms: length is not level; judge the words chosen, not the number of them.
8 NON-BLOCKING — P1 had no effect size, seat pool or aggregation rule. Accepted. Mean-of-seat-means, minimum 2 surviving seats, P1 bar set at > +0.20.
9 NON-BLOCKING — the prompt called the renderings "independent" when some are dependence-flagged. Accepted. The word is removed from the prompt.

One finding was NOT the critic's and is recorded so the amendment list is honest: the site sha256 scramble keys on site_id only, so rewriting the glosses did not change any arm's label assignment, and the blinding is unaffected by the amendment.