Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260805f-translated-register/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260805f-translated-register
statusfrozen
created2026-08-05
updated2026-08-05
sensesnaturalness
internal-judgment-onlytrue
provisionaltrue
linkswiki/arms/ARM-translated-register.md, wiki/goodness-senses.md, wiki/base/anchors/README.md, workshop/translations/quijote-I20/R04-v1/translation.md, workshop/translations/quijote-I20/R06-v1/translation.md, config/models.md

E-20260805f — is translated English a register of its own?

AMENDED 2026-08-05 after the pre-run critic pass, before measure.py was executed against any text. Eight amendments A1–A8, all from accepted critic findings, are recorded in critic.md and marked inline below. The critic returned NEEDS-AMENDMENT, nine findings, five BLOCKING, and all nine are accepted. A2 is a materials defect the critic's F1 led to and the critic could not see: The House of Souls reprints The Great God Pan whole, so the Machen original arm was counting the same prose twice.

Frozen 2026-08-05 (S114) before any statistic in §5 was computed. What existed before the freeze and is declared here: the corpus sizes and the three confound diagnostics of §2.4 (block counts, dialogue share, mean sentence length), which were computed in order to choose the materials and to size the design. No feature matrix, no difference vector, no cosine and no permutation was computed before this file was committed.

1. Question

Does an English literary translation sit at a different distance from the English of its own moment than the same translator's own original English does — and, if it does, do independent translators displace in a shared direction?

naturalness is defined as distance from the target-language norm, and an evaluation invoking it must name one of three register anchors. All three are built from original English writing; every object the project scores with them is a translation. Nothing on the shelf establishes that those are the same kind of object. That is the gap this run measures.

2. Materials

2.1 The four treatment cells — within-translator

Each cell is one hand that both translated literary prose and wrote literary prose in English, within a few years, all public domain, all from Project Gutenberg. Holding the hand fixed removes the author, which is by far the largest source of variation in the feature space used.

cell translator TRANSLATED (source lang, year) ORIGINAL English (year) genre
SMOL Tobias Smollett Don Quixote (ES, 1755) The Adventures of Roderick Random (1748) matched — picaresque novel both sides
MACH Arthur Machen Memoirs of Casanova (FR, 1894) The Great God Pan (1894), The Three Impostors (1895) (A2: The House of Souls removed — it reprints Pan whole) mismatched — memoir vs weird fiction
HEARN Lafcadio Hearn Gautier, One of Cleopatra's Nights (FR, 1882) Chita (1889) matched — fantastic tales vs novella
HAPG Isabel F. Hapgood Hugo, Les Misérables (FR, 1887) Russian Rambles (1895) mismatched — novel vs travel essays

Three source languages, 139 years, four hands, two of the four genre-matched.

2.2 The negative control — original vs original, same hand

If two originals by the same hand differ in the same direction that a translation differs from an original, the direction is not about translating. Both control cells reuse texts from §2.1 by construction: the control asks what the same hand's variation looks like without translation in it.

A1: every control text appears in NO treatment arm. Sharing a baseline arm couples d_ctrl and d_treat algebraically, so the control could fire or fail for reasons unrelated to translating.

cell slot standing in for TRANSLATED slot standing for ORIGINAL
SMOL-C Peregrine Pickle (1751) Ferdinand Count Fathom (1753)
MACH-C The Hill of Dreams (1907) Far Off Things (1922)

2.3 The positive control — dialogue against narration

A register difference that certainly exists in English and that the machinery must recover: within each of the four treatment cells' translated text, quoted speech against narration. If the pipeline cannot find a shared direction there, it cannot find one anywhere, and a null in §5 would be uninterpretable.

2.4 Confound diagnostics, computed before the freeze

Dialogue share is higher in the translation in all four cells (SMOL 44.9 vs 16.9, MACH 29.0 vs 13.7, HEARN 20.9 vs 5.9, HAPG 18.4 vs 13.3, per cent of tokens inside quotation marks). This is a systematic confound running in one direction, and it is the reason the primary analysis is narration-only: quoted spans are removed before blocking. The full-text analysis is reported as a secondary and its disagreement with the primary is a registered failure criterion (FC5).

Mean sentence length: SMOL 41.9 / 36.7, MACH 20.6 / 21.7, HEARN 25.4 / 19.2, HAPG 15.6 / 23.4 — no consistent direction, which is worth recording because it is the variable an armchair account of translationese would reach for first.

2.5 Provenance

Every text is a Project Gutenberg plain-text file fetched 2026-08-05, byte sizes and SHA-256 in materials/corpus-manifest.json. The files are not committed (23 MB, public domain, re-fetchable from the manifest's URLs); the extracted analysis blocks are.

3. Procedure

  1. Strip the Gutenberg header and footer; for the primary, strip quoted spans (corpus.split_speech, with a 4,000-character guard so an unterminated quotation mark cannot swallow a chapter).
  2. Tokenise to lowercase word forms; cut non-overlapping blocks of 2,000 tokens, never straddling two works; discard each work's trailing remainder.
  3. Thin each arm evenly to at most 40 blocks so no cell dominates by length.
  4. Feature space: the 100 most frequent word types across all analysis blocks pooled — the Burrows's Delta convention, which is overwhelmingly function words and is the standard instrument for exactly this kind of register/authorship question.
  5. Per block, relative frequency of each feature; z-standardise each feature across the four treatment cells' blocks (Delta standardisation). A7: the feature inventory and the μ, σ are computed ONCE, outside the permutation loop; the null relabels blocks only. The control blocks do not shape the space the control is then measured in. The critic's alternative (standardise on original arms only) is not adopted and is recorded as an unexplored sensitivity.
  6. Per cell, the difference vector d = mean(z over TRANSLATED blocks) − mean(z over ORIGINAL blocks), in 100 dimensions.

4. The statistic

S = the mean cosine similarity over the six unordered pairs of the four treatment cells' difference vectors. S is high only if independent hands displace in the same direction; it is near zero if each hand's translated prose differs from its own original prose idiosyncratically.

Null: within each cell, pool its two arms' blocks and randomly relabel them preserving arm sizes; recompute every d; recompute S. B = 10,000 permutations, seed 20260805. P = (1 + #{S_null ≥ S_obs}) / (B + 1).

This null holds the cell structure, the block lengths and the feature space fixed and destroys only the translated/original labelling, which is the thing under test.

5. Predictions, registered

A8, second part — the announceable sentence is bound in advance, whatever the numbers. The strongest claim this run may make is: these four hands, translating from Spanish and French into English between 1755 and 1895, displace their narration-only 100-MFW profile from their own original English in a shared direction. It may not say "translated English is a register", and it may not say the three shelf anchors are the wrong yardstick in general.

6. Failure criteria, registered

7. What this run cannot do

8. Cost

$0 for the measurement — it is arithmetic over public-domain text, standard library only. The only API call is the independent pre-run critic pass (charter §8), a non-juror seat; there are no jurors in this run.