Repository path: workshop/experiments/E-20260805f-translated-register/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260805f-translated-register |
| status | frozen |
| created | 2026-08-05 |
| updated | 2026-08-05 |
| senses | naturalness |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-translated-register.md, wiki/goodness-senses.md, wiki/base/anchors/README.md, workshop/translations/quijote-I20/R04-v1/translation.md, workshop/translations/quijote-I20/R06-v1/translation.md, config/models.md |
E-20260805f — is translated English a register of its own?
AMENDED 2026-08-05 after the pre-run critic pass, before measure.py was executed against any
text. Eight amendments A1–A8, all from accepted critic findings, are recorded in critic.md and
marked inline below. The critic returned NEEDS-AMENDMENT, nine findings, five BLOCKING, and all
nine are accepted. A2 is a materials defect the critic's F1 led to and the critic could not see:
The House of Souls reprints The Great God Pan whole, so the Machen original arm was counting the
same prose twice.
Frozen 2026-08-05 (S114) before any statistic in §5 was computed. What existed before the freeze and is declared here: the corpus sizes and the three confound diagnostics of §2.4 (block counts, dialogue share, mean sentence length), which were computed in order to choose the materials and to size the design. No feature matrix, no difference vector, no cosine and no permutation was computed before this file was committed.
1. Question
Does an English literary translation sit at a different distance from the English of its own moment than the same translator's own original English does — and, if it does, do independent translators displace in a shared direction?
naturalness is defined as distance from the target-language norm, and an evaluation invoking it
must name one of three register anchors. All three are built from original English writing;
every object the project scores with them is a translation. Nothing on the shelf establishes
that those are the same kind of object. That is the gap this run measures.
2. Materials
2.1 The four treatment cells — within-translator
Each cell is one hand that both translated literary prose and wrote literary prose in English, within a few years, all public domain, all from Project Gutenberg. Holding the hand fixed removes the author, which is by far the largest source of variation in the feature space used.
| cell | translator | TRANSLATED (source lang, year) | ORIGINAL English (year) | genre |
|---|---|---|---|---|
SMOL |
Tobias Smollett | Don Quixote (ES, 1755) | The Adventures of Roderick Random (1748) | matched — picaresque novel both sides |
MACH |
Arthur Machen | Memoirs of Casanova (FR, 1894) | The Great God Pan (1894), The Three Impostors (1895) (A2: The House of Souls removed — it reprints Pan whole) | mismatched — memoir vs weird fiction |
HEARN |
Lafcadio Hearn | Gautier, One of Cleopatra's Nights (FR, 1882) | Chita (1889) | matched — fantastic tales vs novella |
HAPG |
Isabel F. Hapgood | Hugo, Les Misérables (FR, 1887) | Russian Rambles (1895) | mismatched — novel vs travel essays |
Three source languages, 139 years, four hands, two of the four genre-matched.
2.2 The negative control — original vs original, same hand
If two originals by the same hand differ in the same direction that a translation differs from an original, the direction is not about translating. Both control cells reuse texts from §2.1 by construction: the control asks what the same hand's variation looks like without translation in it.
A1: every control text appears in NO treatment arm. Sharing a baseline arm couples d_ctrl and
d_treat algebraically, so the control could fire or fail for reasons unrelated to translating.
| cell | slot standing in for TRANSLATED | slot standing for ORIGINAL |
|---|---|---|
SMOL-C |
Peregrine Pickle (1751) | Ferdinand Count Fathom (1753) |
MACH-C |
The Hill of Dreams (1907) | Far Off Things (1922) |
2.3 The positive control — dialogue against narration
A register difference that certainly exists in English and that the machinery must recover: within each of the four treatment cells' translated text, quoted speech against narration. If the pipeline cannot find a shared direction there, it cannot find one anywhere, and a null in §5 would be uninterpretable.
2.4 Confound diagnostics, computed before the freeze
Dialogue share is higher in the translation in all four cells (SMOL 44.9 vs 16.9, MACH 29.0 vs 13.7, HEARN 20.9 vs 5.9, HAPG 18.4 vs 13.3, per cent of tokens inside quotation marks). This is a systematic confound running in one direction, and it is the reason the primary analysis is narration-only: quoted spans are removed before blocking. The full-text analysis is reported as a secondary and its disagreement with the primary is a registered failure criterion (FC5).
Mean sentence length: SMOL 41.9 / 36.7, MACH 20.6 / 21.7, HEARN 25.4 / 19.2, HAPG 15.6 / 23.4 — no consistent direction, which is worth recording because it is the variable an armchair account of translationese would reach for first.
2.5 Provenance
Every text is a Project Gutenberg plain-text file fetched 2026-08-05, byte sizes and SHA-256 in
materials/corpus-manifest.json. The files are not committed (23 MB, public domain,
re-fetchable from the manifest's URLs); the extracted analysis blocks are.
3. Procedure
- Strip the Gutenberg header and footer; for the primary, strip quoted spans (
corpus.split_speech, with a 4,000-character guard so an unterminated quotation mark cannot swallow a chapter). - Tokenise to lowercase word forms; cut non-overlapping blocks of 2,000 tokens, never straddling two works; discard each work's trailing remainder.
- Thin each arm evenly to at most 40 blocks so no cell dominates by length.
- Feature space: the 100 most frequent word types across all analysis blocks pooled — the Burrows's Delta convention, which is overwhelmingly function words and is the standard instrument for exactly this kind of register/authorship question.
- Per block, relative frequency of each feature; z-standardise each feature across the four
treatment cells' blocks (Delta standardisation). A7: the feature inventory and the
μ, σare computed ONCE, outside the permutation loop; the null relabels blocks only. The control blocks do not shape the space the control is then measured in. The critic's alternative (standardise on original arms only) is not adopted and is recorded as an unexplored sensitivity. - Per cell, the difference vector
d = mean(z over TRANSLATED blocks) − mean(z over ORIGINAL blocks), in 100 dimensions.
4. The statistic
S = the mean cosine similarity over the six unordered pairs of the four treatment cells'
difference vectors. S is high only if independent hands displace in the same direction; it is
near zero if each hand's translated prose differs from its own original prose idiosyncratically.
Null: within each cell, pool its two arms' blocks and randomly relabel them preserving arm
sizes; recompute every d; recompute S. B = 10,000 permutations, seed 20260805.
P = (1 + #{S_null ≥ S_obs}) / (B + 1).
This null holds the cell structure, the block lengths and the feature space fixed and destroys only the translated/original labelling, which is the thing under test.
5. Predictions, registered
- P1 (primary).
S_obs > 0andP < 0.05. - P2 (the wire to the translation limb). A4: evaluated on the FULL-TEXT SECONDARY analysis, and
only there — the primary strips quoted speech, and second-person address lives overwhelmingly
inside it, so on the primary P2 could "fail" because the signal was deleted. A5, the carrier rule
pre-specified exactly: filter to features whose sign agrees in 4 of 4 cells → rank by
|mean d|→ take up to 15. If that set is empty, P2 FAILS. The prediction: at least one member of{you, your}is in the carrier set and the translations are higher on it.thou/thee/thy/thine/yeare excluded from P2 and reported separately as period-confounded (Smollett 1755 against three 1880s–90s hands). This is registered because the lead's own translator's log named it first:T-quijote-I20-R06-v1D2 andT-quijote-I20-R04-v1R1 record that Spanish makes deference with a third-person honorific used as a pronoun, that English has no such pronoun, and that the translator is therefore forced to choose between plainyouand a period costume. All three source languages here grammaticalise a T/V distinction English lacks. The log was frozen and committed at72acb43and2a5d1cbbefore this design was written. - P3. A8: a tested statistic, not a bare sign.
S₂ = cos(d_SMOL, d_HEARN)with its own permutation null on those two cells (B = 10,000, same seed). P3 holds ifS₂ > 0and itsP < 0.05.
A8, second part — the announceable sentence is bound in advance, whatever the numbers. The strongest claim this run may make is: these four hands, translating from Spanish and French into English between 1755 and 1895, displace their narration-only 100-MFW profile from their own original English in a shared direction. It may not say "translated English is a register", and it may not say the three shelf anchors are the wrong yardstick in general.
6. Failure criteria, registered
- FC1. Any cell with fewer than 10 blocks in either arm is dropped from the primary and the
drop is reported. (Checked before freezing: the smallest arm is
HEARN/original at 12.) - FC2 — the control, and it can withhold the primary. A3: both sides of the threshold are now the
same object.
c̄is the mean of the eight individual pairwise cosinescos(d_ctrl, d_treat)over {2 control cells} × {4 treatment cells} — the same quantitySaverages. (The first version compared a cosine-to-centroid against a mean pairwise cosine, which is systematically larger for agreeing vectors and put the threshold in the wrong place.) Ifc̄ ≥ S_obs, P1 is withheld. Reported but not thresholded: the two control cells' cosine with each other, and the minimum treatment pairwise cosine. - FC3. If narration-stripping removes < 5% of tokens for any text, the stripper failed on that text and the cell is reported as effectively unstripped.
- FC4. If
P ≥ 0.05, no shared direction is claimed and the anchor is not built. This is a real possible outcome and it is a result: it would say the three original-writing register points are not systematically the wrong yardstick. - FC5. If the secondary full-text analysis reverses the sign of
S, the finding is reported as dialogue-dependent and P1 is qualified accordingly. - A6 sensitivity, not a criterion. A leave-
HEARN-outSis reported,HEARNbeing the smallest cell; the primary null is unchanged. - FC6 — the positive control. If §2.3's dialogue-against-narration contrast does not return
P < 0.05with a clearly positiveS, the machinery is not demonstrated to detect a known register difference and no null in §5 may be reported as evidence of absence.
7. What this run cannot do
- It is six hands' worth of prose, four of them translating, and the cells are not a sample of anything. A shared direction across four would be a finding about these four.
- Two of four cells are genre-mismatched, which P3 addresses only partially.
- Function-word profiles are a coarse instrument. A shared direction says that translated prose differs, and what carries it; it does not say the difference is audible to a reader, and nothing here measures whether any reader hears it.
- The critic's three non-translation paths to a high
Sare not excluded by this design and are reported as limits: (a)HAPGandMACHboth contrast long narrative prose against a non-novel original in a parallel way; (b) residual unstripped speech, given the one-directional dialogue asymmetry of §2.4; (c) shared Romance-source pressure — all four cells translate from Spanish or French, so a "translated English" reading and a "Romance-into-English" reading are not separable here at all. - FC6 is a floor on the machinery, not a power calculation. Passing it shows the pipeline can find a register difference that certainly exists; it does not show the design is powered against the primary's much smaller between-work effect.
- Nothing here is a jury verdict. No panel model is called for judgment,
config/models.mdreads NOT CALIBRATED, and Tier D remains NOT PASSED. Every evaluative sentence downstream carriesinternal-judgment-onlyandprovisional. - The lead's own rendering (
T-quijote-I20-R04-v1) is not an analysis block and enters no statistic. It is 830 words, far below one block, and its function-word rates would be noise. Its role is the translator's log, which is where P2 came from.
8. Cost
$0 for the measurement — it is arithmetic over public-domain text, standard library only. The only API call is the independent pre-run critic pass (charter §8), a non-juror seat; there are no jurors in this run.