Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: wiki/findings/results/RS-20260806c-source-grammar.md · rendered 2026-09-09

Page metadata (front matter)
typeresult
idRS-20260806c-source-grammar
statusresolved
created2026-08-06
updated2026-08-06
sensesnaturalness
internal-judgment-onlytrue
provisionaltrue
linksworkshop/experiments/E-20260806c-source-grammar/design.md, workshop/experiments/E-20260806c-source-grammar/critic.md, wiki/arms/ARM-translated-register.md, wiki/findings/results/RS-20260805f-translated-register.md, wiki/goodness-senses.md, wiki/base/anchors/README.md, workshop/translations/immensee-elisabeth/R06-v1/translation.md, workshop/translations/immensee-elisabeth/R04-v1/translation.md, workshop/experiments/E-20260806c-source-grammar/materials/census-quijote-es.md, config/models.md

The second person replicates across three source families, and the null it was meant to explain does not

E-20260806c-source-grammar, ARM-translated-register step 2, the arm's closing step. S119, 2026-08-06. $0.0418704, one API call — the pre-run critic. The measurement is arithmetic over public-domain text and cost nothing; the translation is the lead's own and is never ledgered.

Verification: 227 checks, 0 failures, four mutation tests, all four caught, by a verifier that reimplements the tokenisation, the speech split, the blocking, the thinning, the rates, the pooled statistics, the exact decomposition and both census parsers, and imports nothing from the measurement. Key-usage delta matches the per-request cost to 0.000000000. Pre-run critic NEEDS-AMENDMENT, ten findings, six BLOCKING, all ten accepted, nine amendments applied before measure.py was written.

internal-judgment-only and provisional: Tier D is NOT PASSED, config/models.md reads NOT CALIBRATED. No jury was convened and nothing was judged. No sense was scored.

1. What the arm asked, and the short answer

RS-20260805f found no profile-level translated-English register in four hands translating from Spanish and French, and one thing that did move: you and your, in all four hands, at P = 0.0001. Step 2 asked whether a naturalness anchor can hold a feature-level displacement with no register behind it — which first required ruling in or out the two alternatives that page's own limits named: the source family (all four cells were Romance) and dialogue share (you and your are address words, and the figure came from the full-text analysis).

Both alternatives fail, and the displacement is real. Across nine within-translator cells spanning German, Russian, Italian, Spanish and French sources, an English literary translation carries more second-person forms than the same translator's own English: positive in 8 of 9 cells, and on the four cells that are both non-Romance and independent of each other, pooled +1.82 per 1,000 tokens precision-weighted (P = 0.0011) and +5.90 unweighted (P = 0.0001). It is not carried by dialogue volume, it survives with quoted speech stripped out, and a declared post-hoc probe finds it in the three cells that set fiction against fiction as well.

And the finding the arm was built on does not survive. P4 predicted the profile-level null would replicate. It does not: on the German and Russian cells S = 0.321, P = 0.0001, where the four Romance cells gave −0.0021 at P = 0.497. "There is no translated-English register" was a statement about four Romance cells and it does not generalise.

The run's own positive control missed its bar — 0.4813 against a declared ≥ 0.50 — and that is reported in §5 rather than smoothed over.

2. Design in one paragraph

Five new within-translator cells, all public domain, every translator attribution read off the Project Gutenberg file's own header: CARL (Goethe's Wilhelm Meister tr. Carlyle 1824/1827 against Sartor Resartus), DUFF (Meinhold's Amber Witch tr. Lady Duff Gordon 1844 against Letters from Egypt), HAPG-RU (Tolstoy's Sevastopol tr. Hapgood against Russian Rambles), DOLE-RU (two Tolstoy volumes tr. Dole against The Spell of Switzerland) and DOLE-ROM (Verga and Valdés tr. Dole against the same Spell of Switzerland). The last two, and HAPG-RU against RS-20260805f's HAPG, hold the hand and the original-arm baseline identical and vary only the source language — the tightest contrast the question admits without commissioning translations. Corpus 1 is re-declared unchanged and asserted by SHA-256 against its frozen manifest. Displacement is the difference in rate per 1,000 tokens between the translated arm's 2,000-token blocks and the original arm's, thinned to 40 per arm; the null relabels blocks within each cell, 10,000 times.

3. P1 — the replication, and the three ways it could have been an artifact

cell family source {you, your} Δ, full Δ, narration only possessive-class Δ share fraction
SMOL Romance ES +4.91 +4.94 −13.56 1.65
MACH Romance FR +3.09 −6.45 +15.29 −1.94
HEARN Romance FR +3.79 −0.57 +5.85 0.71
HAPG Romance FR +4.45 +0.79 −0.34 0.24
CARL Germanic DE +4.21 +0.80 +17.01 −0.12
DUFF Germanic DE −2.10 −3.58 +17.16 0.31
HAPG-RU Slavic RU +11.31 +7.34 −0.11 0.21
DOLE-RU Slavic RU +10.17 +4.55 +12.31 0.04
DOLE-ROM Romance IT+ES +13.55 +3.73 +15.47 0.05

(Bold: the four cells P1 is stated over. Rates per 1,000 tokens. FC3: DOLE-ROM shares its original arm with DOLE-RU and HAPG-RU with HAPG, so those pairs are never counted as independent.)

Pooled over the four bold cells: precision-weighted +1.825, P = 0.0011; unweighted +5.901, P = 0.0001; positive in 3 of 4. All four clear the ≥ 10-blocks-per-arm floor A6 imposed. P1 holds.

Three gates could have taken it away and none fired:

3.1 A declared post-hoc probe on the genre alternative, because FC8 only tests one cell

The critic's F5 is right that no cell in corpus 2 pairs narrative fiction with narrative fiction, and you/your are address features that novels carry more of than travel books. Corpus 1 has three cells whose original arm is narrative fiction. Splitting all nine that way:

original arm cells mean {you, your} Δ positive
narrative fiction SMOL, MACH, HEARN +3.93 3 of 3
non-fiction / non-narrative HAPG, CARL, DUFF, HAPG-RU, DOLE-RU, DOLE-ROM +6.93 5 of 6

The displacement is present in every fiction-against-fiction cell. It is larger where the original arm is non-fiction, which is what a genre contribution would look like, but it does not depend on one. Labelled a probe; the primary was not re-decided on it, and FC3 forbids reading the second row as 5 of 6 independent cells.

4. P4 fails, and that is the run's largest result

P4 predicted that the profile-level statistic would again be indistinguishable from its null on the new cells. It is not.

cells S P null 95th pct |S|
RS-20260805f, four Romance cells SMOL MACH HEARN HAPG −0.0021 0.497 —
this run, four non-Romance cells CARL DUFF HAPG-RU DOLE-RU +0.321 0.0001 0.113

Same recipe, same block length, same 100-MFW Delta-z construction, same permutation null, narration only. The declared minimum-detectable-effect floor (A4) is met: the null's 95th percentile is 0.113, below the 0.15 floor, so the design could have called an effect of that size and this is not a power artifact in either direction.

What this licenses and what it does not. It licenses: four hands translating German and Russian literary prose into English displace their narration-only function-word profile from their own original English in a shared direction. It does not license "translated English is a register" — S = 0.321 is well under the 0.739 that separates speech from narration in these very texts, and the second declared probe shows where much of it comes from.

4.1 A declared post-hoc probe on where the shared direction lives

Pairwise cosines of the four difference vectors, with RS-20260805f's HAPG added — a Romance cell whose original arm is also travel prose:

DUFF HAPG-RU DOLE-RU HAPG
CARL 0.443 0.290 0.274 0.121
DUFF — 0.059 0.297 0.048
HAPG-RU — 0.591 0.584
DOLE-RU — 0.258

The two largest non-coupled values are the two within-family pairs: CARL~DUFF (both German, different authors) at 0.443 and HAPG-RU~DOLE-RU at 0.591. (HAPG~HAPG-RU, italicised, shares an original arm and is algebraically coupled — it is not evidence of anything.) But HAPG-RU and DOLE-RU both translate Tolstoy, so the highest informative cosine in the table is confounded between Russian and the same source author, and this design cannot separate them. Adding HAPG gives S = 0.297 at P = 0.0005 over five cells; the Romance cell aligns with the German ones at 0.121 and 0.048, which is weak.

5. The positive control missed its bar

FC1 required speech against narration, inside corpus 2's translated arms, to return S ≥ 0.50 at P < 0.01. It returned S = 0.4813 at P = 0.0001 — highly significant, and 0.019 below the threshold the design set before the run. HAPG-RU was dropped from the control by FC2 for having only 5 speech blocks against a floor of 6.

The gate's stated consequence is that every null in this run is withheld. This run reports no null: P1 is a positive result, P4's registered null failed and the observed statistic is positive, and FC4/FC4b are an exact algebraic identity and a sign comparison, neither of which is a permutation null. So nothing is withheld, and the reason is that there was nothing of that kind to withhold — not that the control passed. It did not pass. A reader who wants the conservative reading should note that the direction of the miss cuts against over-reading rather than for it: an instrument shown to be less sensitive than required, which then finds an effect, is not made weaker by its under-sensitivity. Had P4 come back null, that null would have been unreportable.

6. The mechanism: the census predicted it, the group orderings agree, and the one test with force splits

The translation limb produced two censuses (§7). From them the design registered, as exploratory hypotheses with no arm-closing power (A1, on the critic's F1):

prediction observed leave-one-out
P2a pro-drop sources displace {you, your} more than German +9.23 vs +1.06 — holds no flip
P2b German sources displace the possessive class more than Romance +17.09 vs +4.54 — holds no flip
P2b⁻ (the form frozen in the R06 log) German below Romance refuted —

The prediction frozen inside the translator's log was wrong, and the thing that overturned it was a second census done before any corpus number existed. The R06 log reasoned from German alone and guessed that Romance would supply possessives more freely; coding the lead's frozen Spanish translation showed the opposite, the design registered the reversed form as P2b, and the corpus agrees with the census rather than with the log.

But the only source-family test that keeps its force fails. P3 compares cells with the hand and the original-arm baseline identical:

P3 does not hold, and P2a/P2b cannot stand in for it: they compare different books by different people, which is the confound P3 exists to remove. The honest state of the mechanism question is open.

7. The translation limb, and the census that reversed a frozen prediction

Storm's Immensee, chapter «Elisabeth» — 1,260 German words, 33 paragraphs, the novella's final meeting — rendered R06 (draft, frozen) then R04 (self-revision): T-immensee-elisabeth-R06-v1, T-immensee-elisabeth-R04-v1. Paragraphing 33 to 33.

The R06 log carries a census of every {you, your, my, his, her, its, our, their} token in the draft, coded against the German; materials/census-quijote-es.md applies it to the lead's frozen Spanish rendering from S114. Both parsed by analysis/census.py, both re-parsed by the verifier:

second person supplied possessive class supplied
German — Storm 0 of 5 23 of 44
Spanish — Cervantes 6 of 14 1 of 9

The site of obligatory supply is opposite. Spanish drops the subject pronoun, so you has to be put in; Spanish carries its possessives overtly (mi pretensión, su amo, sus ojos) and they are merely copied. German never drops the subject; German takes the bare article on body parts and belongings (den Kopf, die Augen, um den Hals, in die Hand), and the possessive is supplied at half the sites.

Contamination measured, not asserted, against the one freely reachable published English (PG 6650), after both artifacts were frozen:

pair shared 7-grams 12-grams 15-grams longest run
lead R04 ~ PG 6650 60 3 0 13
lead R06 ~ PG 6650 61 3 0 13

The 13-token run is "the garden room he did not go in to him he stood still". No published–published baseline exists: only one English Immensee is freely reachable, so this figure has nothing to be compared against on this locus, and is reported without one. contamination: suspected stands. The lead's prose enters no statistic in this run — its role is the log and the census.

8. What the arm does with this — the anchor decision

ARM-translated-register's Done when asks whether a Tier 1 anchor for the translated-English register can be built, or must be declared unbuildable. No anchor is built, and the reason has changed.

RS-20260805f closed step 1 believing there was nothing for a fourth anchor to point at. There is something: a stable, replicated, dialogue-independent displacement in two closed word classes, and — on German and Russian sources — a profile-level shared direction as well. What there is not is something an anchor could be. The three register anchors on the shelf are texts that an evaluation names and reads a translation against; this is a rate per 1,000 tokens on function words, and no evaluation procedure in this project consults one. An anchor that could hold it would have to be a different kind of object than the shelf's other three, and nothing in the framework would call it.

So: naturalness keeps its three original-writing register points, and gains a written note saying that the translated/original mismatch was measured twice, that the mismatch is real at the level of individual forms and of the German/Russian profile, and that it is not currently something the sense can select. The mechanism question — whether the displacement is the source language's grammar showing through — is left open and named, per A2's separation of the anchor decision from the scientific claim.

9. A correction to the predecessor, per amendment A7

E-20260805f design §2.4 and RS-20260805f §2 state that dialogue share is higher in the translation in all four cells, giving MACH as 29.0 against 13.7. Recomputed per work — which is how that analysis blocks its texts, one work at a time — Machen's original arm is 37.5 (Pan 43.0, The Three Impostors 35.5) and the asymmetry in that cell runs the other way. Concatenating the two original works and choosing a single quotation convention for the pair gives 25.9, under which the published direction holds; the analysis-consistent computation is per work.

The sentence is false as published. The critic was right that calling it "two ways of computing the same diagnostic" is policy rather than epistemology. No statistic in RS-20260805f changes: the diagnostic was used to justify choosing narration-only as the primary analysis, and that choice does not require the confound to run one way. The result page is not reopened; the correction lives here and on the arm page.

10. Limits