Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260810-legend-lexis/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260810-legend-lexis
statusfrozen
created2026-08-10
updated2026-08-10
sensesstyle-correspondence, voice
internal-judgment-onlytrue
provisionaltrue
linkswiki/arms/ARM-legend.md, workshop/translations/szent-peter-esernyoje/R05-v1/translation.md, workshop/translations/szent-peter-esernyoje/register.md, wiki/findings/results/RS-20260809e-legend-layer.md, workshop/experiments/E-20260809e-legend-layer/design.md, config/models.md

E-20260810-legend-lexis — does Mikszáth's archaic layer carry Károli's Bible, or only Károli's verb endings?

Frozen 2026-08-10 (S149) before any statistic on the split was computed. The translation limb it is wired to was frozen first, at commit f7e9ac8.

0. The wire, in one sentence

Span E rendered the false miracle at ¶240–¶244 in King James English on register rule V12's premise that Mikszáth's archaic layer is carried by scriptural lexis and syntax; this experiment tests that premise on the Hungarian, and the translation's own erratum 1 — which would flatten span D's ¶162–¶164 and span E's ¶240–¶244 together — is the declared price of a null (D47).

1. Question, and why it is a question about the novel

RS-20260809e (S144) demoted V12 and required span E to re-decide it. What that experiment refuted was a distributional claim — that the marked verb morphology clusters into paragraphs — at p = 0.1253, with 40 of 55 marked tokens immovable under the null, i.e. substantially a power failure.

That is not what V12 asserts. V12 is a claim about a carrier: that the layer is marked by scriptural lexis and syntax rather than by bent verb morphology, and that English should therefore reach for the King James Bible rather than for quoth and -eth. That claim has never been tested, and it is testable, because Hungarian has a received biblical register with one dominant source — Károli Gáspár's Bible — and the forms Mikszáth reaches for are in it: járulának (7 occurrences in Károli), jövének (97), vagynak (1), hozának (19), rendelé (165), elveszté (15), lőn (526).

The question. Given a narration paragraph of this novel, does the presence of marked verb morphology predict Károli-Bible phrasing in the rest of that paragraph — the marked forms themselves masked out?

What this teaches about translating literature (subject rule): it decides which English register a translator should reach for at a specific set of paragraphs in a specific novel, and it decides it against the source rather than against the translator's ear. It is not about this project's instruments, statistics or raters.

2. Materials, all free, all already reachable

id text provenance size
N «Szent Péter esernyője», whole novel, Révai 1910 PG #68911, already segmented into 2,133 paragraphs with part and configuration labels at E-20260809e/runs/paragraphs.tsv 53,470 words
C the S144 occurrence census E-20260809e/runs/occurrences.tsv — 992 past-tense tokens over 76 lemmas, each labelled MARKED/UNMARKED and NARRATION/TAG/MIXED 181 MARKED
K Károli Gáspár, Szent Biblia MEK-00161, 929 chapter files, fetched this session 485,318 tokens after normalisation
J Jókai Mór, five works — «A jövő század regénye» I, «Erdély aranykora», «Megtörtént regék», «A kis királyok» I, «Életemből» II PG #55911, #56189, #57029, #57233, #64502 495,519 tokens after normalisation

J is the secular period control and it is the reason this experiment can say anything about the Bible rather than about old formal Hungarian. Jókai is the ambient literary Hungarian Mikszáth wrote inside; if the marked paragraphs are elevated against K and equally against J, the effect is register-general and V12's King James claim is not licensed.

Declared limit on K. MEK-00161 gives no revision year. The digital Károli in circulation is almost certainly the 1908 revision, which postdates the novel by thirteen years; Mikszáth's readers knew an earlier one. The morphology and the clause patterns are continuous across the revisions, but individual verse wordings are not. The direction of this error is conservative: a revision that changes wording reduces n-gram overlap and biases the test toward the null.

Declared limit on N. One witness (R). No page image. collation.md §6, written this session, records two ⚑ rows that were false; neither falls in the census's material.

3. Instrument

Normalisation (identical for N, K and J, and the reason it is needed is that the three texts are in three orthographies — R has cz and 1910 vowel length, PG-Jókai has cz, MEK-Károli is modernised): lowercase → cz → c → fold vowel length only (á→a é→e í→i ó→o ő→ö ú→u ű→ü, which preserves the o/ö and u/ü distinctions because those are different vowels) → tokens are maximal runs of [a-zöüäëß].

Masking — the control that makes the test non-circular. The census's 80 accepted MARKED surface types are deleted from N, from K and from J alike, and no bigram is allowed to span a deleted position (the token stream is cut there). A marked paragraph therefore cannot score by containing lőn, and Károli cannot supply a match through its own 526 lőns. Masking is applied to all MARKED types including TAG ones, which is the conservative choice.

Score. For a paragraph p and comparator X:

overlap(p, X) = |bigrams(p) ∩ bigrams(X)| / |bigrams(p)|

where bigrams are adjacent normalised token pairs after masking, taken as sets (a bigram counts once per paragraph however often it repeats).

Split-blind calibration already run, and reported here because it was run before the split: over the 387 NARRATION paragraphs of ≥30 words, overlap(·, K) has mean 0.2113, median 0.2045, range 0.0444–0.5429, and no paragraph scores zero. The measure is not degenerate at this paragraph length. No contrast was computed.

4. Population and split

The 30-word floor is registered as primary; 20 and 40 are registered sensitivities (MARKED n = 39 and 28). Short paragraphs give the ratio a noisy denominator and lose proportionally more bigrams to masking.

5. Strata — the two confounds, controlled by construction

The permutation shuffles the MARKED label within strata, never across the whole population. Strata = length quartile × religious-topic indicator, computed on the population before the split is used.

  1. Length. Quartiles of paragraph word count within the population.
  2. Religious topic — the confound that would otherwise decide the result. Paragraphs about God, church, death and burial resemble the Bible whatever their verb morphology, and the marked paragraphs are disproportionately about exactly that. The indicator is 1 if the paragraph contains ≥ 1 token matching the declared lexicon below.

The lexicon, declared here in full and built blind to the split, by prefix on normalised tokens, with the collisions found by inspecting the novel's own vocabulary and excluded by name:

isten* szent* templom* egyhaz* ima imad* angyal* oltar* koporso* temet* halott*
mise* gyon* aldas* predikal* jezus* krisztus pokol* üdvöz* zsolozsma* ereklye*
apostol* plebanos* püspök* kapolna* sekrestye* harang* vallasos bün*
pap*   EXCEPT papir* papa* papu* papr* papl*
menny* EXCEPT mennyi* mennyire mennyit mennyivel mennyiert mennyiseg*
kereszt* EXCEPT keresztül*
csoda csodak csodat csodaja csodatetel* csodatevö*   (miracle)
       EXCEPT csodalkoz* csodalatos* csodalattal     (astonishment — and `csodálkozék` is itself a MARKED form)

The csoda split is semantic and deliberate: miracle is topic, astonishment is not, and one of the astonishment forms is in the masked set, so including it would let the indicator proxy the outcome.

6. Predictions, registered

P1 — primary. mean overlap(·, K) | MARKED > mean overlap(·, K) | UNMARKED, one-sided, by within-stratum permutation of the MARKED label, 10,000 permutations, α = 0.05.

P2 — specificity. The standardised effect (difference ÷ pooled SD) against K exceeds the standardised effect against J. Absolute overlap levels are not comparable across comparators — K has 230,167 distinct masked bigrams and J has 336,186, so J matches more of everything — which is why P2 is stated on standardised within-comparator contrasts and why a density-matched sensitivity (J's bigram set deterministically subsampled to K's size, seed 20260810) is registered alongside it.

P3 — locality. P1 survives with Part I excluded (MARKED n = 23). This is RS-20260809e's amendment A9, imported into the design rather than added by a critic, because Part I is the passage that generated V12 and the last experiment's whole result turned on it.

7. Power floors, checked before freezing

MARKED n ≥ 25 in the primary (33, met) and ≥ 15 outside Part I (23, met). If a floor had failed the test would not have been run and this page would say so.

8. Failure criteria and the decision table — written before the run

outcome V12 the translation
P1 fails struck erratum 1 flattens span D ¶162–¶164 and span E ¶240–¶244
P1 holds, P3 fails struck — the effect is again inside the passage that generated it erratum 1, same scope
P1 and P3 hold, P2 fails re-founded in weakened form: the layer is marked and English should mark it, but the King James is not licensed as the carrier — the effect is register-general no erratum; V12 is rewritten to say what was and was not shown
P1, P2 and P3 all hold re-founded as a lexical rule with a measurement behind it no erratum; span D's and span E's marking stand

This table is the point of the design. The last time V12 was tested, the consequence was argued after the number arrived. Here the consequence is fixed first, in a frozen artifact, and the translation that will pay for it is already committed.

9. What this cannot show, stated before the run

10. Procedure

  1. Design frozen and committed (this file), translation already frozen at f7e9ac8.
  2. Independent pre-run critic pass, one non-Anthropic panel model, on this page and the built lexicon. Findings recorded with accept/overrule in §11 before any contrast is computed.
  3. run.py — builds the comparators, masks, scores, permutes; raw outputs to runs/.
  4. verify.py — an independent path that rebuilds the scores from the stored paragraph text, re-implements the permutation by index sampling rather than label shuffling, and runs mutation tests: (i) drop a masked type, (ii) mask random UNMARKED tokens in UNMARKED paragraphs at the MARKED arm's rate — the masking-artefact control — (iii) collapse the strata, (iv) truncate the comparator. Each must move the statistic in the predicted direction.
  5. Result page; register and translation updated per §8's table; only then is Worswick opened for the contamination measurement and D44a's prediction.

11. Critic findings and amendments

Pre-run critic: openai/gpt-5.6-terra (P1), one call, $0.0293675, provider OpenAI, finish_reason: stop. Verdict NEEDS-REDESIGN, 12 findings (6 BLOCKING, 5 SERIOUS, 1 MINOR). Raw at runs/critic-out.md. Eleven accepted, one accepted-in-part with the overrule written. No contrast had been computed when the pass was made and none was computed until every amendment below was implemented. Sections 3–8 above are superseded where an amendment says so.

A1 (finding 1, BLOCKING — ACCEPTED). The mask becomes status-independent, and the population becomes the paragraph-level opportunity set. The critic is right that masking only the MARKED forms cuts only MARKED paragraphs, so the two arms' residual bigram sets are not comparable by construction. Replaced by: every census past-tense token type — MARKED and UNMARKED, 134 types in union — is deleted from N and from every comparator, so the deletion rule cannot see the arm. The population is restricted to narration paragraphs holding ≥ 1 census past-tense token, which is the paragraph-level form of RS-20260809e's own opportunity set and is the "matched sets with actual overlap" finding 2 asks for. New population: 237 (MARKED 33, UNMARKED 204), at ≥30 words. The mask removes 1.9% of the novel's tokens and 1.6% of Károli's.

A consequence worth recording: the union mask deletes monda and kérde from all corpora — the homographic 3sg indefinite slot that RS-20260809e §2 had to exclude. It is now out of the score entirely. It is still not out of the arm assignment (finding 11).

A2 (finding 1 + 2 — ACCEPTED). Deletion count becomes a stratification variable and the balance table is published. MARKED paragraphs still lose more tokens on average than UNMARKED ones, so paragraphs are compared only within the same deletion band. Bands: 1 and ≥2. Observed: MARKED 10 / 23, UNMARKED 120 / 84 — both bands populated in both arms.

A3 (finding 2, BLOCKING — ACCEPTED IN PART; the overrule is written). The full stratum table, including movable-label counts, is published before the primary is read. What is overruled: the demand to add part, chapter and a rich topic model as stratification variables. With 33 MARKED paragraphs, a five-level part factor crossed with the rest produces cells of size 0 and 1 and destroys the very movable-label count the same finding asks to see. Part enters as P3's heterogeneity analysis instead, which is where it belongs. Strata are therefore deletion band (2) × religious-rate tertile (3) = 6, and the result states plainly that the null is exchangeability given those two variables and nothing else, and that narrative content, character, and local sequence are not controlled. Length is handled by a registered sensitivity (middle length tertile only; and overlap residualised on log word count) rather than by a seventh and eighth cell.

A4 (finding 3, BLOCKING — ACCEPTED IN PART; the overrule is written). The topic measure becomes continuous and the lexicon is extended. What is overruled: blinded human coders assigning multi-label scene annotations. This project has no human coders and cannot obtain them; the amendment is not available at any price, and pretending otherwise would be worse than saying so. What is accepted: (i) the indicator becomes a continuous religious-term rate per 100 tokens, tertiled, replacing the binary flag; (ii) the lexicon is extended with the classes the critic named and that occur in this novel — ördög* pokol* prof* tanitvany* evangel* iras szentiras paradicsom* feltamad* gyontat* atok* megatkoz* özvegy* arva* barat* kolostor* zarand* bucsu* mennyorszag* kalvaria* mennybolt* — each checked against the novel's own vocabulary for collisions before use; (iii) the residual-confounding limit is stated in the result as a primary limit, not a footnote.

A5 (finding 4, BLOCKING — ACCEPTED IN FULL, and it is the best finding). P2 is redefined as a paired interaction with a real test. The old P2 — "the standardised effect against K exceeds the standardised effect against J" — had no null, no α and no treatment of the dependence between two scores from the same paragraph. Replaced by: per paragraph d = overlap(p,K) − overlap(p,H); statistic mean(d | MARKED) − mean(d | UNMARKED), tested by the same within-stratum permutation as P1, one-sided, α = 0.05, with a 95% permutation interval. The comparators' different bigram densities shift d by a constant and cancel in a difference of means, which also disposes of the subsampling kludge the old design needed.

A6 (finding 5, BLOCKING — ACCEPTED). A second, sharper secular control. One Jókai collection cannot separate Bible from old formal Hungarian. Added: H, an archaizing-secular pool — Kemény Zsigmond «Zord idő» I–II and «A rajongók» I, Jósika «A csehek Magyarországban» II, Gárdonyi «A láthatatlan ember» (PG #69380, #69381, #68979, #67794, #57477), 379,028 tokens, five historical novels whose whole business is elevated archaic Hungarian that is not scripture. All three comparators are systematically subsampled to 379,028 tokens so the levels are comparable. P2's primary control is H, not J — J is reported alongside. What is narrowed: with three comparators there is a range of secular contrasts, not a distribution, and the result says so.

A7 (finding 6 + 7, SERIOUS — ACCEPTED). A type-by-type mask audit is published, with tokens removed from N, K, J and H for each of the 134 types, ambiguous types flagged by name. The claim that masking's effect on the comparator contrast is "simply conservative" is withdrawn; the audit is reported and the direction is left open.

A8 (finding 8, SERIOUS — ACCEPTED IN FULL). The "conservative" claim about the 1908 revision is struck. §2's sentence "The direction of this error is conservative" is withdrawn: a revision can move overlap either way, and it touches K and not the secular controls, so it can move P1 and P2 in either direction. No pre-1897 Károli witness is reachable to this project; the conclusion is limited to this digital Károli and says so.

A9 (finding 9, BLOCKING — ACCEPTED IN FULL, and it overturns the design's own centrepiece). §8's table treated non-rejection as refutation and let a p-value above 0.05 delete two spans of frozen translation. That is invalid and the critic is right. §8 is replaced by §12 below. Two changes: (i) a null is recorded as failure to found, never as refutation, and V12 — which is already demoted — simply stays demoted, with the register recording that two designs have now failed to found it; (ii) a population-level association cannot prescribe a change to two particular paragraphs, so a span-level descriptive analysis is added — pid 163, 164 (span D's ¶162–¶163) and pid 241, 242, 245 (span E's ¶240, ¶241, ¶244) reported with their own scores and percentiles, labelled descriptive and not a test. The translator's frozen log at D47 declared that a null would force erratum 1. That declaration was itself a statistical error, and it is superseded here rather than acted on. The log is not edited (V2); the supersession is recorded in the register, in this design and in the result.

A10 (finding 10, SERIOUS — ACCEPTED). P3 is recast as heterogeneity, not replication. Part-I and non-Part-I effects are reported with permutation intervals plus an explicit interaction test. The phrase "the effect is again inside the passage that generated it" is withdrawn as a conclusion a subset analysis can support; it becomes a description of where the estimate sits.

A11 (finding 11, SERIOUS — ACCEPTED). The directional claim about the undetectable 3sg slot is withdrawn, and replaced by a measurement: every UNMARKED narration paragraph in the population is searched for candidate archaic 3sg indefinite forms from a declared list, and the contamination count is reported. No direction is asserted without it.

A12 (finding 12, MINOR — ACCEPTED). Mutation tests become diagnostics, each with a pre-declared expected failure mode and a reported magnitude, and none of them may alter the primary conclusion.

12. Decision table, replacing §8 (amendment A9)

outcome what it licenses about V12 what happens to the translation
P1 founds it (p < 0.05, positive estimate) and P2 founds it re-founded as a lexical rule, with the King James licensed as the carrier of a Károli allusion span D's and span E's marking stand, with a measurement behind them for the first time
P1 founds it, P2 does not re-founded in weakened form: the marked paragraphs carry elevated phrasing, but the Bible is not shown to be the source rather than archaizing literary Hungarian generally marking stands; V12 rewritten to claim only what was shown
P1 does not found it stays demoted — a second design has failed to found it, which is not the same as refuting it, and §7's floors are sample-size thresholds, not a power calculation against a registered minimum effect no automatic erratum. The register relabels the marking at ¶162–¶164 and ¶240–¶244 as an explicitly unsupported translator's judgment; the craft report must then defend it on non-statistical grounds or issue erratum 1

The span-level scores are descriptive in every row. They may not be used to license a change to any span, and they exist because finding 9 is right that a population association says nothing about five particular paragraphs.