Repository path: workshop/experiments/E-20260810-legend-lexis/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260810-legend-lexis |
| status | frozen |
| created | 2026-08-10 |
| updated | 2026-08-10 |
| senses | style-correspondence, voice |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-legend.md, workshop/translations/szent-peter-esernyoje/R05-v1/translation.md, workshop/translations/szent-peter-esernyoje/register.md, wiki/findings/results/RS-20260809e-legend-layer.md, workshop/experiments/E-20260809e-legend-layer/design.md, config/models.md |
E-20260810-legend-lexis — does Mikszáth's archaic layer carry Károli's Bible, or only Károli's verb endings?
Frozen 2026-08-10 (S149) before any statistic on the split was computed. The translation limb it
is wired to was frozen first, at commit f7e9ac8.
0. The wire, in one sentence
Span E rendered the false miracle at ¶240–¶244 in King James English on register rule V12's
premise that Mikszáth's archaic layer is carried by scriptural lexis and syntax; this experiment
tests that premise on the Hungarian, and the translation's own erratum 1 — which would flatten
span D's ¶162–¶164 and span E's ¶240–¶244 together — is the declared price of a null (D47).
1. Question, and why it is a question about the novel
RS-20260809e (S144) demoted V12 and required span E to re-decide it. What that experiment
refuted was a distributional claim — that the marked verb morphology clusters into paragraphs
— at p = 0.1253, with 40 of 55 marked tokens immovable under the null, i.e. substantially a power
failure.
That is not what V12 asserts. V12 is a claim about a carrier: that the layer is marked by
scriptural lexis and syntax rather than by bent verb morphology, and that English should therefore
reach for the King James Bible rather than for quoth and -eth. That claim has never been tested,
and it is testable, because Hungarian has a received biblical register with one dominant source —
Károli Gáspár's Bible — and the forms Mikszáth reaches for are in it: járulának (7 occurrences
in Károli), jövének (97), vagynak (1), hozának (19), rendelé (165), elveszté (15), lőn
(526).
The question. Given a narration paragraph of this novel, does the presence of marked verb morphology predict Károli-Bible phrasing in the rest of that paragraph — the marked forms themselves masked out?
What this teaches about translating literature (subject rule): it decides which English register a translator should reach for at a specific set of paragraphs in a specific novel, and it decides it against the source rather than against the translator's ear. It is not about this project's instruments, statistics or raters.
2. Materials, all free, all already reachable
| id | text | provenance | size |
|---|---|---|---|
| N | «Szent Péter esernyője», whole novel, Révai 1910 | PG #68911, already segmented into 2,133 paragraphs with part and configuration labels at E-20260809e/runs/paragraphs.tsv |
53,470 words |
| C | the S144 occurrence census | E-20260809e/runs/occurrences.tsv — 992 past-tense tokens over 76 lemmas, each labelled MARKED/UNMARKED and NARRATION/TAG/MIXED |
181 MARKED |
| K | Károli Gáspár, Szent Biblia | MEK-00161, 929 chapter files, fetched this session | 485,318 tokens after normalisation |
| J | Jókai Mór, five works — «A jövő század regénye» I, «Erdély aranykora», «Megtörtént regék», «A kis királyok» I, «Életemből» II | PG #55911, #56189, #57029, #57233, #64502 | 495,519 tokens after normalisation |
J is the secular period control and it is the reason this experiment can say anything about the Bible rather than about old formal Hungarian. Jókai is the ambient literary Hungarian Mikszáth wrote inside; if the marked paragraphs are elevated against K and equally against J, the effect is register-general and V12's King James claim is not licensed.
Declared limit on K. MEK-00161 gives no revision year. The digital Károli in circulation is almost certainly the 1908 revision, which postdates the novel by thirteen years; Mikszáth's readers knew an earlier one. The morphology and the clause patterns are continuous across the revisions, but individual verse wordings are not. The direction of this error is conservative: a revision that changes wording reduces n-gram overlap and biases the test toward the null.
Declared limit on N. One witness (R). No page image. collation.md §6, written this session,
records two ⚑ rows that were false; neither falls in the census's material.
3. Instrument
Normalisation (identical for N, K and J, and the reason it is needed is that the three texts are
in three orthographies — R has cz and 1910 vowel length, PG-Jókai has cz, MEK-Károli is
modernised): lowercase → cz → c → fold vowel length only (á→a é→e í→i ó→o ő→ö ú→u ű→ü,
which preserves the o/ö and u/ü distinctions because those are different vowels) → tokens
are maximal runs of [a-zöüäëß].
Masking — the control that makes the test non-circular. The census's 80 accepted MARKED
surface types are deleted from N, from K and from J alike, and no bigram is allowed to span a
deleted position (the token stream is cut there). A marked paragraph therefore cannot score by
containing lőn, and Károli cannot supply a match through its own 526 lőns. Masking is applied to
all MARKED types including TAG ones, which is the conservative choice.
Score. For a paragraph p and comparator X:
overlap(p, X) = |bigrams(p) ∩ bigrams(X)| / |bigrams(p)|
where bigrams are adjacent normalised token pairs after masking, taken as sets (a bigram counts once per paragraph however often it repeats).
Split-blind calibration already run, and reported here because it was run before the split: over
the 387 NARRATION paragraphs of ≥30 words, overlap(·, K) has mean 0.2113, median 0.2045, range
0.0444–0.5429, and no paragraph scores zero. The measure is not degenerate at this paragraph
length. No contrast was computed.
4. Population and split
- Population: paragraphs with
config == NARRATIONinparagraphs.tsvand ≥ 30 words. N = 387. - MARKED: paragraph contains ≥ 1 token labelled
MARKEDandNARRATIONin the census. n = 33 (10 Part I, and 23 outside it). - UNMARKED: paragraph contains no MARKED token of any configuration. n = 354.
- Paragraphs holding only MARKED TAG or MIXED tokens are excluded from both arms, not pooled into UNMARKED.
The 30-word floor is registered as primary; 20 and 40 are registered sensitivities (MARKED n = 39 and 28). Short paragraphs give the ratio a noisy denominator and lose proportionally more bigrams to masking.
5. Strata — the two confounds, controlled by construction
The permutation shuffles the MARKED label within strata, never across the whole population. Strata = length quartile × religious-topic indicator, computed on the population before the split is used.
- Length. Quartiles of paragraph word count within the population.
- Religious topic — the confound that would otherwise decide the result. Paragraphs about God, church, death and burial resemble the Bible whatever their verb morphology, and the marked paragraphs are disproportionately about exactly that. The indicator is 1 if the paragraph contains ≥ 1 token matching the declared lexicon below.
The lexicon, declared here in full and built blind to the split, by prefix on normalised tokens, with the collisions found by inspecting the novel's own vocabulary and excluded by name:
isten* szent* templom* egyhaz* ima imad* angyal* oltar* koporso* temet* halott*
mise* gyon* aldas* predikal* jezus* krisztus pokol* üdvöz* zsolozsma* ereklye*
apostol* plebanos* püspök* kapolna* sekrestye* harang* vallasos bün*
pap* EXCEPT papir* papa* papu* papr* papl*
menny* EXCEPT mennyi* mennyire mennyit mennyivel mennyiert mennyiseg*
kereszt* EXCEPT keresztül*
csoda csodak csodat csodaja csodatetel* csodatevö* (miracle)
EXCEPT csodalkoz* csodalatos* csodalattal (astonishment — and `csodálkozék` is itself a MARKED form)
The csoda split is semantic and deliberate: miracle is topic, astonishment is not, and one of
the astonishment forms is in the masked set, so including it would let the indicator proxy the
outcome.
6. Predictions, registered
P1 — primary. mean overlap(·, K) | MARKED > mean overlap(·, K) | UNMARKED, one-sided,
by within-stratum permutation of the MARKED label, 10,000 permutations, α = 0.05.
P2 — specificity. The standardised effect (difference ÷ pooled SD) against K exceeds the standardised effect against J. Absolute overlap levels are not comparable across comparators — K has 230,167 distinct masked bigrams and J has 336,186, so J matches more of everything — which is why P2 is stated on standardised within-comparator contrasts and why a density-matched sensitivity (J's bigram set deterministically subsampled to K's size, seed 20260810) is registered alongside it.
P3 — locality. P1 survives with Part I excluded (MARKED n = 23). This is RS-20260809e's
amendment A9, imported into the design rather than added by a critic, because Part I is the passage
that generated V12 and the last experiment's whole result turned on it.
7. Power floors, checked before freezing
MARKED n ≥ 25 in the primary (33, met) and ≥ 15 outside Part I (23, met). If a floor had failed the test would not have been run and this page would say so.
8. Failure criteria and the decision table — written before the run
| outcome | V12 | the translation |
|---|---|---|
| P1 fails | struck | erratum 1 flattens span D ¶162–¶164 and span E ¶240–¶244 |
| P1 holds, P3 fails | struck — the effect is again inside the passage that generated it | erratum 1, same scope |
| P1 and P3 hold, P2 fails | re-founded in weakened form: the layer is marked and English should mark it, but the King James is not licensed as the carrier — the effect is register-general | no erratum; V12 is rewritten to say what was and was not shown |
| P1, P2 and P3 all hold | re-founded as a lexical rule with a measurement behind it | no erratum; span D's and span E's marking stand |
This table is the point of the design. The last time V12 was tested, the consequence was argued after the number arrived. Here the consequence is fixed first, in a frozen artifact, and the translation that will pay for it is already committed.
9. What this cannot show, stated before the run
- Nothing about English. Every outcome is about the Hungarian. That King James English is a
good carrier for a Károli allusion is a translator's judgment and stays
internal-judgment-only; the experiment can only say whether there is an allusion to carry. - Bigram overlap is not allusion. Two texts share bigrams for many reasons. What the design can separate is biblical sharing from period-literary sharing (P2); it cannot separate deliberate echo from unconscious idiom.
- The census is a floor (
RS-20260809e§2): the 3sg indefinite slot (monda,kérde) is mechanically undetectable in Hungarian and is excluded, so some MARKED paragraphs are sitting in the UNMARKED arm. This biases toward the null. - One author, one novel, one comparator each side. Nothing here generalises to Hungarian prose.
- Tier D has not passed; every evaluative sentence downstream carries
provisional: true.
10. Procedure
- Design frozen and committed (this file), translation already frozen at
f7e9ac8. - Independent pre-run critic pass, one non-Anthropic panel model, on this page and the built lexicon. Findings recorded with accept/overrule in §11 before any contrast is computed.
run.py— builds the comparators, masks, scores, permutes; raw outputs toruns/.verify.py— an independent path that rebuilds the scores from the stored paragraph text, re-implements the permutation by index sampling rather than label shuffling, and runs mutation tests: (i) drop a masked type, (ii) mask random UNMARKED tokens in UNMARKED paragraphs at the MARKED arm's rate — the masking-artefact control — (iii) collapse the strata, (iv) truncate the comparator. Each must move the statistic in the predicted direction.- Result page; register and translation updated per §8's table; only then is Worswick opened
for the contamination measurement and
D44a's prediction.
11. Critic findings and amendments
Pre-run critic: openai/gpt-5.6-terra (P1), one call, $0.0293675, provider OpenAI, finish_reason:
stop. Verdict NEEDS-REDESIGN, 12 findings (6 BLOCKING, 5 SERIOUS, 1 MINOR). Raw at
runs/critic-out.md. Eleven accepted, one accepted-in-part with the overrule written. No
contrast had been computed when the pass was made and none was computed until every amendment below
was implemented. Sections 3–8 above are superseded where an amendment says so.
A1 (finding 1, BLOCKING — ACCEPTED). The mask becomes status-independent, and the population
becomes the paragraph-level opportunity set. The critic is right that masking only the MARKED
forms cuts only MARKED paragraphs, so the two arms' residual bigram sets are not comparable by
construction. Replaced by: every census past-tense token type — MARKED and UNMARKED, 134 types
in union — is deleted from N and from every comparator, so the deletion rule cannot see the arm.
The population is restricted to narration paragraphs holding ≥ 1 census past-tense token, which
is the paragraph-level form of RS-20260809e's own opportunity set and is the "matched sets with
actual overlap" finding 2 asks for. New population: 237 (MARKED 33, UNMARKED 204), at
≥30 words. The mask removes 1.9% of the novel's tokens and 1.6% of Károli's.
A consequence worth recording: the union mask deletes monda and kérde from all corpora — the
homographic 3sg indefinite slot that RS-20260809e §2 had to exclude. It is now out of the score
entirely. It is still not out of the arm assignment (finding 11).
A2 (finding 1 + 2 — ACCEPTED). Deletion count becomes a stratification variable and the balance table is published. MARKED paragraphs still lose more tokens on average than UNMARKED ones, so paragraphs are compared only within the same deletion band. Bands: 1 and ≥2. Observed: MARKED 10 / 23, UNMARKED 120 / 84 — both bands populated in both arms.
A3 (finding 2, BLOCKING — ACCEPTED IN PART; the overrule is written). The full stratum table, including movable-label counts, is published before the primary is read. What is overruled: the demand to add part, chapter and a rich topic model as stratification variables. With 33 MARKED paragraphs, a five-level part factor crossed with the rest produces cells of size 0 and 1 and destroys the very movable-label count the same finding asks to see. Part enters as P3's heterogeneity analysis instead, which is where it belongs. Strata are therefore deletion band (2) × religious-rate tertile (3) = 6, and the result states plainly that the null is exchangeability given those two variables and nothing else, and that narrative content, character, and local sequence are not controlled. Length is handled by a registered sensitivity (middle length tertile only; and overlap residualised on log word count) rather than by a seventh and eighth cell.
A4 (finding 3, BLOCKING — ACCEPTED IN PART; the overrule is written). The topic measure becomes
continuous and the lexicon is extended. What is overruled: blinded human coders assigning
multi-label scene annotations. This project has no human coders and cannot obtain them; the
amendment is not available at any price, and pretending otherwise would be worse than saying so.
What is accepted: (i) the indicator becomes a continuous religious-term rate per 100 tokens,
tertiled, replacing the binary flag; (ii) the lexicon is extended with the classes the critic named
and that occur in this novel — ördög* pokol* prof* tanitvany* evangel* iras szentiras paradicsom*
feltamad* gyontat* atok* megatkoz* özvegy* arva* barat* kolostor* zarand* bucsu* mennyorszag*
kalvaria* mennybolt* — each checked against the novel's own vocabulary for collisions before use;
(iii) the residual-confounding limit is stated in the result as a primary limit, not a footnote.
A5 (finding 4, BLOCKING — ACCEPTED IN FULL, and it is the best finding). P2 is redefined as a
paired interaction with a real test. The old P2 — "the standardised effect against K exceeds the
standardised effect against J" — had no null, no α and no treatment of the dependence between two
scores from the same paragraph. Replaced by: per paragraph d = overlap(p,K) − overlap(p,H);
statistic mean(d | MARKED) − mean(d | UNMARKED), tested by the same within-stratum permutation
as P1, one-sided, α = 0.05, with a 95% permutation interval. The comparators' different bigram
densities shift d by a constant and cancel in a difference of means, which also disposes of the
subsampling kludge the old design needed.
A6 (finding 5, BLOCKING — ACCEPTED). A second, sharper secular control. One Jókai collection cannot separate Bible from old formal Hungarian. Added: H, an archaizing-secular pool — Kemény Zsigmond «Zord idő» I–II and «A rajongók» I, Jósika «A csehek Magyarországban» II, Gárdonyi «A láthatatlan ember» (PG #69380, #69381, #68979, #67794, #57477), 379,028 tokens, five historical novels whose whole business is elevated archaic Hungarian that is not scripture. All three comparators are systematically subsampled to 379,028 tokens so the levels are comparable. P2's primary control is H, not J — J is reported alongside. What is narrowed: with three comparators there is a range of secular contrasts, not a distribution, and the result says so.
A7 (finding 6 + 7, SERIOUS — ACCEPTED). A type-by-type mask audit is published, with tokens removed from N, K, J and H for each of the 134 types, ambiguous types flagged by name. The claim that masking's effect on the comparator contrast is "simply conservative" is withdrawn; the audit is reported and the direction is left open.
A8 (finding 8, SERIOUS — ACCEPTED IN FULL). The "conservative" claim about the 1908 revision is struck. §2's sentence "The direction of this error is conservative" is withdrawn: a revision can move overlap either way, and it touches K and not the secular controls, so it can move P1 and P2 in either direction. No pre-1897 Károli witness is reachable to this project; the conclusion is limited to this digital Károli and says so.
A9 (finding 9, BLOCKING — ACCEPTED IN FULL, and it overturns the design's own centrepiece).
§8's table treated non-rejection as refutation and let a p-value above 0.05 delete two spans of
frozen translation. That is invalid and the critic is right. §8 is replaced by §12 below. Two
changes: (i) a null is recorded as failure to found, never as refutation, and V12 — which is
already demoted — simply stays demoted, with the register recording that two designs have now
failed to found it; (ii) a population-level association cannot prescribe a change to two particular
paragraphs, so a span-level descriptive analysis is added — pid 163, 164 (span D's ¶162–¶163)
and pid 241, 242, 245 (span E's ¶240, ¶241, ¶244) reported with their own scores and percentiles,
labelled descriptive and not a test. The translator's frozen log at D47 declared that a null
would force erratum 1. That declaration was itself a statistical error, and it is superseded here
rather than acted on. The log is not edited (V2); the supersession is recorded in the register,
in this design and in the result.
A10 (finding 10, SERIOUS — ACCEPTED). P3 is recast as heterogeneity, not replication. Part-I and non-Part-I effects are reported with permutation intervals plus an explicit interaction test. The phrase "the effect is again inside the passage that generated it" is withdrawn as a conclusion a subset analysis can support; it becomes a description of where the estimate sits.
A11 (finding 11, SERIOUS — ACCEPTED). The directional claim about the undetectable 3sg slot is withdrawn, and replaced by a measurement: every UNMARKED narration paragraph in the population is searched for candidate archaic 3sg indefinite forms from a declared list, and the contamination count is reported. No direction is asserted without it.
A12 (finding 12, MINOR — ACCEPTED). Mutation tests become diagnostics, each with a pre-declared expected failure mode and a reported magnitude, and none of them may alter the primary conclusion.
12. Decision table, replacing §8 (amendment A9)
| outcome | what it licenses about V12 | what happens to the translation |
|---|---|---|
| P1 founds it (p < 0.05, positive estimate) and P2 founds it | re-founded as a lexical rule, with the King James licensed as the carrier of a Károli allusion | span D's and span E's marking stand, with a measurement behind them for the first time |
| P1 founds it, P2 does not | re-founded in weakened form: the marked paragraphs carry elevated phrasing, but the Bible is not shown to be the source rather than archaizing literary Hungarian generally | marking stands; V12 rewritten to claim only what was shown |
| P1 does not found it | stays demoted — a second design has failed to found it, which is not the same as refuting it, and §7's floors are sample-size thresholds, not a power calculation against a registered minimum effect | no automatic erratum. The register relabels the marking at ¶162–¶164 and ¶240–¶244 as an explicitly unsupported translator's judgment; the craft report must then defend it on non-statistical grounds or issue erratum 1 |
The span-level scores are descriptive in every row. They may not be used to license a change to any span, and they exist because finding 9 is right that a population association says nothing about five particular paragraphs.