Repository path: workshop/experiments/E-20260809e-legend-layer/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260809e-legend-layer |
| status | frozen |
| created | 2026-08-09 |
| updated | 2026-08-09 |
| senses | style-correspondence, voice |
| internal-judgment-only | true |
| provisional | true |
| links | workshop/translations/szent-peter-esernyoje/R05-v1/translation.md, workshop/translations/szent-peter-esernyoje/register.md, wiki/findings/results/RS-20260808g-two-pasts.md, wiki/arms/ARM-legend.md |
E-20260809e — is Mikszáth's marked past one device or two?
Frozen before any counting was run. The translation limb it hangs off — span D, chapter IV
¶141–¶210 — was committed at a4cdf0b before this page existed.
1. The question, and why it is a question about literature
Register rule V6 renders every marked speech tag in this novel with a plain English past, and
RS-20260808g licensed that as a measured non-loss: four independent hands rendered mondá and
mondta identically at 11 of 11 sites.
Span D found the same morphology doing something else. At ¶162–¶164 the narrator, having just
called Mrs Adamecz's vision stupid chatter, narrates it in járulának, lépkedének,
eltemettetének, jövének, vagynak, ki for aki, Mivelhogy — devotional printed Hungarian,
switched on exactly where the village's supernatural claim is being reported. These are not speech
tags. The translator marked them (register V12: scriptural lexis and syntax, no bent morphology)
and wrote V6a to say that V6's evidence is about tags and nothing else.
V12 currently rests on the translator's reading of two paragraphs, and span D logged a
counter-instance against it before this design existed — ¶198's kerekedék, an archaic past in
plain weather narration with no legend, no vision and no villager reporting anything (D30).
So: is the narration-internal marked past a device, or is it scatter? A translator has to decide, and the two answers give opposite instructions. If the forms are dispersed at random through the narration, V12 is marking noise and must come out of the register. If they are clumped, the clumps are something a translator is obliged to see, and V6's non-loss cannot be stretched over them.
Subject rule (wiki/tracks.md). This asks what Mikszáth's prose does and what an English
translator can carry of it. It is not about the project's instruments; RS-20260808g is cited as a
scope limit on a translation rule, not audited.
2. Materials
- R — Révai Testvérek, Budapest, 1910, Project Gutenberg #68911, the whole novel: five parts, 53,470 words. Public domain; the same witness that is the arm's copy-text (V1).
- Cost $0 on the counting. One API call is budgeted for the pre-run critic (§6).
The census is run on the whole novel and not on Part I, because Part I is 8,064 words and the device under test is rare; the arm has translated only Part I but the question is about Mikszáth.
3. Detection — the instrument, published whole
A mechanical sweep of every word type in the novel against the marked-form patterns:
| class | pattern | examples expected |
|---|---|---|
A 3sg past -á/-é |
verb stem + á/é, word-final |
mondá, kérdé, ismétlé, kiáltá, rendelé, terjeszté |
B 3pl past -ának/-ének |
word-final | járulának, lépkedének, jövének, hozának |
C -ék/-ák past |
word-final | kérdék, mondák, mutatkozék, fáradozék |
| D archaic passive | -taték, -teték, -tatának, -tetének |
eltemettetének |
| E closed list | exact match | lőn, vala, valának, valék, vagynak, levének |
Every candidate type produced by the sweep is inspected by eye and accepted or rejected, and the
full candidate list including every rejection is written to runs/candidates.tsv. The patterns
over-generate badly — Hungarian has many nouns and adverbs ending in -á/-é (hozzá, felé,
ané, accusatives, -é possessive), and -ának/-ének is also the 3pl possessive-dative. The
inspection is the instrument, the regex is only the sieve, and publishing the rejects is what makes
the instrument checkable.
The form inventory above was written from chapters III and IV, which the translator has read.
That is declared rather than hidden: it means the inventory is not blind, and it is why §5's
first check is against RS-20260808g's independently produced census figure.
4. Classification — two axes, both mechanical
Axis 1 — stem class. A token is DICENDI if its stem is on the frozen list below, else NON-DICENDI. The list is fixed here and may not be extended after the run:
mond, kérd, kérdez, felel, válaszol, szól, szólal, szólít, kiált, felkiált, ismétel, tódít,
jegyez, dadog, sóhajt, humorizál, mormog, suttog, dörmög, hebeg, rikolt, viszonoz, folytat,
közbeszól, pattan, vág, terjeszt, beszél, mesél, kérdezősköd, feddj, int, magyaráz
Axis 2 — paragraph type. A paragraph is DIALOGUE if it begins with a dialogue dash – or
contains a dash-delimited attribution; else PROSE.
The 2×2 is reported. The test population is NON-DICENDI tokens in PROSE paragraphs — the population V12 governs and V6 does not.
5. Procedure
- Extract the novel body from the Gutenberg file; strip front and back matter; split into paragraphs; record each paragraph's word count and part.
- Run the sweep (§3); inspect; write
runs/candidates.tsvwith every accept and reject. - Check against the published figure before anything else.
collation.md§4 reports a S138 census of 103 marked forms in R across the whole novel. If this detector accepts fewer than 80 or more than 130, the discrepancy is resolved and reported before any test is run. - Classify (§4); report the 2×2.
- Primary test — dispersion. Null: each NON-DICENDI/PROSE token is assigned independently to a paragraph with probability proportional to that paragraph's word count. Statistic S = the number of paragraphs containing ≥2 such tokens. 10,000 permutations, fixed seed 20260809, two-sided p from the null distribution of S.
- Secondary — M = the maximum number of such tokens in any one paragraph, same null.
- Descriptive, never primary — what the largest clumps narrate, quoted in Hungarian with the English of any that fall inside a translated span.
6. Gates
- Pre-run critic. An independent non-Anthropic panel model reads this page before the run and is
asked for the ways the design is wrong. Its objections are recorded and answered in the result,
and where an objection is overruled the overruling is written down.
(blb)fired last session on exactly this: the critic finding S143 overruled is the one that came true. - Post-run verification recomputes every reported number from the raw outputs by a second path.
- No judgment is involved anywhere in this design, so no jury, no blinding, and no Tier D question
arises. Every evaluative sentence about the translation remains
internal-judgment-only.
7. Predictions, registered
- P1 (primary). Observed S exceeds the null's 97.5th percentile — the narration-internal marked forms are clumped.
- P2. DICENDI tokens outnumber NON-DICENDI tokens: V6's material is the bulk of the device.
- P3 (descriptive). The largest clumps fall in passages narrating the supernatural, the legendary, or the ceremonial.
8. Failure criteria, written before the run
- If observed S falls inside the null's central 95%, P1 fails: the layer is scatter, V12 is
withdrawn from the register as unsupported by anything but a reading of two paragraphs, and the
marking in span D ¶162–¶164 goes into the errata discussion for the craft report. This outcome is
live —
D30already logged a counter-instance. - If the NON-DICENDI/PROSE population is smaller than 10 tokens, the test is underpowered and no p-value is reported at all. The finding is then that V12 governs a negligible layer, which also argues for withdrawing it, and the result says so instead of reporting a test.
- If §5 step 3's check fails, the detector is wrong or the S138 figure is, and that is settled
and reported before any test — a published figure of this project's own is at stake, which is the
one condition under which method work may run inside a unit (
continue-prompt.md§4.5). - A clumped result does not license V12 as written. It licenses only that there is something
there. Whether scriptural English is the right carrier is a craft judgment and stays
internal-judgment-only.
§9 Amendments after the pre-run critic — the design above is SUPERSEDED where this section says so
Critic: openai/gpt-5.6-terra (P1), one call, stop, 15,712 chars, $0.0288515, verdict
NEEDS-REDESIGN, 14 findings (3 BLOCKING, 8 SERIOUS, 3 MINOR). Raw at runs/critic.json,
prompt at runs/prompt-critic.txt. All 14 are accepted; two of them replace the experiment with
a better one, and nothing is overruled. Note (blb) — the critic finding S143 overruled is the
one that came true — is the reason this section exists in this form.
Amendments, in the order the findings came:
A1 (BLOCKING 1 + 2 — the null was wrong, and so was the statistic). §5.5's null placed tokens into paragraphs in proportion to word count. That models where marked forms fall and cannot tell a device from ordinary clause chaining, one repeated lemma, or a long paragraph. Replaced by an opportunity-set design. Every past-tense token of every lemma that occurs at least once in a marked form is collected; each token is MARKED or UNMARKED. The question becomes the right one — given that Mikszáth is using this verb in the past, when does he reach for the archaic form? — and paragraph length, lemma frequency and verb density all drop out because they are carried by the opportunity set itself.
A2 (BLOCKING 2 — the statistic). Primary statistic is now S₂ = the number of narration paragraphs containing marked tokens from ≥ 2 distinct lemmas. A paragraph cannot enter it by repeating one word. Old S (paragraphs with ≥2 marked tokens) and M (the maximum in one paragraph) are demoted to descriptive and reported without a p-value.
A3 (BLOCKING 3 — the DICENDI list is not a partition). The frozen stem list is withdrawn
entirely; feddj was not even a stem. Treatment is now defined by configuration, not by verb
meaning: a token is TAG if it stands in a paragraph that opens with a dialogue dash, else
NARRATION. Lemma class survives only as a descriptive column.
A4 (SERIOUS 4 — paragraph-level dialogue classification is crude). Accepted conservatively: mixed paragraphs — prose paragraphs containing an internal dash-delimited span — are excluded from the primary population and reported separately, rather than parsed with a span parser this design cannot validate.
A5 (SERIOUS 5 — types vs tokens). Everything is occurrence-level. runs/occurrences.tsv
carries one row per token: raw token, normalised token, paragraph id, part, sentence index, context
window, proposed lemma, MARKED/UNMARKED, TAG/NARRATION/MIXED, accepted/rejected, and the rejection
reason. Every reported number is computed from that file. Matching is case-folded and NFC-normalised;
punctuation and quotation dashes are stripped at the token edge.
A6 (SERIOUS 6 — the inventory was hypothesis-shaped). The form inventory is re-derived from the paradigm, not from chapters III–IV. The target is the Hungarian narrative past (elbeszélő múlt) in the third person, plus the archaic copula:
| paradigm slot | ending | e.g. |
|---|---|---|
| 3sg definite | -á / -é |
mondá, kérdé, ismétlé |
| 3pl indefinite | -ának / -ének |
járulának, jövének |
| 3pl definite | -ák / -ék |
mondák, kérdék |
| 3sg ikes | -ék |
mutatkozék, kerekedék |
| archaic passive | -taték, -teték, -tatának, -tetének |
eltemettetének |
| copula, closed | exact | vala, valának, valék, vagynak, lőn, levének |
Declared under-coverage, stated rather than hidden: the 3sg indefinite slot (-a/-e:
monda, kérde) is not detectable mechanically — it is homographic with thousands of ordinary
nouns, including, unhelpfully, monda "legend". It is excluded, and the census is therefore a
floor. Prefixed verbs count as distinct lemmas (megkérdé ≠ kérdé), declared here.
A7 (SERIOUS 7 — "no judgment is involved" was false). It is withdrawn. Accepting or rejecting a
candidate occurrence is morphological judgment made by the agent that formed the hypothesis.
What is done about it: every decision carries a code and a context window in occurrences.tsv; the
acceptance pass is completed and committed before any statistic is computed; and the verifier
recomputes the tables from the file by a second path.
A8 (SERIOUS 8 — the 103 check was not validation). Demoted to a smoke test. The 80–130 band is too wide to validate a detector, S138 published no occurrence-level list, and it is this project's own figure rather than an independent one. A mismatch is reported; it certifies nothing.
A9 (SERIOUS 9 — the five parts). Part becomes a blocking variable. Permutation is within part × lemma. Counts, opportunity totals and marked share are reported per part before any test. Sensitivity analysis: the primary is recomputed with Part I excluded, since Part I contains the paragraphs that motivated the inventory.
A10 (SERIOUS 10 — the test was one-sided but declared two-sided). It is one-sided, upper tail; excess clustering is the only predicted direction. p = (1 + #{simulated ≥ observed}) / (B + 1), B = 10,000, seed 20260809, all simulated statistics retained. A low-tail result is reported as a separate descriptive dispersion finding, not as the primary.
A11 (SERIOUS 10 — the failure consequence overreached). Absence of paragraph-scale clustering does not prove absence of a device; a register can be dispersed. The failure consequence is narrowed: if P1 fails, V12 is demoted on the register from a bound rule to a translator's reading carrying "unsupported by the census", and span E is required to re-decide it. It is not silently deleted, and no claim is made that the layer is stylistically insignificant.
A12 (SERIOUS 11 — clumping is not meaning). The three claims are separated and only the first is tested: (i) distributional clustering — tested; (ii) association with legend or ceremony — not tested, and P3 is withdrawn as a prediction; the largest clumps are quoted so a reader can see them, with no label scheme and no confirmation claimed; (iii) translation consequence — not established by anything here. A positive result licenses no English register and no rule.
A13 (BLOCKING 1's chaining concern, by a different mechanism). Clause chaining is handled not by blocking on clause chains — which this design cannot segment reliably — but by a sentence-level sensitivity: the primary is recomputed counting at most one marked token per lemma per sentence, so a run of coordinated verbs cannot manufacture S₂. Reported alongside the primary; if the two disagree the disagreement is the finding.
A14 (MINOR 13, 14 — extraction and simulation reporting). Body anchors are the Gutenberg
*** START / *** END markers; headings, part titles and the transcriber's notes are removed by an
explicit rule and the removals are counted; paragraphs are blank-line separated, given stable ids,
and written out whole to runs/paragraphs.tsv so any contributing paragraph can be checked
against the source.
§10 Predictions and failure criteria, as amended
- P1 (primary, one-sided). Observed S₂ exceeds the 97.5th percentile of the null formed by permuting MARKED/UNMARKED labels within part × lemma.
- P2. TAG tokens outnumber NARRATION tokens.
- ~~P3~~ withdrawn (A12).
- Failure: P1 not met → V12 demoted per A11. Fewer than 10 NARRATION marked tokens, or fewer than 4 distinct NARRATION lemmas → no p-value is reported at all; the finding is descriptive.
- The §8 bullets stand, as narrowed by A11 and A12.