Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260809e-legend-layer/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260809e-legend-layer
statusfrozen
created2026-08-09
updated2026-08-09
sensesstyle-correspondence, voice
internal-judgment-onlytrue
provisionaltrue
linksworkshop/translations/szent-peter-esernyoje/R05-v1/translation.md, workshop/translations/szent-peter-esernyoje/register.md, wiki/findings/results/RS-20260808g-two-pasts.md, wiki/arms/ARM-legend.md

E-20260809e — is Mikszáth's marked past one device or two?

Frozen before any counting was run. The translation limb it hangs off — span D, chapter IV ¶141–¶210 — was committed at a4cdf0b before this page existed.

1. The question, and why it is a question about literature

Register rule V6 renders every marked speech tag in this novel with a plain English past, and RS-20260808g licensed that as a measured non-loss: four independent hands rendered mondá and mondta identically at 11 of 11 sites.

Span D found the same morphology doing something else. At ¶162–¶164 the narrator, having just called Mrs Adamecz's vision stupid chatter, narrates it in járulának, lépkedének, eltemettetének, jövének, vagynak, ki for aki, Mivelhogy — devotional printed Hungarian, switched on exactly where the village's supernatural claim is being reported. These are not speech tags. The translator marked them (register V12: scriptural lexis and syntax, no bent morphology) and wrote V6a to say that V6's evidence is about tags and nothing else.

V12 currently rests on the translator's reading of two paragraphs, and span D logged a counter-instance against it before this design existed — ¶198's kerekedék, an archaic past in plain weather narration with no legend, no vision and no villager reporting anything (D30).

So: is the narration-internal marked past a device, or is it scatter? A translator has to decide, and the two answers give opposite instructions. If the forms are dispersed at random through the narration, V12 is marking noise and must come out of the register. If they are clumped, the clumps are something a translator is obliged to see, and V6's non-loss cannot be stretched over them.

Subject rule (wiki/tracks.md). This asks what Mikszáth's prose does and what an English translator can carry of it. It is not about the project's instruments; RS-20260808g is cited as a scope limit on a translation rule, not audited.

2. Materials

The census is run on the whole novel and not on Part I, because Part I is 8,064 words and the device under test is rare; the arm has translated only Part I but the question is about Mikszáth.

3. Detection — the instrument, published whole

A mechanical sweep of every word type in the novel against the marked-form patterns:

class pattern examples expected
A 3sg past -á/-é verb stem + á/é, word-final mondá, kérdé, ismétlé, kiáltá, rendelé, terjeszté
B 3pl past -ának/-ének word-final járulának, lépkedének, jövének, hozának
C -ék/-ák past word-final kérdék, mondák, mutatkozék, fáradozék
D archaic passive -taték, -teték, -tatának, -tetének eltemettetének
E closed list exact match lőn, vala, valának, valék, vagynak, levének

Every candidate type produced by the sweep is inspected by eye and accepted or rejected, and the full candidate list including every rejection is written to runs/candidates.tsv. The patterns over-generate badly — Hungarian has many nouns and adverbs ending in -á/-é (hozzá, felé, ané, accusatives, -é possessive), and -ának/-ének is also the 3pl possessive-dative. The inspection is the instrument, the regex is only the sieve, and publishing the rejects is what makes the instrument checkable.

The form inventory above was written from chapters III and IV, which the translator has read. That is declared rather than hidden: it means the inventory is not blind, and it is why §5's first check is against RS-20260808g's independently produced census figure.

4. Classification — two axes, both mechanical

Axis 1 — stem class. A token is DICENDI if its stem is on the frozen list below, else NON-DICENDI. The list is fixed here and may not be extended after the run:

mond, kérd, kérdez, felel, válaszol, szól, szólal, szólít, kiált, felkiált, ismétel, tódít, jegyez, dadog, sóhajt, humorizál, mormog, suttog, dörmög, hebeg, rikolt, viszonoz, folytat, közbeszól, pattan, vág, terjeszt, beszél, mesél, kérdezősköd, feddj, int, magyaráz

Axis 2 — paragraph type. A paragraph is DIALOGUE if it begins with a dialogue dash – or contains a dash-delimited attribution; else PROSE.

The 2×2 is reported. The test population is NON-DICENDI tokens in PROSE paragraphs — the population V12 governs and V6 does not.

5. Procedure

  1. Extract the novel body from the Gutenberg file; strip front and back matter; split into paragraphs; record each paragraph's word count and part.
  2. Run the sweep (§3); inspect; write runs/candidates.tsv with every accept and reject.
  3. Check against the published figure before anything else. collation.md §4 reports a S138 census of 103 marked forms in R across the whole novel. If this detector accepts fewer than 80 or more than 130, the discrepancy is resolved and reported before any test is run.
  4. Classify (§4); report the 2×2.
  5. Primary test — dispersion. Null: each NON-DICENDI/PROSE token is assigned independently to a paragraph with probability proportional to that paragraph's word count. Statistic S = the number of paragraphs containing ≥2 such tokens. 10,000 permutations, fixed seed 20260809, two-sided p from the null distribution of S.
  6. Secondary — M = the maximum number of such tokens in any one paragraph, same null.
  7. Descriptive, never primary — what the largest clumps narrate, quoted in Hungarian with the English of any that fall inside a translated span.

6. Gates

7. Predictions, registered

8. Failure criteria, written before the run


§9 Amendments after the pre-run critic — the design above is SUPERSEDED where this section says so

Critic: openai/gpt-5.6-terra (P1), one call, stop, 15,712 chars, $0.0288515, verdict NEEDS-REDESIGN, 14 findings (3 BLOCKING, 8 SERIOUS, 3 MINOR). Raw at runs/critic.json, prompt at runs/prompt-critic.txt. All 14 are accepted; two of them replace the experiment with a better one, and nothing is overruled. Note (blb) — the critic finding S143 overruled is the one that came true — is the reason this section exists in this form.

Amendments, in the order the findings came:

A1 (BLOCKING 1 + 2 — the null was wrong, and so was the statistic). §5.5's null placed tokens into paragraphs in proportion to word count. That models where marked forms fall and cannot tell a device from ordinary clause chaining, one repeated lemma, or a long paragraph. Replaced by an opportunity-set design. Every past-tense token of every lemma that occurs at least once in a marked form is collected; each token is MARKED or UNMARKED. The question becomes the right one — given that Mikszáth is using this verb in the past, when does he reach for the archaic form? — and paragraph length, lemma frequency and verb density all drop out because they are carried by the opportunity set itself.

A2 (BLOCKING 2 — the statistic). Primary statistic is now S₂ = the number of narration paragraphs containing marked tokens from ≥ 2 distinct lemmas. A paragraph cannot enter it by repeating one word. Old S (paragraphs with ≥2 marked tokens) and M (the maximum in one paragraph) are demoted to descriptive and reported without a p-value.

A3 (BLOCKING 3 — the DICENDI list is not a partition). The frozen stem list is withdrawn entirely; feddj was not even a stem. Treatment is now defined by configuration, not by verb meaning: a token is TAG if it stands in a paragraph that opens with a dialogue dash, else NARRATION. Lemma class survives only as a descriptive column.

A4 (SERIOUS 4 — paragraph-level dialogue classification is crude). Accepted conservatively: mixed paragraphs — prose paragraphs containing an internal dash-delimited span — are excluded from the primary population and reported separately, rather than parsed with a span parser this design cannot validate.

A5 (SERIOUS 5 — types vs tokens). Everything is occurrence-level. runs/occurrences.tsv carries one row per token: raw token, normalised token, paragraph id, part, sentence index, context window, proposed lemma, MARKED/UNMARKED, TAG/NARRATION/MIXED, accepted/rejected, and the rejection reason. Every reported number is computed from that file. Matching is case-folded and NFC-normalised; punctuation and quotation dashes are stripped at the token edge.

A6 (SERIOUS 6 — the inventory was hypothesis-shaped). The form inventory is re-derived from the paradigm, not from chapters III–IV. The target is the Hungarian narrative past (elbeszélő múlt) in the third person, plus the archaic copula:

paradigm slot ending e.g.
3sg definite -á / -é mondá, kérdé, ismétlé
3pl indefinite -ának / -ének járulának, jövének
3pl definite -ák / -ék mondák, kérdék
3sg ikes -ék mutatkozék, kerekedék
archaic passive -taték, -teték, -tatának, -tetének eltemettetének
copula, closed exact vala, valának, valék, vagynak, lőn, levének

Declared under-coverage, stated rather than hidden: the 3sg indefinite slot (-a/-e: monda, kérde) is not detectable mechanically — it is homographic with thousands of ordinary nouns, including, unhelpfully, monda "legend". It is excluded, and the census is therefore a floor. Prefixed verbs count as distinct lemmas (megkérdé ≠ kérdé), declared here.

A7 (SERIOUS 7 — "no judgment is involved" was false). It is withdrawn. Accepting or rejecting a candidate occurrence is morphological judgment made by the agent that formed the hypothesis. What is done about it: every decision carries a code and a context window in occurrences.tsv; the acceptance pass is completed and committed before any statistic is computed; and the verifier recomputes the tables from the file by a second path.

A8 (SERIOUS 8 — the 103 check was not validation). Demoted to a smoke test. The 80–130 band is too wide to validate a detector, S138 published no occurrence-level list, and it is this project's own figure rather than an independent one. A mismatch is reported; it certifies nothing.

A9 (SERIOUS 9 — the five parts). Part becomes a blocking variable. Permutation is within part × lemma. Counts, opportunity totals and marked share are reported per part before any test. Sensitivity analysis: the primary is recomputed with Part I excluded, since Part I contains the paragraphs that motivated the inventory.

A10 (SERIOUS 10 — the test was one-sided but declared two-sided). It is one-sided, upper tail; excess clustering is the only predicted direction. p = (1 + #{simulated ≥ observed}) / (B + 1), B = 10,000, seed 20260809, all simulated statistics retained. A low-tail result is reported as a separate descriptive dispersion finding, not as the primary.

A11 (SERIOUS 10 — the failure consequence overreached). Absence of paragraph-scale clustering does not prove absence of a device; a register can be dispersed. The failure consequence is narrowed: if P1 fails, V12 is demoted on the register from a bound rule to a translator's reading carrying "unsupported by the census", and span E is required to re-decide it. It is not silently deleted, and no claim is made that the layer is stylistically insignificant.

A12 (SERIOUS 11 — clumping is not meaning). The three claims are separated and only the first is tested: (i) distributional clustering — tested; (ii) association with legend or ceremony — not tested, and P3 is withdrawn as a prediction; the largest clumps are quoted so a reader can see them, with no label scheme and no confirmation claimed; (iii) translation consequence — not established by anything here. A positive result licenses no English register and no rule.

A13 (BLOCKING 1's chaining concern, by a different mechanism). Clause chaining is handled not by blocking on clause chains — which this design cannot segment reliably — but by a sentence-level sensitivity: the primary is recomputed counting at most one marked token per lemma per sentence, so a run of coordinated verbs cannot manufacture S₂. Reported alongside the primary; if the two disagree the disagreement is the finding.

A14 (MINOR 13, 14 — extraction and simulation reporting). Body anchors are the Gutenberg *** START / *** END markers; headings, part titles and the transcriber's notes are removed by an explicit rule and the removals are counted; paragraphs are blank-line separated, given stable ids, and written out whole to runs/paragraphs.tsv so any contributing paragraph can be checked against the source.

§10 Predictions and failure criteria, as amended