Repository path: workshop/experiments/E-20260810-legend-lexis/runs/critic-out.md · rendered 2026-09-09
critic-out.md
NEEDS-REDESIGN
- BLOCKING — Masking does not remove the main construction artefact.
MARKED paragraphs are selected because they contain the tokens that are then deleted. They will therefore have more cuts, fewer eligible bigrams, and differently located missing contexts than UNMARKED paragraphs. Since deletion suppresses both adjacent bigrams, it preferentially removes syntactically informative material around the marked verb. The remaining denominator and remaining bigram set are not comparable by arm. Masking the same surface types in K and J does not repair this paragraph-level asymmetry.
The registered random-UNMARKED-token mutation is not an adequate control: random tokens do not reproduce the identity, grammatical position, adjacency structure, lemma distribution, or clustering of marked finite verbs. It may move in the “predicted” direction even when the primary effect is entirely induced by selective deletion.
Amendment: Redefine the primary representation so that both arms undergo the same syntactic deletion rule independent of MARKED status—for example, remove all finite verbs (or all tokens in a predeclared verb-position class) in every N paragraph and comparator, with no bigrams across the resulting boundaries. Alternatively, score fixed left/right context windows around every finite-verb position while excluding the verb itself, sampling equivalent positions in UNMARKED paragraphs. Demonstrate balance in eligible-bigram count, cut count, and context-window count by arm before testing.
- BLOCKING — The stated permutation null is not justified by the strata.
Within-stratum label permutation tests conditional exchangeability of MARKED status given only length quartile and a binary keyword flag. Marked morphology is plausibly associated with part, narrative event, speaker proximity, temporal distance, character, sentence/paragraph syntax, lexical density, and local narrative sequence. These are not randomized and are not controlled by the two strata. A significant result would be an association under an implausible artificial reassignment of marked verbs, not evidence that morphology predicts biblical phrasing “given a narration paragraph” in the claimed sense.
This is especially acute because the primary arm has only 33 paragraphs. Some of the eight length-by-topic cells may have few or zero MARKED paragraphs; cells with no label variation contribute nothing, while small cells make the randomization distribution unstable and highly dependent on arbitrary quartile boundaries.
Amendment: Publish the complete pre-run stratum table, including total, MARKED, and UNMARKED counts and the number of movable labels per cell. Replace the two-variable stratification with a predeclared matched or conditional model that includes at minimum part, paragraph length continuously, local chapter/section, and a richer topic measure. Restrict inference to matched sets with actual overlap. If no adequate controls exist for a MARKED paragraph, exclude it from the estimand rather than pretending it is exchangeable with all same-quartile paragraphs.
- BLOCKING — The religious-topic control is too weak to support the claimed “Bible rather than topic” interpretation.
A one-keyword binary flag cannot control topic. A paragraph containing one occurrence ofpap*is treated identically to a burial, miracle, liturgy, or biblical-history paragraph saturated with religious language. Conversely, the lexicon omits many likely Bible-associated topics and names: Satan/devil, prophet, disciple, Gospel, Scripture, Israel, Jews, Pharisees, David, Moses, Abraham, paradise, hell variants, sin inflections not covered bybün*, sacrament, confession, resurrection, widow/orphan, curse, covenant, and many non-religious biblical narrative formulae. It also cannot capture a paragraph whose biblical resemblance is syntactic or narrative rather than lexical.
The control therefore risks leaving religious scene content in the MARKED coefficient while describing the residual as a lexical/syntactic carrier independent of topic.
Amendment: Have blinded human coders assign predeclared multi-label scene/topic annotations at paragraph level, with categories such as church/liturgy, death/burial, miracle/legend, biblical history, clergy, rural administration, domestic scene, and narration of action. Use these labels, plus a continuous predeclared religious-term count, in matching/adjustment. Retain the current lexicon only as a sensitivity analysis. Document inter-coder disagreement and adjudication before any outcome contrast is computed.
- BLOCKING — P2 has no valid specificity test.
“The standardized effect against K exceeds the standardized effect against J” is not supplied with a test statistic, null distribution, direction threshold, alpha, or treatment of the dependence between the two scores from the same paragraphs. Comparing two separately standardized mean differences is not a valid demonstration that K is more associated than J. A numerical ordering alone could be sampling noise.
Amendment: Define P2 as a single paired interaction statistic, for example the MARKED–UNMARKED difference in (overlap_K - overlap_J), estimated within the same matched sets/strata and tested by the same valid conditional randomization or model-based procedure. Register its one- or two-sided direction, alpha, confidence interval, and minimum substantively meaningful interaction. Do not call P2 a pass without this inferential procedure.
- BLOCKING — The K–J comparison does not isolate “Bible” from “old formal Hungarian.”
K and J differ not merely in bigram inventory size but in author, genre, textual organization, edition history, formulaic repetition, vocabulary breadth, and internal heterogeneity. K is a highly repetitive multi-book translation; J is five selected prose works by one author. Set-based paragraph overlap rewards membership in a comparator’s inventory and discards precisely the comparator frequencies and repetitiveness that distinguish these corpora. Subsampling J’s distinct bigram set to K’s size creates an arbitrary altered J rather than a comparably repetitive secular corpus. It does not make a standardized K–J contrast a Bible-specific effect.
Amendment: Use multiple predeclared secular comparators matched as far as possible for late-nineteenth-century prose, narrative genre, corpus size, and orthographic treatment, and partition K into comparable units/books. Estimate K’s contrast against the distribution of contrasts from secular corpora, not against one selected J collection. If only J is available, narrow the claim to “K differs from this Jókai collection,” not “biblical rather than period-literary.”
- SERIOUS — Deleting marked surface types from K may erase the very Károli signal under investigation, and the direction is not simply conservative.
The argument is intended to test lexical/syntactic material apart from marked endings, but all 80 MARKED surface types are removed globally from K, where several are common and potentially participate in Károli-specific collocations. This changes K’s bigram inventory disproportionately relative to J. It can lower true K overlap, but it can also selectively remove generic archaic material from K and change the apparent K–J specificity in either direction. “Mask it everywhere” ensures symmetry of the operation, not neutrality of its effect on the comparator contrast.
Amendment: Pre-register two distinct estimands: (a) residual-context overlap after a status-independent all-finite-verb masking rule, and (b) full-text overlap with marked forms excluded only from the paragraph score via fixed context windows. Report both as sensitivity analyses and do not interpret disagreement as merely conservative. Quantify for each comparator the proportion of tokens and distinct bigrams removed by masking.
- SERIOUS — Surface-type masking may delete non-equivalent forms and homographs.
The procedure deletes every occurrence of each of 80 strings in N, K, and J. A surface string can occur in another morphological function, another lemma, quotation, or editorial material. The census labels occurrences in N; it does not establish that every same string in the comparators is the same marked verb phenomenon. Any such mismatch changes K and J differently and invalidates the assertion that the marked morphology has been removed comparably.
Amendment: Provide a type-by-type audit of all 80 forms in all three corpora, including ambiguous forms and token counts removed. Use morphological tagging plus manual audit for ambiguous strings, or exclude ambiguous types from the primary mask and place them in a registered sensitivity analysis.
- SERIOUS — The 1908-revision limitation is incorrectly described as necessarily conservative.
A later Károli revision can reduce overlap through modernization or altered wording, but it can also increase overlap with Mikszáth through modernization, editorial normalization, or adoption of forms closer to the novel’s text. It may affect K but not J, so it can alter both P1 and P2 in either direction. The design cannot claim a one-directional bias without comparing editions.
Amendment: Remove the “conservative” assertion. Either obtain and use a pre-1897 Károli witness, or add a preregistered edition-robustness analysis using at least two documented Károli editions. Limit conclusions to the particular digital Károli edition if no historical witness is available.
- BLOCKING — The decision table treats non-rejection as refutation and overextends group-level results to translation spans.
P1 failure at α=.05 does not show V12 is false. The listed “power floors” are sample-size thresholds, not power calculations for a registered minimum effect, and no equivalence margin is defined. Thus “P1 fails → V12 struck → flatten D and E” licenses a drastic textual consequence from an inconclusive result. P2 failure likewise does not establish that the effect is “register-general”; it may reflect an underpowered or invalid K–J interaction. P3 failure shows only that the estimated association is concentrated in Part I, not that it is invalid in the passages that generated V12.
Further, P1–P3 concern an aggregate population of paragraphs. They do not establish that either ¶162–164 or ¶240–244 has the relevant residual Károli signal. A global association cannot by itself prescribe alteration of two particular translation spans.
Amendment: Replace “fails/holds” consequences with effect-size and interval criteria. Define a smallest effect compatible with retaining V12, an equivalence margin for striking it, and a separately powered test or uncertainty criterion for P2. Add a predeclared span-level descriptive analysis, using only information permitted after the global test, before any change to D or E. If span-level evidence is not part of the design, remove automatic translation consequences entirely.
-
SERIOUS — P3 is not an independent locality confirmation and cannot bear the stated interpretation.
P3 reuses 23 of the 33 primary MARKED paragraphs and their controls. It is a subset analysis, not a replication. Its reduced sample and changed stratum composition make failure compatible with lost precision, while success does not show the effect is homogeneous across the novel. The table’s statement that P3 failure means the effect is “again inside the passage that generated it” is stronger than the test can show.Amendment: Define P3 as an effect-estimation heterogeneity analysis: report Part-I and non-Part-I effects with confidence intervals and their predeclared interaction test. Do not characterize a non-significant outside-Part-I result as confinement to Part I unless a direct interaction and an equivalence criterion support that conclusion.
-
SERIOUS — The claim that missed 3sg forms “bias toward the null” is not secure.
Misclassified MARKED paragraphs enter UNMARKED only if the undetectable forms occur there without another detected MARKED token. If those forms are especially common in biblical or legendary passages, contamination can distort topic balance and either attenuate or, through altered control composition, amplify an association. The direction depends on where the missed forms occur and on their contexts.Amendment: Hand-annotate a random, predeclared sample of putatively UNMARKED narration paragraphs for 3sg indefinite archaic forms and estimate contamination. Run a sensitivity analysis that reassigns likely contaminated paragraphs under pessimistic and optimistic assumptions. Remove the directional-bias claim unless empirically established.
-
MINOR — The mutation-test requirement is not a valid validation criterion as written.
“Each must move the statistic in the predicted direction” has no registered expected magnitude, sampling tolerance, or rationale for why comparator truncation and collapsed strata must have a particular direction. It can become an unprincipled post hoc veto or confirmation mechanism.Amendment: Recast mutation tests as diagnostic reports with predeclared expected failure modes, effect-size outputs, and interpretation rules. Do not make direction alone a pass/fail condition, and do not alter the primary conclusion based on an unregistered diagnostic threshold.