Repository path: workshop/experiments/E-20260810x-legend-again/runs/critic-out.md · rendered 2026-09-09
critic-out.md
NEEDS-REDESIGN
- BLOCKING — “the register is not what moved it” is not supported by this design.
The re-render differs from the earlier version in more than register state: copy-text, elapsed translator experience, knowledge of the work, and the act of re-rendering are all changed. The causal classes are retrospective explanations by the sole translator/coder, not experimentally isolated causes. A count of five sites assigned exclusively to (B) does not identify the register’s total effect.
Proposed amendment: Remove the headline conclusion that “the register is not what moved it.” Replace it with: “Under the coder’s exclusive-cause scheme, five divergence sites were assigned primarily to later register rules; this does not estimate the register’s total causal effect.”
- BLOCKING — The precedence rule mechanically understates (B), so “only five” is a misleading number.
The report explicitly identifies four sites that are governed by later rules but coded as (A) or (C): ¶42 and ¶44 are governed by V20/V21 but counted as (C); ¶20 and ¶11 are governed by V8/V4 but counted as (A). Since precedence is E > A > C > B > D, (B)=5 is an exclusive residual, not the number of sites affected by later rules. The text repeatedly treats it as the latter: “the register’s measurable purchase … is five sites,” “only five are traceable,” and “the register is not what moved it.”
The likely direction of error is unambiguous: (B) is biased downward, while (A) and (C) are inflated as sole explanations whenever a register rule also shaped the English.
Proposed amendment: Report at least two non-exclusive measures: “5 sites primarily coded B” and “9 sites on which a post-S132 rule governed the new wording.” Delete all wording that equates the former with the register’s “purchase,” “effect,” or total number of traceable rule sites.
- SERIOUS — (D)=141 is not a demonstrated class of “free variation”; it is an unaudited residual.
The coder wrote both versions and assigned every class. The residual class absorbs every disagreement not recognized as copy-text, later rule, whole-work recurrence, or correction. It therefore likely includes missed source corrections, missed recurrences, unrecognized rule effects, and ordinary revision. The report admits that the 141 count was not blind-coded or independently reviewed, yet uses it as the dominant explanatory result.
The most likely direction is upward: strict treatment of (E), broad but undocumented “free variation,” and precedence-based exclusion of overlapping causes all enlarge (D). Conversely, some deliberate translator changes could be incorrectly represented as spontaneous variation.
Proposed amendment: Rename the class everywhere to “residual/unattributed variation (coder-assigned)” and replace “141 are free variation” with “141 were not assigned a higher-precedence cause by the translator-coder.” Do not calculate or interpret the 7.4× ratio as a substantive causal result without independent coding.
- SERIOUS — The correction class is deliberately censored downward and cannot support a clean causal partition.
Section 5 says three candidate errors were rejected because they were “under-translation rather than error,” while the same translator judges whether earlier choices were errors. That distinction is evaluative and unstable. If the retained correction definition is intentionally narrow, then (E)=2 is not “the correction count”; it is merely the count meeting this coder’s narrow threshold. The rejected cases then disappear into (D), inflating “free variation.”
Proposed amendment: Add a separate “possible source-adequacy revision” class or table listing all five disputed cases. State that the E/D boundary is not reliable enough to support the claimed distribution of causes.
- BLOCKING — “23 of 24 named decision sites agree” is selection on the outcome, not a valid convergence statistic.
The 24 items were apparently named after both translations were available and are chosen precisely because they are memorable, lexically distinctive, and mostly identical. The list includes repeated instances (“Slovak ×5,” Western name order ×4, four speech tags) while calling them “decision sites,” but does not define whether its denominator is types, occurrences, phrase families, or all translation decisions. The sentence “every one of them a place where a translator had a real choice” is not a sampling frame. There is no basis for treating 23/24 as a rate of convergence over the prose.
The honest denominator would be all eligible, pre-specified decision opportunities in the span under a reproducible inclusion rule—not a hand-selected list of salient agreements. No numerical denominator is presently supplied.
Proposed amendment: Delete the 23/24 statistic and every inference drawn from it. If retained, publish a pre-specified inventory rule, enumerate every eligible occurrence in ¶1–47 before checking agreement, distinguish types from tokens, and include both agreements and disagreements.
- SERIOUS — The 23/24 list is internally noncommensurable even as a descriptive list.
It mixes individual words, multiword renderings, recurrent lexical choices, naming conventions, several person titles, four occurrences of speech tags, and a one-character plural/hyphen difference. A phrase selected because it survived identically is not comparable to a recurring policy or to a lexical choice with several possible English renderings. “23 of 24” gives all these unlike objects equal weight.
Proposed amendment: Present these as illustrative examples only, without a numerator/denominator claim. If a quantitative analysis is wanted, separate lexical-token, lexical-type, phrase, and policy-level decisions.
- BLOCKING — 8.2 versus 8.5 does not establish “the same rate,” “convergence,” or a replication-like result.
The figures differ by 0.3 sites per 100 words: 160/1,963 versus 234/2,764. Even under an inappropriate independence approximation, that difference is far smaller than count uncertainty; the sites are also clustered, alignment-dependent, and not independent observations. More fundamentally, these are two whole-text comparisons in different language conditions, at different work intervals, and under a different translator/coder state. They are not repeated measurements of a common quantity. The page itself concedes “two points,” but the title and opening nevertheless frame them as showing a shared rate.
Proposed amendment: Remove “diverges at the same rate as the first,” “against the only comparable measurement,” and any replication/convergence language. Replace with: “This span’s descriptive site density was 8.2 per 100 new-English words; the earlier Canth exercise reported 8.5 under a similar diff procedure. The two values are not evidence of a stable cross-work rate.”
- SERIOUS — The site-rate comparison is especially fragile because the alignment structure differs.
One old paragraph is aligned to two new paragraphs, the translations have different English lengths, and the site unit is created bySequenceMatcherplus a discretionary merging rule across up to two matching tokens. A one-to-two paragraph mapping and changes in sentence/paragraph structure can alter the number of maximal runs independently of translation divergence. The page treats 160 as though it were a directly comparable count of decisions.
Proposed amendment: Add a sensitivity analysis varying the merge threshold and treating the ¶18/¶19 split separately. If this is not done, state that the 8.2 figure is a property of this particular alignment and site-construction rule, not a robust decision-rate estimate.
- BLOCKING — P-B should not be presented as merely “NOT DECIDED”; the registered prediction is operationally invalid, or fails under the declared analysis rule.
The page has one exclusive coding scheme for the experiment, under which (B)=5 and the 6–15 prediction band fails. It then creates a second, overlapping count of 9 after seeing the assignments, yielding a pass. The fact that the registration did not specify how overlapping causes would be counted does not make the outcome genuinely undecidable; it makes the prediction untestable as registered. Calling it “not decided” shields it from the failure that follows the stated precedence rule while preserving a post-hoc passing alternative.
Proposed amendment: Record P-B as: “INVALID AS REGISTERED: the outcome metric was not operationally specified before analysis. Under the subsequently adopted exclusive-cause rule it would be a failure, 5 versus 6–15; the overlapping count of 9 is descriptive only.” Do not present the two readings as equally confirmatory tests.
-
BLOCKING — The Section 8 denominator is arithmetically and conceptually unclear.
The text says “one of nine bigrams,” “one site of five,” and “five of the nine forms are in Károli.” But the displayed table has eleven marked-form rows, several rows have two queried bigrams, and more than five listed forms have nonzero Károli form counts. It is impossible for a reader to reconstruct the claimed nine-form/nine-bigram denominator from the table. This makes the central numerical conclusion unauditable.Proposed amendment: Before publication, define exactly: (a) the number of marked narration sites, (b) the number of marked-form tokens, (c) the number of distinct form types, and (d) the number of pre-specified collocation tests. Recompute every numerator and denominator from that definition, and provide the underlying query list.
-
BLOCKING — One observed bigram, without an archaizing secular comparator, cannot establish a biblical signal or invert V12.
The sole positive phrase result islőn nagyat ¶244. A formula occurring seven times in Károli may be compatible with biblical echo, but it does not show that the formula is distinctively biblical rather than ordinary nineteenth-century archaizing diction. The page expressly acknowledges this absence of a comparator. That admission is not discharged by calling ¶244 “best-supported,” saying V12 “is inverted,” or claiming that the morphology “is” the layer’s marker.At most, the result refutes a categorical claim that morphology can never occur in a relevant biblical-looking construction, if V12 truly makes that categorical claim. It does not demonstrate that the five marked passages are scriptural, that morphology rather than lexis/syntax marks a biblical layer, or that the English rendering is positively justified.
Proposed amendment: Replace “V12 as written is inverted” with: “The ¶244
lőn nagyformula is a candidate Károli-linked construction. Without an archaizing-secular comparator, this check cannot distinguish biblical association from general archaism and cannot establish the mechanism claimed by V12.” Remove “best-supported rendering” and “positive reason” claims pending the specified comparator design. -
SERIOUS — The use of a likely 1908 revision and boilerplate-contaminated corpus further weakens the Károli claims.
The corpus includes substantial navigation boilerplate and is “almost certainly” a later revision with documented modernization of the copula. Form and phrase frequencies from that witness cannot straightforwardly characterize Károli 1590 or the historical source relevant to the claim. The limits section notes this, but Section 8 still uses raw counts such as 691, 47, and 28 rhetorically as evidence of scriptural morphology.Proposed amendment: Stop presenting raw form frequencies from this corpus as evidence about Károli’s historical stylistic layer. Identify and use an edition-verified, de-boilerplated corpus, or label all current counts as provisional witness-specific retrieval results.
-
SERIOUS — “The paragraph-level design could not have found this device” overstates what the masking result shows.
The result shows that this particular masked paragraph-level instrument excludedlőnandhozának, and therefore could not detect overlap carried by those tokens. It does not establish that the paragraph-level design had no evidential value about other kinds of lexical or syntactic overlap, nor that its null was “not evidence about the novel.” The latter is an inference about the absent unmasked test, not a measurement.Proposed amendment: Rewrite: “Because the masking procedure removed the candidate archaizing forms, its null result does not test overlap carried by those forms. It remains informative only about the unmasked remainder of the paragraph.”
-
SERIOUS — The limits section mentions major defects but does not retract the claims those defects invalidate.
The page correctly states that the sole coder authored both texts, that self-similarity is inseparable from rule-following, that the Károli check has no comparator, and that the corpus may be the wrong revision. But the headline and conclusion continue to make strong claims that depend on precisely those unresolved defects. A limitations paragraph does not license contradictory certainty elsewhere.Proposed amendment: Conform the title, one-sentence result, Sections 6–8 conclusions, and prediction summary to the stated limits. Claims that remain unsupported after the caveats should be removed rather than merely qualified in §11.