Repository path: workshop/experiments/E-20260728j-classb-marking/runs/critic.response.md · rendered 2026-09-09
critic.response.md
- TASK A
A1. No. The worst non-mechanical clause is: “where the propositional content does not require it.” The lead decides what the content “requires,” after seeing its own rendering. Likewise m3’s “attributable to a speaker rather than to a neutral narrator” requires an attribution judgment, not form detection. A list of surface forms does not make those exclusion clauses mechanical.
A2. As run, P3 is not a meaningful failure test. The lead selects the forced wording and then decides whether it is “episodic rather than generic”; it can make P3 true by choosing generic wording, or call any counterexample non-episodic. Change required: freeze a referent/event test and have an independent blinded grader apply it, or make P3 a strictly textual criterion (e.g., specified generic quantifiers/classes versus preserved named/event-bound referents).
A3. No. P4 conflates a present-tense English simile with carrying the Italian out-of-sequence mark. Dole’s present in “leaves do” may be ordinary English simile idiom, not a deliberate FID/gnomic carry. The design has no “source-mark correspondence” field or rule. It needs separate codes: English present present; independently idiomatic/required by English construction; and carries the source’s marked sequence contrast. MARKED alone cannot answer the stated question.
A4. No prediction outcome is allowed to threaten the intended closing use. Section 7 turns P4 confirmation into “the pair genuinely lacks the resource,” despite P4 observing one comparator and despite the forced limb potentially supplying alternatives; either falsification route becomes a useful downgrade or erratum. That is not a neutral outcome map. In particular, P4 confirmation cannot license its stated pair-level conclusion.
- TASK B
B1. No. Frozen translation prose is not the only contaminable material. Dole can contaminate the lead’s grading rules, forced-rendering alternatives and rationales, continuous ¶42 retranslation, interpretation of “construction,” NO-COUNTERPART decisions, and the closing report. Stage 1 must include and freeze all lead classifications and analyses that could be revised after Dole, not merely the prose.
B2. Replace discretionary NO-COUNTERPART with a pre-specified alignment procedure: quote Dole’s enclosing sentence/paragraph; identify the source proposition(s) retained, omitted, or redistributed; require a named independently checkable lexical/event anchor; and treat uncertainty as a scored AMBIGUOUS outcome that counts against P4/P5 reportability, not as an exclusion. The lead must not be able to rescue P4 by declaring an inconvenient marked conditional non-corresponding.
B3. Section 7 exceeds §6.3. “One published English «Jeli» does X at these nine sites” is the limit. Section 7’s “the pair genuinely lacks the resource” and “regardless of policy” generalize from Dole’s choices to English–Italian pair capacity, and conflict with the forced-rendering limb’s very purpose.
- TASK C
- P3 cannot genuinely fail under the present self-production/self-classification procedure.
- P6 is stated on a quantity the procedure does not produce. “added lexical items with no counterpart in the Italian (m3 material, counted as tokens added)” requires a frozen token/alignment policy: what counts as a token, multiword particles, function words required by English syntax, distributed paraphrase, and whether an Italian item is a “counterpart.”
- P6 also predicts clustering but its failure condition tests only the total: “Falsified if … 1 or 0.” A result with two additions at gnomic sites satisfies the numeric condition while falsifying “cluster at the conditional sites.”
- P4’s denominator is unstable: exclusions under NO-COUNTERPART are permitted, but P4 remains worded “4 of the 5” and “none of the three.”
- TASK D
A verifier recomputing stored codes, even with wholly separate code, will still miss:
- incorrect or self-serving MARKED/UNMARKED/FRAME classifications;
- m1’s “does not require it” and m3’s attribution judgments;
- source-to-Dole and Italian-to-English correspondence/alignment errors;
- silent changes in tokenization or “same construction” judgments;
- exclusions that alter P4/P5 denominators;
- whether a Dole present is idiomatic English rather than source-mark carriage;
- whether Stage 1’s notes/classifications were actually fixed before comparator exposure.
Independent arithmetic is not independent semantic verification.
- TASK E
- Section 6’s “Fewer than four: the affected sites are excluded” is unenforceable without an adjudicated correspondence rule, and makes the P4/P5 thresholds unreachable or undefined after exclusions.
- P4’s “at least 4 of 5” and “none of three” must specify what happens at every possible exclusion count; “state the exclusion count” is not a decision rule.
- Section 9 does not provide an actual worst-case cost: it names “max_tokens = 8000 at list out-price plus the prompt at list in-price” but gives no model-specific prices, prompt-token maximum, reserve-model maximum, or total cap. It therefore cannot verify the claimed budget.
- The reserve condition “on an empty or
lengthreturn” is not operationally defined: no schema defines empty, no requirement preserves the failed response, and no rule handles other malformed/incomplete output. - The claimed stage ordering is not technically enforced by stated commit hashes, access controls, or a pre-Dole manifest of all Stage-1 judgments.
VERDICT: NEEDS-REDESIGN