Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260728j-classb-marking/runs/critic.response.md · rendered 2026-09-09

critic.response.md

  1. TASK A

A1. No. The worst non-mechanical clause is: “where the propositional content does not require it.” The lead decides what the content “requires,” after seeing its own rendering. Likewise m3’s “attributable to a speaker rather than to a neutral narrator” requires an attribution judgment, not form detection. A list of surface forms does not make those exclusion clauses mechanical.

A2. As run, P3 is not a meaningful failure test. The lead selects the forced wording and then decides whether it is “episodic rather than generic”; it can make P3 true by choosing generic wording, or call any counterexample non-episodic. Change required: freeze a referent/event test and have an independent blinded grader apply it, or make P3 a strictly textual criterion (e.g., specified generic quantifiers/classes versus preserved named/event-bound referents).

A3. No. P4 conflates a present-tense English simile with carrying the Italian out-of-sequence mark. Dole’s present in “leaves do” may be ordinary English simile idiom, not a deliberate FID/gnomic carry. The design has no “source-mark correspondence” field or rule. It needs separate codes: English present present; independently idiomatic/required by English construction; and carries the source’s marked sequence contrast. MARKED alone cannot answer the stated question.

A4. No prediction outcome is allowed to threaten the intended closing use. Section 7 turns P4 confirmation into “the pair genuinely lacks the resource,” despite P4 observing one comparator and despite the forced limb potentially supplying alternatives; either falsification route becomes a useful downgrade or erratum. That is not a neutral outcome map. In particular, P4 confirmation cannot license its stated pair-level conclusion.

  1. TASK B

B1. No. Frozen translation prose is not the only contaminable material. Dole can contaminate the lead’s grading rules, forced-rendering alternatives and rationales, continuous ¶42 retranslation, interpretation of “construction,” NO-COUNTERPART decisions, and the closing report. Stage 1 must include and freeze all lead classifications and analyses that could be revised after Dole, not merely the prose.

B2. Replace discretionary NO-COUNTERPART with a pre-specified alignment procedure: quote Dole’s enclosing sentence/paragraph; identify the source proposition(s) retained, omitted, or redistributed; require a named independently checkable lexical/event anchor; and treat uncertainty as a scored AMBIGUOUS outcome that counts against P4/P5 reportability, not as an exclusion. The lead must not be able to rescue P4 by declaring an inconvenient marked conditional non-corresponding.

B3. Section 7 exceeds §6.3. “One published English «Jeli» does X at these nine sites” is the limit. Section 7’s “the pair genuinely lacks the resource” and “regardless of policy” generalize from Dole’s choices to English–Italian pair capacity, and conflict with the forced-rendering limb’s very purpose.

  1. TASK C
  1. TASK D

A verifier recomputing stored codes, even with wholly separate code, will still miss:

Independent arithmetic is not independent semantic verification.

  1. TASK E

VERDICT: NEEDS-REDESIGN