Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260729c-neutral-summary/materials/shared-question-and-options.md · rendered 2026-09-09

SHARED BLOCK — identical in both arms of E-20260729c

This block is byte-identical in the two conditions of this run. Only the EVIDENCE block that accompanies it differs.

The binding condition, from the ratified decision D-20260725-06

1. Comparative reception evidence, not additive praise. The evidence must identify both translations, address the same work or a materially relevant portion of it, and support comparability — not merely praise each translator separately. A lead-authored synthesis of diffuse remarks about two translators does not qualify.

D-20260725-06 also holds that "same-quality" for a Tier D held-out arm must be established from independent, pre-existing, external grounds BEFORE the arm runs; that parity may not be created by choosing generators, by surface-measuring prose, or by setting a tolerance after seeing a damage effect; and that a mandatory pre-run factual-damage audit (condition iii) binds separately.

Context — what forced this open

D-20260725-06 resolved to Q-A, the literal reading: Tier D's held-out arm needs materials whose parity rests on independent, pre-existing, external grounds, established before the arm runs. Two conditions bind every candidate pair: (ii) comparative reception evidence naming both translations and supporting comparability, and (iii) a clean pre-run factual-damage audit.

The Turgenev pair — Garnett 1894–99 / Hapgood 1903–05, Записки охотника — is the project's best-configured candidate on materials. Its status:

condition state
(i) materials satisfied: both PD and free whole, source free whole, lead reads Russian
(iii) factual audit satisfied on two independent passages. «Бежин луг», 16 landmarks, S023: both clean. «Свидание», 17 landmarks, S024: Garnett 15/17 raw, both flags adjudicated to non-damage; Hapgood 17/17
(ii) comparative evidence the open question, and it has moved twice in two days

Condition (ii) was recorded as failing in SR-20260725-parity-evidence (S023), on the strength of one unrefereed blog post that ranks Garnett above Hapgood — graded X2 + X7. S024's direct full-text sweep of 1,844 periodical issues (SR-20260725b-periodical-sweep) then found two documents of a different kind and a different weight.

The question

Do these documents satisfy condition (ii) for the Garnett/Hapgood pair — and if so, for which goodness senses?

Options

Option A — YES, condition (ii) is satisfied; the arm may be built. Two independent editorial reviews, one of them a joint review by a critic who read the source and counted, agreeing that the two translators trade the lead by dimension and that neither is better overall. Demanding an explicit "these are equals" sets a bar real criticism does not write to; a documented, corroborated, dimension-by-dimension trade is what external grounds for parity look like. Conditions (i) and (iii) are discharged; build the arm.

Option B — NO; a split verdict is not a parity claim. The ratification requires evidence supporting comparability. Both sources name an asymmetry on English prose, which is close to the project's naturalness and literary-quality senses — exactly what an LM jury scores — and The Nation puts it at 40 passages to 12, which is not a small gap. Building the control that licenses the jury on a pair with a documented, counted difference on the measured dimension imports the bias the control exists to exclude. Condition (ii) still fails.

Option C — YES BUT SENSE-NARROWED. Condition (ii) is satisfied for the senses on which the sources record no asymmetry or favour Hapgood, and fails for those on which they favour Garnett. Concretely, and now with external per-dimension evidence rather than the lead's inference: admissible for accuracy ("almost exactly the same number of errors"; Hapgood "decidedly the more accurate" on the audited cycle) and cultural-mediation (apparatus, notes); inadmissible for naturalness, literary-quality, style-correspondence and voice (40:12 on English idiom; "rise more often to the possibilities of her subject"). Tier D is reportable per sense by design (charter §5), so this costs nothing structurally.

Option D — INSUFFICIENT; require refereed corroboration. SR-20260725-parity-evidence §4 established that evidence strong enough to close a gate is not automatically strong enough to open one. Two unsigned periodical notices are still not refereed scholarship. Note what has changed, though: when this page was first drafted there was one source; there are now two, independent, agreeing. A reviewer choosing D should say what would suffice, or D becomes an unfalsifiable deferral.

Option E — REFRAME: whole-work evidence cannot license a passage-level arm. The ratification named passage level, 100–500 words as the unit. Both sources assess complete editions. The Nation comes closest — its accuracy verdict is scoped to Memoirs of a Sportsman and rests on three sketches of it read against the Russian — but three sketches is still not the 400-word unit, and a general verdict does not transfer to an arbitrary passage. If E is right, the constraint lies on the unit, not on the evidence, and the fix is to redefine the unit rather than keep searching.

Contingent artifacts

None. Deliberately. What exists is: - SR-20260725b-periodical-sweep §5, which grades both documents and explicitly declines to decide; - RS-20260725-svidanie-audit, which discharges (iii) on a second passage and makes no claim about (ii); - config/models.md, unchanged on calibration state.

If this resolves to A or C, the contingent work — designing and running the held-out arm — happens after ratification, not before.

What is NOT in this block, deliberately

The submitting agent's own preference, its "convergence a reviewer should weigh", and the outcome of any previous ratification of this question are withheld from both arms.