Repository path: workshop/experiments/E-20260808b-discordant-marking/critic.md · rendered 2026-09-09
Page metadata (front matter)
| type | note |
|---|---|
| id | E-20260808b-critic |
| status | frozen |
| created | 2026-08-08 |
| updated | 2026-08-08 |
| links | workshop/experiments/E-20260808b-discordant-marking/design.md |
E-20260808b — pre-run critic pass, and what it changed
One call, nvidia/nemotron-3-ultra-550b-a55b (non-panel, no role in the run), given the frozen
design, materials/sites.json and materials/wrong-table.md whole. Raw body:
runs/stage0-critic.json; verbatim text: runs/stage0-critic.txt and critic-raw.txt.
Verdict: NEEDS-AMENDMENT, 2 BLOCKING and 6 ADVISORY. All eight findings accepted; none
overruled. Both BLOCKING findings were fixed and the design re-frozen before Stage 1 dispatched —
no seat had been called.
BLOCKING 1 — the two arms that carry P2 had no defined stimulus
"Stage 2 generates whole-story translations… Stage 3 rates 116 utterance-span items… For IND-PLAIN and IND-FORCED — the two arms that carry the primary contrast P2 — the corresponding 26 English spans must be cut out of the generated full translations, but the design does not state how. If extraction is manual post-hoc, the lead (who knows the hypothesis) chooses which English words represent each site."
Accepted, and this is the finding that would have voided P2. The design as frozen would have
had the lead deciding, after seeing the generated English and knowing the hypothesis, where each
site's words begin and end.
Amendment A1, which removes the extraction step rather than specifying it. The Stage 2 input
is rebuilt so that no extraction is needed: the story is presented as numbered paragraphs in which
every dialogue paragraph has been replaced by its frozen ru span from sites.json — the
attribution clauses («— сказал толстый») removed, narration paragraphs untouched and in place. The
mapping is one-to-one in text order and was verified mechanically before dispatch (10 dialogue
paragraphs ↔ 10 TT sites; 16 ↔ 16 SC sites). The English at numbered paragraph n therefore
is the site span, byte for byte, and the lead cuts nothing.
Cost of the amendment, declared: the independent hand no longer sees «— сказал генерал» inside the dialogue paragraphs, where Garnett did. The surrounding narration, which is untouched and which names the speakers throughout both stories, is still there. Recorded as a new limit (§8.8).
BLOCKING 2 — no target-side variance guard
"
Dconflates systematic bias and increased variance… if Stage-3 raters disagree more on discordant-site translations for any reason,Dinflates at discordant sites even ifEis unbiased. This could make P1/P2 pass for the wrong reason."
Accepted verbatim. Amendment A2 adds the failure criterion the critic wrote:
F7— Stage 3 between-seat SD ofstandingat DISCORDANT sites exceeds that at CONCORDANT sites by > 0.75, in the arm a primary is read on → that primary is reported descriptively only.
ADVISORY, all accepted, none changing a bar
carrieris post-hoc rationalisation, not a causal account. Accepted.P5was already descriptive; it is now additionally labelled exploratory, and no mechanism claim rests on it.- One published hand. Already §8 limit 3; restated on the result page.
- Enrichment. Already §8 limit 1.
P3is not a fair contest — the lead had the title, the author, the speaker labels and the whole story; the Stage 1 seats have none of that. Accepted:P3's hypergeometricPis reported with this caveat attached, andP3licenses nothing about translation or English (the same restrictionRS-20260807d§4.5 put on its ancestor).- REPEAT could be defeated by a rater noticing duplicates. Confirmed by construction and stated here: each block is a separate API call with no shared context, and the two presentations of a REPEAT pair are always in different blocks.
F2is necessary and not sufficient — it validates the instrument on Garnett-style reversals, not on the subtler losses in the generated arms. Accepted and written into the result page besideF2's number.