Repository path: workshop/experiments/E-20260724-r01r02-selfrevise/critic.md · rendered 2026-09-09
Page metadata (front matter)
| type | note |
|---|---|
| id | E-20260724-r01r02-selfrevise-critic |
| status | active |
| created | 2026-07-24 |
| updated | 2026-07-24 |
| links | workshop/experiments/E-20260724-r01r02-selfrevise/design.md |
Independent pre-run critic pass — E-20260724-r01r02-selfrevise
Experiment discipline step 2 (experiments README). An independent adversarial critic — a separate agent instantiation, not the orchestrator that drafted design.md — pressure-tested the design before freeze. Verdict: FREEZE-AFTER-FIXES (mandatory: B1 temperature confound; A1/A2 framing; E2 position-firing bug; E1 tie-rate failure bug; F1/E3 length diagnostic). The critique is recorded verbatim below, followed by the disposition of each point. Design v2 incorporates every mandatory fix and the cheap recommended ones; a run may proceed only against v2.
Critique (verbatim)
Independent pre-run critic pass — E-20260724-r01r02-selfrevise
Experiment discipline step 2. An independent adversarial critic (not the agent that drafted
design.md) pressure-tested the frozen design before freeze. The design is careful, honestly hedged, and has clearly absorbed the calibration-v1 critic's lessons (pseudo-replication averaging, "position-bias not stability" naming, descriptive-only inter-juror agreement, correct provisional/internal-judgment-only handling). It is not a redesign candidate. But it has one interpretive confound that makes its core bet uninterpretable as stated, one inference-framing overclaim, and two failure-criterion logic bugs that would suppress exactly the honest null the design says it wants. These are cheap to fix but mandatory.A. What the paired contrast actually measures
A1. "R01 vs R02 as regimes" is not what the paired preference estimates — blocker (framing). The design is right that E[quality(revision) − quality(draft)] is an unbiased estimator of the regime-level mean difference, because the draft distribution is the R01 distribution. But M1 does not report a mean-quality difference; it reports a forced-choice preference rate, P(revision_i ≻ draft_i) on the same item i. That functional is not equal to the regime-level quantity a real R01-vs-R02 deployment would face, which is P(R02_j ≻ R01_k) for independent draws j, k. Concretely: if revision reliably adds a tiny +ε to whatever draft it receives, the paired preference → ~1.0, while the marginal distributions of R01 and R02 are nearly identical, so an independent-samples R01-vs-R02 forced choice → ~0.5. The paired rate therefore systematically overstates regime-level discriminability whenever draft-to-draft variance is nonzero (which is the whole reason for k=3). What M1 legitimately measures is the conditional / pipeline question — "given a draft in hand, does revising it help?" — which is a fine and arguably more decision-relevant quantity, but it is not "R01 vs R02 as regimes," and the title, §1, and the finding must not present the preference rate as a regime-marginal comparison. Fix: reframe the deliverable throughout as the conditional revision effect (does the revision pass improve on its own draft), state explicitly that the paired preference rate is not the independent-samples regime preference and will over-read the latter under draft variance, and note that a marginal R01-vs-R02 comparison would require an (unpaired, higher-variance) design.
A2. "Isolates exactly the thing R02 adds — the revision pass" is false — blocker (sets up B1). §1 asserts the paired contrast holds the draft fixed and isolates the revision pass. It does not. Draft and revision differ in four ways at once: (i) the revision act (checking against source), (ii) temperature 0.7→0.4, (iii) a different prompt, (iv) different input (revision sees draft+source). The contrast isolates the bundle, not the revision act. This is fine for a regime/pipeline comparison (the bundle is R02) but fatal for any mechanistic "the revision pass buys X" claim — which §1's question and P1/P2's rationales repeatedly make. Fix: delete every "isolates the revision pass" claim; scope all mechanistic language to "the R01→R02 delta, which bundles a revision pass, a temperature drop, and a prompt change." See B1 for the sharpest consequence.
B. Temperature confound (the important one)
B1. The predicted P1/P2 signature is identical to the predicted signature of merely lowering temperature — blocker. Draft is sampled at 0.7, revision at 0.4. Lowering temperature shifts token selection toward higher-probability, more conventional, more idiomatic, "safer" choices — which reads as more natural and less idiosyncratically voiced/textured. That is exactly the pattern the experiment predicts and would treat as its headline result: P1 (naturalness/accuracy up) and P2 (voice/style-correspondence down). So a "confirmation" of the core bet is fully consistent with the revision act doing nothing and the effect being a pure temperature artifact. The experiment as designed cannot distinguish "one self-revision pass improves fluency at some cost to voice" from "temp 0.4 prose reads smoother and flatter than temp 0.7 prose." For a bare decision "use R02 or not," temperature is legitimately part of R02 and this is not a confound. For the mechanistic reading the design actually foregrounds, it is disqualifying. Fix (pick one): (a) cheapest and best — add a small control arm: regenerate each draft at temp 0.4 with no revision (12 extra near-free calls, mostly P5), and check whether draft@0.4-vs-draft@0.7 already produces the naturalness-up/voice-down signature; report how much of any observed effect is attributable to temperature before crediting the revision act; or (b) run the revision at 0.7 to match (changes R02's frozen spec — needs a regime version bump, so not this run); or (c) if neither, hard-scope every finding as "R01→R02 delta including a 0.7→0.4 temperature drop; this experiment attributes nothing to the revision act specifically," and strike all "what the revision pass buys" language including from §1's Question. Option (a) is within budget and turns a confound into a measured quantity — strongly preferred.
C. Statistics and independence structure
C1. "12 independent paired items" mis-states the nesting — should-fix. The 12 items are 2 works × 2 translators × 3 reps. Only the 3 reps within a (work×translator) cell are genuine independent replication (independent draft samples). Work and translator are fixed factors with 2 hand-picked levels each, not random draws from a population. A bootstrap that resamples 12 items treats work and translator variance as sampling noise and understates uncertainty about generalization while conflating three variance sources. If the revision effect is large for one work/translator and null for another, pooling hides it and the CI is simply wrong. Fix: make the per-cell breakdown (4 cells × per-sense) the primary object; if bootstrapping, use a cluster/hierarchical resample (resample cells, then reps within), and state plainly that with 2 levels per fixed factor nothing is estimable with a CI — the "12-item CI" is descriptive only. (The design's "wide CI, don't editorialize" hedge is good but the "12 independent items" label contradicts it.)
C2. Juror pooling into the headline — should-fix. M1 says "per juror→item→sense … pool across the 12 items (… per juror as breakdown)." If the headline rate pools both jurors, that is 24 correlated observations (2 jurors × 12 items) counted as if adding independent information. Fix: keep juror as a stratum, report per-juror M1 as primary, and never pool 2 jurors into a 24-denominator rate.
C3. P1's bar is a coin-flip under the null — should-fix. P1 is met if pooled revision-preference point estimate is "> 0.5" with no requirement that a CI exclude 0.5. Under a true null, a point estimate above 0.5 occurs ~half the time; "P1 passes" then carries almost no information, and combined with B1's built-in upward bias it is close to unfalsifiable-to-fail. Fix: require the (cluster-)bootstrap lower bound > 0.5 for P1 to count as met, or explicitly downgrade P1 to "directional signal only, not a test."
D. Predictions — falsifiability
D1. P2 ("voice/style not reliably > 0.5") passes trivially by lack of power — should-fix. This is the mirror of calibration-critic B1: at k=3/2-jurors almost nothing will be "reliably > 0.5," so P2 is nearly guaranteed to "pass" regardless of the truth, and a pass tells you little. A prediction confirmed by your own underpowering is not informative. Fix: reframe P2 as a within-experiment sense-class contrast on the same items — predict revision-preference on {voice, style-correspondence} is lower than on {naturalness, accuracy}. That contrast differences out much of the shared power limitation and is genuinely falsifiable; the paired same-item structure makes it available almost for free. (It also becomes the honest way to detect a cost even at low power.)
D2. P1 conjunction is ambiguous — minor. "On naturalness AND accuracy … pooled revision-preference > 0.5": unclear whether both senses must individually exceed 0.5 or the pooled-over-both rate. Accuracy is the sense most likely to genuinely tie out (light revisions rarely change propositional content detectably), so lumping it with naturalness could sink or inflate P1 arbitrarily. Fix: state P1 per-sense, separately.
D3. P3 thresholds and metric — minor. The 0.02–0.35 band is defended but the 0.35 upper edge is arbitrary: a thorough, legitimate revision of a stiff draft can exceed 0.35 edit fraction without being a "wholesale rewrite." Normalized word-level Levenshtein also scores paragraph reordering as heavy editing. Fix: keep P3 as a descriptive sanity check (it is one), widen/annotate the upper bound, and report % paragraphs touched alongside (already planned) so a high edit fraction from reordering is distinguishable.
E. Failure criteria — two logic bugs
E1. Tie > 70% pooled = instrument failure contradicts the design's own "correct tie-ing is not failure" and regresses the accepted F4 fix — should-fix (blocker for the honest-null goal). §10.1 and §6 correctly note that near-identical draft/revision pairs should produce high ties, and threat #1 says high ties are "correct, not instrument failure." Yet §7 hard-codes "tie on > 70% pooled across senses → instrument/mechanics failure, no numbers cited." So a clean run in which the revision genuinely changes little would trip the failure switch and get thrown out as broken — the design would discard its own honest null. This also reverts calibration-critic F4, which the project accepted (tie-rate failure should apply only to senses where ties are not expected). Fix: gate the tie-rate failure on M4 — it fires only when edits are substantial (M4 above some floor) yet jurors still tie at high rate; where M4 shows near-zero edits, high ties are the expected correct answer and are reported as the null, not a failure. Apply per-sense, not pooled.
E2. Flip-rate > 1/3 → "position-confounded" conflates genuine indifference with position bias and will spuriously suppress real nulls — blocker. On a sense where draft and revision are genuinely near-identical (the expected case, per threat #1), a juror is near-indifferent and its forced pick is ~random, so it will "flip" between the two orderings ~50% of the time — exceeding the 1/3 firing threshold from indifference alone, with no position bias present. M2 would then stamp the sense "position-confounded" and caveat away M1 precisely on the senses where the true answer is "no reliable difference." The design notices the ambiguity ("a flip could be order-bias or genuine indifference") but then still keys the firing rule on flip-rate, so the noticing doesn't reach the rule. Fix: separate the two signals M2 already computes. Position bias = the raw "prefers Passage 1 regardless of role" tendency (systematic slot preference). Fire "position-confounded" only when flip-rate is high and there is a systematic slot preference. High flip-rate without slot preference = indifference → report as the null/tie, not a confound.
E3. A length/polish-cued juror passes every failure criterion while measuring nothing — should-fix. A juror that simply prefers the longer or more-polished-looking passage tracks content, not slot, so it does not flip, shows revision-preference ~1.0, and clears the tie-rate and position rules — while measuring only a length/polish heuristic (see F1). Nothing in §7 catches this. Fix: add a length-vs-preference diagnostic (below) and treat a strong preference–length correlation as a confound flag.
F. Blinding
F1. Role is blinded, but length/polish leakage may make the preference a "polish detector" rather than a quality signal — should-fix. Blinding to which is the revision is correctly handled and is the thing that matters formally. But draft and revision are near-twins from the same model/source; the revision is systematically likely to be longer and (especially at temp 0.4) smoother, giving a stable non-positional cue. A juror picking "the more polished-looking one" would produce revision-preference > 0.5 on every sense — including accuracy, where more polished ≠ more accurate — corrupting those judgments. Compounded with B1 (temp 0.4 = smoother) this is a plausible single alternative explanation for the whole predicted pattern. This threat is not in the §10 list. Fix (near-free, make mandatory): compute and report, per sense, the correlation between revision-preference and the draft→revision length delta (chars/words); if preference tracks length, report the result as length/polish-cued, not a quality finding. Add length/polish leakage to §10.
G. Senses
G1. Global senses have low resolution on within-pair near-twins — should-fix. The typology defines
voiceas "global and cumulative" andliterary-qualityas a whole-text judgment. Forcing a pairwise pick on these between two passages that differ by a few local edits asks the instrument to resolve a global property below its resolution; a null there is expected from the comparison design, not informative about voice/quality regression.style-correspondence(explicitly "local and formal") andnaturalnessare much better suited to a near-twin pairwise call. Fix: state that on global senses the within-pair comparison has low resolution, so a null is uninformative (not evidence of "no regression"); consider that the cost-detection for voice is better carried by the D1 sense-class contrast than by an absolute voice-preference rate.G2.
literary-qualityis judged but appears in no prediction, and the risk-sense set is inconsistent — minor. §4 lists literary-quality among "senses most at risk of regression," but P2 (the cost check) names only voice and style-correspondence, omitting literary-quality. Either it's a cost-check sense (then P2 should include it) or it's exploratory (then say so). This is the calibration-critic's G1 pattern (front-matter/prediction mismatch). Fix: add literary-quality to the P2/D1 cost contrast or label it explicitly exploratory-no-pass/fail.H. Budget and scope
H1. Jury token estimate likely undercounts input by ~2× — should-fix. §9 estimates 12–18k tokens/jury-call, but each call carries the whole JP source (8.6k / 10k chars — Japanese often tokenizes above 1 token/char) plus two full English translations of a whole work (each ~4.5–6k words ≈ 6–8k tokens). Realistic input is ~25–35k tokens/call, so 48 calls could run materially above the $1.0–1.6 jury estimate and the $2.50 fund, risking the $5/day cap when combined with anything else that day. Fix: rebuild the estimate from the actual source+draft+revision token counts (available after the translate step, before any jury spend), reconfirm it fits, and set a hard stop; the procedure already gates jury spend on translation completeness, so insert the recomputed estimate there.
H2. Whole-work pairwise judgment is demanding and feeds the polish heuristic — minor. Asking one call to hold two whole ~5k-word translations and return a global per-sense preference stresses juror discrimination and pushes toward the "pick the smoother one" shortcut (F1); the pilot judged a single section. Whole-work matches the regime, so this is a trade-off, not an error. Fix: at least name it as a threat; optionally have jurors cite a located span per sense-preference (a rationale-anchoring that also helps verification).
I. Conventions and status — compliant
provisional: true+internal-judgment-onlyare applied in front matter and reasserted throughout; the design correctly refuses evidential weight, keeps inter-juror agreement descriptive (no κ/cluster claim — learned from calibration A3/H), forbids the summed scorecard, keeps translators disjoint from jurors (no self-judging), uses valid sense ids, and does not claim calibration. Copyright routing is genuinely clean here (both works PD, project-generated, public tree) — the calibration-critic's E-series does not bite. No fix needed; noted so the disposition record is complete.Single most important weakness
The temperature confound (B1) collapsed onto the mechanistic framing (A2). The experiment's headline result is the P1/P2 signature — naturalness/accuracy up, voice/style down — and that is precisely the signature you would predict from nothing but dropping temperature 0.7→0.4, independent of whether the revision act does any work. Because the design repeatedly claims to "isolate the revision pass," a confirmation would be written up as "one self-revision pass buys fluency at a cost to voice" when it may be "0.4-temperature prose reads smoother and flatter." The confound is a false-positive generator aimed straight at the core bet. It is cheap to convert into a measured quantity (a 12-call draft@0.4 control arm) or to neutralize by hard-scoping every claim to the R01→R02 bundle and deleting the mechanistic language — but until one of those is done, the primary deliverable is uninterpretable in the direction the design most wants to read it.
Mandatory fixes before freeze
- B1 (temperature): add the draft@0.4 control arm, or hard-scope all findings to "the R01→R02 bundle including a 0.7→0.4 temperature drop" and strike every "isolates the revision pass" claim.
- A1/A2 (framing): reframe the deliverable as the conditional/pipeline revision effect; state that the paired forced-choice preference rate over-reads the independent-samples regime comparison and is not "R01 vs R02 as regimes."
- E2 (M2 firing): fire "position-confounded" only on high flip-rate plus a systematic slot preference; treat high-flip/no-slot-preference as indifference (the null), not a confound.
- E1 (tie-rate failure): gate the instrument-failure tie-rate on M4 and apply it per-sense, so an expected high-tie null is not discarded as a broken run.
- F1/E3 (length diagnostic): compute and report per-sense preference-vs-length-delta correlation; flag a length-tracking preference as polish-cued, and add length/polish leakage to §10.
Recommended but not blocking: C1 (cluster/per-cell reporting + drop "12 independent items"), C3 (CI-excludes-0.5 for P1), D1 (reframe P2 as a sense-class contrast), G1/G2, H1 (rebuild the jury token/budget estimate before jury spend).
Verdict: FREEZE-AFTER-FIXES (mandatory: B1, A1/A2, E2, E1, F1/E3).
Disposition (what changed in design v2)
All five mandatory fixes accepted; the cheap recommended fixes (C1, C2, C3, D1, D2, D3, G1, G2, H1, H2) accepted too. Key changes:
- B1 (temperature) — accepted, option (a). Added a third generation arm per cell: a single-pass draft at temp 0.4 (D4), no revision. The jury now judges two comparison sets: MAIN = revision-vs-draft@0.7 (the conditional revision effect, = the R01→R02 bundle), and TEMP = draft@0.4-vs-draft@0.7 (the pure temperature contribution). If the MAIN naturalness-up/voice-down signature also appears in TEMP, it is attributed (wholly or partly) to temperature, not the revision act; the revision-attributable part is reported descriptively as the MAIN-minus-TEMP difference. (§3, §4, §5, §6.)
- A1/A2 (framing) — accepted. The deliverable is reframed as the conditional / pipeline revision effect ("given a draft, does revising it help?"), not a marginal R01-vs-R02 regime comparison. §1 now states that the paired forced-choice preference rate over-reads the independent-samples regime preference under draft variance, and that a marginal comparison would need an unpaired design. Every "isolates the revision pass" claim struck; mechanistic language scoped to "the R01→R02 delta, which bundles a revision act + a 0.7→0.4 temperature drop + a prompt/input change." (§1, §3, throughout.)
- E2 (position firing) — accepted. "Position-confounded" now fires only when flip-rate is high and a systematic slot preference (raw prefers-Passage-1 tendency away from 0.5) is present; high flip-rate with no slot preference is reported as indifference / the null, not a confound. (§6 M2, §7.)
- E1 (tie-rate failure) — accepted. The instrument-failure tie-rate now fires per sense and only where M4 shows substantial edits yet jurors still tie at high rate; where edits are near-zero, high ties are the expected correct null and are reported, not discarded. (§7.)
- F1/E3 (length diagnostic) — accepted, mandatory. M5 added: per-sense correlation between revision-preference and the draft→revision length delta; a preference that tracks length is flagged length/polish-cued, not a quality finding. Length/polish leakage added to §10. (§6 M5, §7, §10.)
- C1 — accepted. Per-cell (4 cells) per-sense breakdown is the primary object; "12 independent items" language dropped; any pooled CI is labeled descriptive only (2 fixed levels per factor → not estimable); bootstrap is a cluster resample (cells, then reps). (§6.)
- C2 — accepted. Juror kept as a stratum; per-juror M1 is primary; the two jurors are never pooled into a 24-denominator (mean-over-jurors-then-over-cells only, clearly labeled). (§6.)
- C3 — accepted. P1 counts as "met" only if the (cluster-)bootstrap lower bound > 0.5; otherwise "directional only." (§5.)
- D1 — accepted. P2 reframed as a within-experiment sense-class contrast: revision-preference on {voice, style-correspondence, literary-quality} < on {naturalness, accuracy}, on the same items. (§5.)
- D2 — accepted. P1 stated per sense. (§5.)
- D3 — accepted. P3 kept as a descriptive sanity check; upper bound annotated (a thorough legitimate revision may exceed it); % paragraphs touched reported alongside. (§5, §6 M4.)
- G1/G2 — accepted. Design now states global senses (voice, literary-quality) have low within-pair resolution → a null there is uninformative; literary-quality folded into the D1 cost contrast (no longer a stray). (§4, §5.)
- H1 — accepted. A jury-cost re-estimate from actual token counts is inserted as a hard gate before any jury spend. (§8, §9.)
- H2 — accepted. Whole-work judgment named as a threat; jurors now cite a short located phrase as evidence per sense (anchors rationale, aids verification, discourages the polish shortcut). (§4, §10.)