Repository path: workshop/experiments/E-20260724-r01r02-selfrevise/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260724-r01r02-selfrevise |
| status | frozen |
| created | 2026-07-24 |
| updated | 2026-07-24 |
| provisional | true |
| internal-judgment-only | true |
| links | workshop/regimes/R01-single-pass.md, workshop/regimes/R02-draft-revise.md, config/models.md, wiki/goodness-senses.md, workshop/canon/kumonoito/manifest.md, workshop/canon/yumejuya-1-2/manifest.md, workshop/experiments/E-20260723-pilot-kumonoito/pilot.md, workshop/experiments/E-20260724-r01r02-selfrevise/critic.md |
| senses | accuracy, naturalness, voice, style-correspondence, literary-quality |
Frozen design (v2) — the conditional self-revision effect (R01→R02)
Status: v1 drafted S010; independent adversarial critic pass returned FREEZE-AFTER-FIXES (critic.md); this v2 incorporates all five mandatory fixes + the cheap recommended ones and is FROZEN for the run. This is the project's first real regime study under full experiment discipline (experiments README steps 1–4), and it doubles as the shakedown of the jury instrument before jury calibration. Jury calibration has NOT passed (config/models.md): no verdict here carries evidential weight — every result is provisional and internal-judgment-only. The deliverables are (a) does the instrument run cleanly, and (b) a provisional read on the conditional self-revision effect and how much of it is a temperature artifact. Neither is a calibrated claim.
1. Question (reframed per critic A1/A2)
Given a single-pass draft in hand, does one self-revision pass (R02's second call) improve it, sense by sense — and how much of any improvement is the revision act versus the temperature drop that R02 bundles in?
This is the conditional / pipeline question, not a marginal regime comparison. R01 and R02 share an identical draft call by spec (R02 v1.0: the draft call is R01 v1.0 verbatim), so each R02 draft is a valid R01 output, and we judge each revision against its own draft (paired). The paired design removes the between-sample sampling variance that made the S001 pilot's n=1 difference "indistinguishable from sampling noise" (pilot obs. 1).
Two framing caveats the critic forced (both load-bearing): - The paired forced-choice preference rate is NOT the marginal R01-vs-R02 preference. If revision reliably adds a small +ε to whatever draft it gets, the paired preference → ~1.0 while the marginal distributions of R01 and R02 outputs stay nearly identical (independent-samples forced choice → ~0.5). So the paired rate over-reads regime-level discriminability whenever draft-to-draft variance is nonzero. We report the conditional effect and never present it as "R01 vs R02 as regimes." A marginal comparison would need an unpaired (higher-variance) design — future work. - The paired contrast does NOT isolate the revision act. Draft→revision differs in four ways at once: (i) the revision act (source-check), (ii) temperature 0.7→0.4, (iii) a different prompt, (iv) different input (revision sees draft+source). It measures the R01→R02 bundle, not the revision act. The §3 temperature-control arm exists precisely to separate the temperature component from the rest.
2. Materials
Two canon works, register-contrasted, short (keeps the run inside one UTC day's budget). Character counts corrected: wc -m on this container returns bytes (no UTF-8 locale); the true Japanese character counts are ~2,850 and ~3,345 (each source is one whole short work).
| work | slug | author (d.) | ~JP chars | register axis |
|---|---|---|---|---|
| 蜘蛛の糸 The Spider's Thread | kumonoito |
Akutagawa (1927) | ~2,850 | formal parable, narrated distance; sustained polite oral-tale です/ます |
| 夢十夜 第一夜・第二夜 | yumejuya-1-2 |
Sōseki (1916) | ~3,345 | dream/uncanny first-person, Meiji diction |
Both are public domain (author death + 70; verified in each canon manifest) and both translations are freshly generated by this project — no third-party copyrighted text is involved, so all artifacts (including raw translations) live in the public tree, no private-texts/ quarantine. Sources: workshop/canon/<slug>/source.txt (whole work) + title/author/year, assembled exactly per R01 v1.0.
Contamination note. Both works are heavily translated into English; a translator model may reproduce remembered phrasings. This does not threaten the paired contrast (draft and revision share any memorization, which differences out), but it means no single output is a clean "from scratch" measurement. The experiment claims only within-pair contrasts, never absolute quality.
3. Regimes, translators, and the temperature-control arm
- Regimes: R01 v1.0 and R02 v1.0 (both frozen 2026-07-24).
- Three generation arms per (work × translator × rep): 1. D7 — single-pass draft at temp 0.7 = the R01 output and the input to the revision. 2. REV — R02 revision of D7 at temp 0.4 = the R02 output. 3. D4 — an independent single-pass draft at temp 0.4, no revision = the temperature control (critic B1 fix, option a).
- Translators (2, disjoint from the jury): P5
deepseek/deepseek-v4-pro(pilot translator, workhorse) and P1openai/gpt-5.6-terra(frontier, terse). Slugs resolved fromconfig/models.mdat run; run records log the resolved slug. - Replication: k = 3 reps per (work × translator) cell. Temperatures per spec: D7 = 0.7, REV = 0.4, D4 = 0.4.
- Cells: 2 works × 2 translators = 4 cells; 3 reps each = 12 rep-triples. Generation calls = 12 × 3 arms = 36.
4. Instrument (jury) and the two comparison sets
The jury judges two blinded pairwise sets on the same items:
- MAIN = REV vs D7 (paired) — the conditional revision effect (= R01→R02 bundle). 12 items.
-
TEMP = D4 vs D7 (both single-pass; independent samples differing only in temperature) — the temperature contribution. 12 items.
-
Jurors (2, disjoint from translators → no model judges its own output): P2
google/gemini-3.6-flash, P3x-ai/grok-4.5. Two jurors is thin; inter-juror agreement is descriptive only (§6 M3) — no cluster/κ claim (that belongs to calibration). - Task per item: the two passages are blinded and shown as "Passage 1"/"Passage 2" with the Japanese source. Independently per sense, the juror returns a forced pairwise preference:
{winner: 1|2|tie, confidence: low|med|high, rationale: one sentence, evidence: a short quoted phrase from the passage it prefers}. Theevidencefield (critic H2) anchors the rationale and aids verification. No summed scorecard; jurors are never asked "which is better overall." - Senses (5):
accuracy,naturalness(what R02's revision prompt targets), plusvoice,style-correspondence,literary-quality. Resolution caveat (critic G1):voiceandliterary-qualityare global, cumulative senses; a forced pairwise pick between two near-twin passages resolves them poorly, so a null on those is uninformative (not evidence of "no regression").style-correspondence(local, formal) andnaturalnesssuit near-twin pairwise judging best. Sense definitions are handed to the juror verbatim (condensed) fromwiki/goodness-senses.md. - Position-bias control: each item is judged twice per juror with the two passages in swapped display order. n=2 detects gross position bias only; named "position-bias check," not "stability" (calibration-critic F1).
- Blinding key: which display passage is which role, per item and ordering, is written to a sealed
runs/blinding-key.json, opened only at analysis. Jurors get no title/author/translator/regime/temperature label. - Sampling: juror temp 0.2. All raw request+response JSON preserved under
runs/.
5. Predictions (frozen before the run)
Chance = 0.5 (paired forced choice). "Revision-preference" on MAIN = fraction of units on which REV is preferred over its own D7 (ties = 0.5), after averaging the two orderings within juror→item→sense (avoids pseudo-replication, calibration-critic A4). Reported per juror and per cell as primary (critic C1/C2); any pooled number carries a descriptive-only cluster-bootstrap CI.
- P1 — the core bet, per sense (critic C3, D2). On naturalness and (separately) accuracy, MAIN revision-preference is > 0.5, counted as met only if the cluster-bootstrap lower bound also exceeds 0.5; otherwise "directional only." Accuracy is expected to tie out more (light revision rarely changes propositional content detectably).
- P2 — the cost check, as a sense-class contrast (critic D1). On MAIN, revision-preference on the global/texture class {voice, style-correspondence, literary-quality} is lower than on the fluency class {naturalness, accuracy}, on the same items. (Differencing out shared power limits; genuinely falsifiable at this n.)
- P-TEMP — the confound test (critic B1). If P1's fluency-up / P2's texture-down signature is real revision work, it should be weaker or absent in the TEMP set (D4 vs D7). Prediction: TEMP shows a smaller fluency-up/texture-down signature than MAIN. If TEMP shows an equal-or-larger signature, the MAIN signature is attributed to temperature, not the revision act, and reported as such.
- P3 — magnitude, descriptive sanity check (critic D3). Mean normalized word-level edit fraction (D7→REV) is > 0.02 (revision not inert). No hard upper bound — a thorough legitimate revision can be large; % paragraphs touched is reported alongside so a high edit fraction from reordering is distinguishable from wholesale rewriting.
- P4 — register / translator (exploratory, per-cell, no pass/fail). Differences across the 4 cells are annotated, not predicted.
None of these are calibrated predictions. The jury is uncalibrated; a met/unmet prediction updates the instrument shakedown and gives a provisional signal only.
6. Metrics (all recomputed from raw at verification)
- M1 — revision-preference per sense, per cell (primary) and per juror. Order-average within juror→item→sense (tie = 0.5). Report the 4 cells × 5 senses grid, and per-juror. Any pooled rate uses a cluster bootstrap (resample the 4 cells, then reps within) and its CI is descriptive only — with 2 fixed levels per factor nothing is estimable; the number is a summary, not an inference. Computed for both MAIN and TEMP sets. Chance 0.5.
- M2 — position-bias (critic E2 fix). Per sense: (i) flip-rate across the two orderings; (ii) slot preference = raw rate of picking "Passage 1" regardless of role (0.5 = unbiased). Firing rule (pre-committed): a sense is flagged position-confounded only if flip-rate is high (> 1/3) AND slot preference departs materially from 0.5 (|slot − 0.5| > 0.2 for a juror). High flip-rate with slot ≈ 0.5 is reported as indifference (the null), not a confound.
- M3 — inter-juror agreement per sense (descriptive). Fraction of items where P2 and P3 pick the same binarized winner (after order-averaging). No κ/cluster claim.
- M4 — revision magnitude. Per rep-triple, normalized word-level edit fraction D7→REV (and % paragraphs touched). Feeds P3 and the E1 tie gate.
- M5 — length/polish leakage (critic F1/E3). Per sense, correlation between the item's revision-preference and its D7→REV length delta (words). A preference that tracks length is flagged length/polish-cued. Also reported for TEMP (D4→D7 length delta).
- Diagnostics: per-juror parse-failure/refusal rate; per-sense tie rate (interpreted against M4, per E1).
7. Failure criteria (frozen before the run)
- Instrument failure (mechanics): a juror refuses or returns unparseable output on > 20% of its calls → the instrument is unusable; report as a mechanics failure, no preference numbers cited.
- Tie-rate failure — per sense, gated on M4 (critic E1). High tie rate on a sense counts as a problem only where M4 shows substantial edits (cell mean edit fraction above ~0.05) yet the juror still ties > 70% on that sense. Where edits are near-zero, high ties are the expected correct null and are reported as the finding, not discarded.
- Position-confounded sense (critic E2): M2 firing rule met (high flip and slot bias) → that sense's M1 is reported as confounded.
- Length-cued result (critic F1/E3): M5 shows preference tracking length on a sense → reported as polish-cued, not a quality finding.
- Under-power is disclosed, not a failure: k=3 / 4 cells / 2 jurors is stated on every M1 number; CIs are descriptive.
- Honest null is first-class (experiments README): "one self-revision pass buys little this instrument can see at this power" is a legitimate, reportable result, written un-smoothed to
wiki/findings/results/.
8. Procedure
- Freeze this v2 (done). Pre-run key snapshot already logged (
config/budget.md, 2026-07-24 13:37 UTC). - Translations run (
tools/run_selfrevise.py translate): generate D7, REV, D4 for each (work × translator × rep); saveruns/translations/<cell>/{d7,revision,d4}.{json,txt}with full request+response+latency+cost. Verify completeness (no truncation/refusal; per R01 §Parameters, a bad output is rerun once, both kept) before any jury spend. - Recompute the jury budget from actual token counts (critic H1) — a hard gate: sum source+D7+REV+D4 tokens, project the 96 jury calls' cost, confirm it fits the day's headroom, set a hard stop. Only then proceed.
- Build blinded items (
tools/run_selfrevise.py build): for MAIN (REV,D7) and TEMP (D4,D7), emit two orderings each, relabel "Passage 1/2"; write sealedruns/blinding-key.json. - Jury run (
tools/run_selfrevise.py judge): each juror (P2,P3) × 24 items × 2 orderings = 96 calls, temp 0.2; save every raw request+response underruns/jury/. - Analysis (
analysis.md): compute M1–M5 + diagnostics from raw for MAIN and TEMP. - Independent verification (
verification.md): a fresh pass (separate agent instantiation / independent recompute) recomputes every number from raw; discrepancies reported, not smoothed. - Numbers-only finding →
wiki/findings/results/(senses named, provenance linked,provisional: true,internal-judgment-only). Updateconfig/budget.mdwith API-returned actuals.
9. Budget (pre-flight estimate)
- Generation: 36 calls (12 each of D7, REV, D4). P5 near-free; P1 terse. Est. ≈ $0.6–1.1.
- Jury: 24 items × 2 jurors × 2 orderings = 96 calls, each carrying JP source (~3–4k tok) + two whole-work translations (~2–2.5k tok each) + rubric ≈ ~9–11k tok input, short output. Est. ≈ $1.3–1.9 (recomputed from actual tokens at the §8 step-3 gate before spending). Note: the works are ~2,850/3,345 JP chars, not the 8.6k/10k the critic's H1 assumed from the byte count — so input is ~10k tok, not 25–35k.
- Total pre-flight estimate ≈ $1.9–3.0; fund $3.00. Fits one UTC day's $5.00 headroom. Snapshot key before (done); record per-response
usage.cost(summed) as primary; cross-check the key delta. Does not run the same day as any calibration spend without re-checking the ledger.
10. Threats to validity (named)
- Temperature confound (critic B1) — the whole reason for the TEMP arm; measured, not assumed away.
- Draft ≈ revision near-twins → high (correct) ties. M4 contextualizes; the E1 gate keeps correct ties from being read as failure.
- Length/polish leakage (critic F1) — a juror preferring the longer/smoother passage would show revision-preference > 0.5 on every sense including accuracy; M5 detects it.
- Global-sense low resolution (critic G1) — nulls on voice/literary-quality are uninformative, not evidence of no regression; the P2 sense-class contrast carries cost-detection instead.
- Whole-work pairwise judgment is demanding (critic H2) — stresses discrimination, invites the polish shortcut; the
evidencefield and M5 mitigate; named as a trade-off (matches the regime unit). - Single self-revision round, self-revision only — by design (R02); no generalization beyond one self-revision pass; cross-model revision is R03+.
- Two jurors, uncalibrated — no evidential weight; agreement descriptive.
- Memorization — differenced out of paired MAIN; flagged for absolute readings (not made).
- k=3 / 4 cells under-power — the dominant limitation, disclosed on every number.
Change log
- 2026-07-24 (S010) — v1 drafted; independent critic pass (
critic.md, FREEZE-AFTER-FIXES). v2 incorporates all five mandatory fixes — temperature-control arm (D4) + TEMP comparison set (B1); conditional/pipeline reframing, marginal-vs-paired caveat, "isolates the revision pass" struck (A1/A2); position-confound firing gated on flip and slot bias (E2); tie-rate failure gated on M4, per-sense (E1); length/polish diagnostic M5 (F1/E3) — plus the recommended fixes (per-cell/per-juror primary + descriptive CI, C1/C2; CI-excludes-0.5 for P1, C3; P2 as sense-class contrast, D1; P1 per sense, D2; P3 descriptive with % paragraphs, D3; global-sense resolution caveat + literary-quality folded into the cost class, G1/G2; jury-budget recompute gate, H1; juror evidence field, H2). Frozen for the run.