Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260724-r01r02-selfrevise/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260724-r01r02-selfrevise
statusfrozen
created2026-07-24
updated2026-07-24
provisionaltrue
internal-judgment-onlytrue
linksworkshop/regimes/R01-single-pass.md, workshop/regimes/R02-draft-revise.md, config/models.md, wiki/goodness-senses.md, workshop/canon/kumonoito/manifest.md, workshop/canon/yumejuya-1-2/manifest.md, workshop/experiments/E-20260723-pilot-kumonoito/pilot.md, workshop/experiments/E-20260724-r01r02-selfrevise/critic.md
sensesaccuracy, naturalness, voice, style-correspondence, literary-quality

Frozen design (v2) — the conditional self-revision effect (R01→R02)

Status: v1 drafted S010; independent adversarial critic pass returned FREEZE-AFTER-FIXES (critic.md); this v2 incorporates all five mandatory fixes + the cheap recommended ones and is FROZEN for the run. This is the project's first real regime study under full experiment discipline (experiments README steps 1–4), and it doubles as the shakedown of the jury instrument before jury calibration. Jury calibration has NOT passed (config/models.md): no verdict here carries evidential weight — every result is provisional and internal-judgment-only. The deliverables are (a) does the instrument run cleanly, and (b) a provisional read on the conditional self-revision effect and how much of it is a temperature artifact. Neither is a calibrated claim.

1. Question (reframed per critic A1/A2)

Given a single-pass draft in hand, does one self-revision pass (R02's second call) improve it, sense by sense — and how much of any improvement is the revision act versus the temperature drop that R02 bundles in?

This is the conditional / pipeline question, not a marginal regime comparison. R01 and R02 share an identical draft call by spec (R02 v1.0: the draft call is R01 v1.0 verbatim), so each R02 draft is a valid R01 output, and we judge each revision against its own draft (paired). The paired design removes the between-sample sampling variance that made the S001 pilot's n=1 difference "indistinguishable from sampling noise" (pilot obs. 1).

Two framing caveats the critic forced (both load-bearing): - The paired forced-choice preference rate is NOT the marginal R01-vs-R02 preference. If revision reliably adds a small +ε to whatever draft it gets, the paired preference → ~1.0 while the marginal distributions of R01 and R02 outputs stay nearly identical (independent-samples forced choice → ~0.5). So the paired rate over-reads regime-level discriminability whenever draft-to-draft variance is nonzero. We report the conditional effect and never present it as "R01 vs R02 as regimes." A marginal comparison would need an unpaired (higher-variance) design — future work. - The paired contrast does NOT isolate the revision act. Draft→revision differs in four ways at once: (i) the revision act (source-check), (ii) temperature 0.7→0.4, (iii) a different prompt, (iv) different input (revision sees draft+source). It measures the R01→R02 bundle, not the revision act. The §3 temperature-control arm exists precisely to separate the temperature component from the rest.

2. Materials

Two canon works, register-contrasted, short (keeps the run inside one UTC day's budget). Character counts corrected: wc -m on this container returns bytes (no UTF-8 locale); the true Japanese character counts are ~2,850 and ~3,345 (each source is one whole short work).

work slug author (d.) ~JP chars register axis
蜘蛛の糸 The Spider's Thread kumonoito Akutagawa (1927) ~2,850 formal parable, narrated distance; sustained polite oral-tale です/ます
夢十夜 第一夜・第二夜 yumejuya-1-2 Sōseki (1916) ~3,345 dream/uncanny first-person, Meiji diction

Both are public domain (author death + 70; verified in each canon manifest) and both translations are freshly generated by this project — no third-party copyrighted text is involved, so all artifacts (including raw translations) live in the public tree, no private-texts/ quarantine. Sources: workshop/canon/<slug>/source.txt (whole work) + title/author/year, assembled exactly per R01 v1.0.

Contamination note. Both works are heavily translated into English; a translator model may reproduce remembered phrasings. This does not threaten the paired contrast (draft and revision share any memorization, which differences out), but it means no single output is a clean "from scratch" measurement. The experiment claims only within-pair contrasts, never absolute quality.

3. Regimes, translators, and the temperature-control arm

4. Instrument (jury) and the two comparison sets

The jury judges two blinded pairwise sets on the same items:

5. Predictions (frozen before the run)

Chance = 0.5 (paired forced choice). "Revision-preference" on MAIN = fraction of units on which REV is preferred over its own D7 (ties = 0.5), after averaging the two orderings within juror→item→sense (avoids pseudo-replication, calibration-critic A4). Reported per juror and per cell as primary (critic C1/C2); any pooled number carries a descriptive-only cluster-bootstrap CI.

None of these are calibrated predictions. The jury is uncalibrated; a met/unmet prediction updates the instrument shakedown and gives a provisional signal only.

6. Metrics (all recomputed from raw at verification)

7. Failure criteria (frozen before the run)

8. Procedure

  1. Freeze this v2 (done). Pre-run key snapshot already logged (config/budget.md, 2026-07-24 13:37 UTC).
  2. Translations run (tools/run_selfrevise.py translate): generate D7, REV, D4 for each (work × translator × rep); save runs/translations/<cell>/{d7,revision,d4}.{json,txt} with full request+response+latency+cost. Verify completeness (no truncation/refusal; per R01 §Parameters, a bad output is rerun once, both kept) before any jury spend.
  3. Recompute the jury budget from actual token counts (critic H1) — a hard gate: sum source+D7+REV+D4 tokens, project the 96 jury calls' cost, confirm it fits the day's headroom, set a hard stop. Only then proceed.
  4. Build blinded items (tools/run_selfrevise.py build): for MAIN (REV,D7) and TEMP (D4,D7), emit two orderings each, relabel "Passage 1/2"; write sealed runs/blinding-key.json.
  5. Jury run (tools/run_selfrevise.py judge): each juror (P2,P3) × 24 items × 2 orderings = 96 calls, temp 0.2; save every raw request+response under runs/jury/.
  6. Analysis (analysis.md): compute M1–M5 + diagnostics from raw for MAIN and TEMP.
  7. Independent verification (verification.md): a fresh pass (separate agent instantiation / independent recompute) recomputes every number from raw; discrepancies reported, not smoothed.
  8. Numbers-only finding → wiki/findings/results/ (senses named, provenance linked, provisional: true, internal-judgment-only). Update config/budget.md with API-returned actuals.

9. Budget (pre-flight estimate)

10. Threats to validity (named)

  1. Temperature confound (critic B1) — the whole reason for the TEMP arm; measured, not assumed away.
  2. Draft ≈ revision near-twins → high (correct) ties. M4 contextualizes; the E1 gate keeps correct ties from being read as failure.
  3. Length/polish leakage (critic F1) — a juror preferring the longer/smoother passage would show revision-preference > 0.5 on every sense including accuracy; M5 detects it.
  4. Global-sense low resolution (critic G1) — nulls on voice/literary-quality are uninformative, not evidence of no regression; the P2 sense-class contrast carries cost-detection instead.
  5. Whole-work pairwise judgment is demanding (critic H2) — stresses discrimination, invites the polish shortcut; the evidence field and M5 mitigate; named as a trade-off (matches the regime unit).
  6. Single self-revision round, self-revision only — by design (R02); no generalization beyond one self-revision pass; cross-model revision is R03+.
  7. Two jurors, uncalibrated — no evidential weight; agreement descriptive.
  8. Memorization — differenced out of paired MAIN; flagged for absolute readings (not made).
  9. k=3 / 4 cells under-power — the dominant limitation, disclosed on every number.

Change log