Repository path: workshop/experiments/E-20260723-pilot-kumonoito/pilot.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260723-pilot-kumonoito |
| status | pilot |
| created | 2026-07-23 |
| updated | 2026-07-23 |
| internal-judgment-only | true |
| provisional | true |
| links | workshop/regimes/R01-single-pass.md, workshop/regimes/R02-draft-revise.md, config/models.md |
Pilot — 蜘蛛の糸 §一 under R01 vs R02
PILOT (charter §10.7): mechanics shakedown, evaluated informally only, never citable as evidence. No frozen design, no critic pass, no jury. Everything below is internal-judgment-only and provisional.
What was run
- Source: Akutagawa Ryūnosuke, 蜘蛛の糸 (1918), section 一 only (802 chars) —
source-sec1.txt. Retrieved 2026-07-23 from Aozora Bunko (card 000879, file92_14545.html, Shift_JIS), ruby stripped by tag removal. PD-Japan: died 1927; 1927 + 70 = 1997 < 2026 ✓. (Full-work canon ingestion is separate, later work — this excerpt exists only for the pilot.) - Regimes: R01 v0.1 (one call) and R02 v0.1 (independent draft call + self-revision call), prompts exactly per spec. Deviation from spec, recorded: the "work" supplied was a section, not the whole text.
- Translator: P5 (
deepseek/deepseek-v4-pro), bound perconfig/models.md; temperature 0.7 draft / 0.4 revision. - Raw outputs:
runs/r01-single-pass.json,runs/r02-draft.json,runs/r02-revision.json(full request+response+latency). - Cost: $0.010446 total (3 calls), ledgered in
config/budget.md.
Informal observations (n=1, temperature-confounded; hypotheses, not findings)
- The revision pass did real, identifiable work. R02's draft opened with a calque — "It was a certain day." for ある日の事でございます — and the revision corrected it to "One day," while also fixing ふと ("suddenly glanced" → "happened to glance"). Self-revision caught the draft's single worst line. But R01's single pass produced "One day" directly: at n=1, regime differences are indistinguishable from sampling noise. A real comparison needs the experiment discipline plus repeated runs.
- Neither regime engaged the storyteller register. The source narrates in polite oral-tale です/ます (ございます、居ります) — a marked, sustained stylistic choice. Both outputs default to neutral written narrative English, and nothing in either regime's process could even notice the problem. Filed as open question
OQ-20260723-target-register(with the broader period-English question from Slate C). - Spec compliance is imperfect: R02's draft emitted a title heading ("The Spider's Thread") despite "no headings"; the revision kept it. Regime prompt wording (metadata block invites a title) needs tightening before freeze.
- Local renderings to watch as future evaluation probes: 玉のように ("white as pearls" vs "pure white like jewels"), 覗き眼鏡 ("peep-box" vs "peep-show viewer"), honorific narration of 御釈迦様 (R01 capitalized "His"; R02 did not) — small, checkable texture points of exactly the kind jury instruments will need.
- Mechanics all pass: Aozora fetch + cp932 decode + ruby strip, spec-faithful prompt assembly, raw-JSON provenance capture, per-response cost accounting.
What this pilot changes
- R01/R02 stay
draft; prompt tightening (observation 3) before any freeze. - Register-targeting question opened.
- Nothing else — no evaluative claim about regimes or models is made or implied.