Repository path: workshop/experiments/E-20260802g-programme-divergence/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260802g-programme-divergence |
| status | frozen |
| created | 2026-08-02 |
| updated | 2026-08-02 |
| senses | naturalness, perceived-source-carriage, style-correspondence, cultural-mediation, accuracy |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/findings/results/RS-20260802g-programme-divergence.md, wiki/base/sources/S-arnold-newman-homer.md, workshop/regimes/R12-arnold-1861.md, workshop/regimes/R13-newman-1856.md, workshop/translations/iliad-priam/R12-v1/translation.md, workshop/translations/iliad-priam/R13-v1/translation.md, wiki/arms/ARM-fluency-record.md, config/budget.md |
E-20260802g — do two published translation programmes decide the same passage differently?
⚠ Discipline breach, declared at the head rather than in a footnote
This design did not receive an independent pre-run critic pass. continue-prompt.md §6 requires
one, and note (rr) had recorded thirty-eight consecutive sessions of compliance. The streak is
broken here and the reason is not a good one: the session's API budget and attention went to the
ratification gate (which overran on a finish_reason: length failure costing $0.2215), and the
coding stage was dispatched without routing the design to a critic first.
What was in place, and it is not nothing: the design below was written and the two controls were
fixed before either coding call was dispatched; the lead's own 22 codes were frozen before the
seats' answers were read; the coders were blind to the hypothesis, to the lead's codes, and to the
category census. What was missing is the adversarial reading that has, in this project, amended
or voided a design in thirty-eight consecutive sessions — including, at S092, a control the critic
computed before dispatch and found failing. A design of this shape is exactly the kind that pass
catches. The result page carries the omission in its limits and NEXT.md's drift check names it.
1. Question
Executed as frozen rule sets by a third hand on one real passage, at how many decision sites does Arnold's 1861 programme forbid Newman's rendering, and Newman's forbid Arnold's? In what grammatical category do the divergences fall?
2. Why it is not method work
What does this unit teach about translating literature? What a translator's stated method actually
governs in the text, and in which grammatical category two opposed methods part company. The
subject is two published programmes and a Greek passage; the seats are instruments and their
reliability is a limit, not the question (wiki/tracks.md §subject rule).
3. Materials
- Greek: Iliad XXIV.486–512 (locus) and XXIV.552–570 (contamination gate), Monro–Allen OCT via
Perseus
tlg0012.tlg001.perseus-grc2, stored atmaterials/source-grc.json. - Regimes:
R12(Arnold, 10 rules),R13(Newman, 11 rules), both frozen before the Greek was opened for translation. - Renderings:
T-iliad-priam-R12-v1,T-iliad-priam-R13-v1, lead as labeled subject, $0, each with a frozen decision log. - Site book:
materials/site-book.json, 22 sites, each carrying the Greek and both renderings. - Coding prompt:
materials/coding-prompt.txt, containing both rule sets with their authors unnamed (called Programme P and Programme Q), all 22 sites, and the two questions.
4. Procedure
- Render the gate span from the Greek alone; run
tools/dependence_check.pyagainst Butler 1898 and Lang–Leaf–Myers 1883; proceed only on 0 shared 12-grams. Ran first, passed. - Render the locus twice, once under each regime, from the Greek alone. Freeze both logs.
- Build the site book from the union of the two logs' decision points.
- Lead codes all 22 sites on both questions; freeze the codes (they are literals in
analysis/verify.py). - Dispatch the site book and both rule sets to two blind non-Anthropic seats, P5 and P3.
- Recompute every reported number in
analysis/verify.py, which imports nothing fromtools/. - Mutation-test the verifier.
5. Controls, fixed before dispatch
- FC1, negative. Site 17: the two renderings are the identical word softly. No rule set can
forbid a rendering identical to its own. Any
yfrom any coder at site 17 is a coder error and the coding is unusable. - FC2, prior positive. Site 2, question A: Newman writes Eld, and Arnold in Last Words names
eld as a bad word in precisely this construction ("from Eld and Death exempted"). A coder
actually applying A4/A5 must answer
y. Fewer than 2 of 3 codings answeringymeans the seats are not applying the rule sets and no divergence count is licensed.
6. Predictions, written before the seats answered
- The two programmes will be found to diverge at more than half the sites. (Held: 19 of 22 majority.)
- Divergences will be predominantly lexical, because Arnold's own concession is that Newman's syntax is Homeric and his diction is not. This prediction is derived from the source read this session, not blind, and is stated as derived-then-tested rather than as a blind forecast. (Held: 14 of 19, with 3 of the 4 syntactic ones downstream of a lexical rule.)
- The lead's divergence count will be within 3 of each seat's. (FAILED against P5: 15 against 19, a gap of 4. Held against P3: 15 against 18, a gap of 3.)
- The two seats will agree with each other less than each agrees with the lead, because the lead wrote the renderings and the rule summaries. (FAILED, and inverted: the seats agree with each other at 0.886 and with the lead at 0.705 and 0.727. The lead is the outlier.)
7. Failure criteria, pre-registered
- F1. Either control fails → no divergence count is reported. (Did not fire.)
- F2. Fewer than 20 of 22 answer lines parse from either seat → that seat is dropped and the primary is reported on the remainder, with the loss stated. (Did not fire; 22/22 both.)
- F3. If the three codings disagree on more than half the sites' divergence status, the scheme is reported as not reproducible and the primary is withheld. (Did not fire: 17 of 22 sites are unanimous on divergence status, 5 split.)
- F4. Any seat returning
finish_reason: lengthis retried once at a raised cap; a second failure drops the seat. (Did not fire.)
8. Budget
Pre-flight worst case $0.20 for the coding stage (2 calls, max_tokens 14,000, sized from note
(bhq) — a seat's measured reasoning appetite plus the answer, not the answer alone). Actual
$0.0862628, 43% of the reservation. Translation, the gate, the site book, all coding by the lead,
analysis/verify.py and both mutation tests are lead work at $0.00; lead translation is never
ledgered (charter §3, A4). tools/dependence_check.py used unmodified; nothing in tools/
changed this session.