Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260816f-night-seam/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260816f-night-seam
statusfrozen
created2026-08-16
updated2026-08-16
senses—
linkswiki/arms/ARM-alf-layla.md, workshop/translations/alf-layla/R05-v1/translation.md, workshop/translations/alf-layla/register.md, workshop/translations/alf-layla/collation.md, wiki/findings/results/RS-20260815f-night-formula.md

Does the night formula carry the work's segmentation into English, or does the frame narrative do it?

Frozen 2026-08-16, after the translation limb it hangs on was frozen and committed — span E, the Tale of the Merchant and the Ifrit complete, log D56–D69, commit e8900b74. The comparators were opened only after that commit.

1. The question, and where it came from

Span E turned up a fact about the copy-text that RS-20260815f could not have seen from one night. The copy-text marks its night boundaries in two different ways and never both at once (D56, V15, collation.md §10):

boundary what the copy-text prints
night 1 → 2 (p. 22/23) dawn formula and frame passage and a heading
night 2 → 3 (p. 27) a heading and nothing else — no formula, no frame passage
night 3 → 4 (p. 28) dawn formula and frame passage and no heading

Witness 2 and Burton's Calcutta II both put one boundary across this stretch, at the p. 28 place, and neither has anything at the p. 27 place. So the copy-text's middle boundary is a typographic boundary with no textual existence: strip the heading and it is gone.

That forces a question the translator cannot answer from the desk. V12 fixes an invariant English sentence for the dawn formula and register.md open question 9 worries about whether its compensating figure wears. Both presuppose that the formula is doing the work of telling a reader where a night ends. It may not be. At every boundary that has a formula, the formula is one sentence inside a seven-sentence frame passage — the sister's praise, the King's resolve, the night in each other's arms, the court convening and breaking up, Dunyazad's request — and any of that would tell a reader that a night had ended.

The question: is the dawn formula load-bearing for segmentation, or is the frame narrative around it sufficient?

What this unit teaches about translating literature (subject rule, wiki/tracks.md): whether a work's structural device reaches its English readers, and which sentence of the seam is carrying it. That is a fact about translating the Nights, not about this project's instruments. It bears directly on a live translation decision — V12, V15, and open question 9 — and on what the framework can say about compensating a device at a seam.

2. Design, as amended by the pre-run critic

The design below is v2. v1 was frozen, sent to an independent pre-run adversarial critic (P1, critic.json, $0.043261) before any body was dispatched, and came back NEEDS REDESIGN with fourteen findings, seven of them BLOCKING. Eleven were accepted and rebuilt into what follows; three are refused in part, with the refusals on the record in §8. v1 was never run. What v1 said, and what changed, is in §8 so that the amendments are auditable rather than invisible.

Five arms, each the same stretch of story: from the opening of the Tale of the Merchant and the Ifrit to the dismissal of the three old men. Two published hands, and one lead rendering in three versions differing by one sentence.

arm text the seam sentence
LEAD-FULL T-alf-layla-R05-v1 spans D+E, 4,267 words, 35 segments V12 as written: And the morning broke upon Shahrazad, and she broke off the talk that was allowed.
LEAD-BRIDGE the same, 4,245 words, 35 segments replaced by And she stopped there. — the cessation kept, the dawn and the figure gone, coherence preserved
LEAD-CUT the same, 4,237 words, 35 segments deleted, which is Lane's published policy applied to this text
BURTON Burton 1885 (PG #3435), 5,292 words, 46 segments his own: And Shahrazad perceived the dawn of day and ceased to say her permitted say.
LANE Lane 1839 (PG #34206), 3,961 words, 33 segments his own at the first boundary; nothing at the second, which he deletes

The three-way lead contrast is the critic's BLOCKING 4 rebuilt. v1 had only FULL and CUT, and the critic was right that deleting the sentence removes a marker and breaks the seam's coherence at once, so a CUT failure would not tell the two apart. BRIDGE separates them: FULL − BRIDGE isolates what the dawn sentence says, holding coherence fixed; BRIDGE − CUT isolates coherence.

Preparation, identical in kind across arms, and logged in full (materials/, manifest.json). Headings, night rubrics and section breaks removed; footnote and note markers stripped ([FN#n], [I_n], [Illustration]); Lane's bracketed editorial paragraph announcing that he will omit the night formulae from here on is removed (§8, refusal 3). The text is then split at sentence boundaries into segments of ~120 words, with a break forced exactly at each printed night boundary — after …the court broke up, and King Shahriyar went into his palace and its equivalents — so that the tested seam falls between segments and never inside one, which is the critic's BLOCKING 2. Segment lengths are held even (medians 113–120 words, range 83–163) so that no truth segment is an outlier in length; v1's forced breaks produced 12-word segments at the seams and would have leaked the answer.

Task, identical across arms. The seat is told the passage is from a story-cycle in which a woman tells a king a story over several nights — she tells for a while, she stops, and later she resumes — and is asked to name every segment after which one night's telling ends and the next begins, quoting the words that told it. The number of boundaries is not given, and the word dawn and the word morning do not appear in the instruction: v1's instruction said she "breaks off at dawn", which told the seat what to hunt for and would have inflated exactly the arms that print it (the critic's MAJOR on demand effects).

Seats: P1 P2 P3 QR. Two replicates on each of the three lead arms (24 bodies), where the primary comparison lives, and one on each published arm (8 bodies), which are descriptive only — 32 bodies, plus 4 memory-assay calls. P4 (note bps), P5 (note bne) and GL (out on any long prompt, NEXT.md) are excluded; these prompts are ~8k tokens, the shape GL died on.

3. Scoring, fixed here before dispatch

The experimental unit is the body, not the body-by-boundary (critic BLOCKING 6: v1's "12 of 16" counted two correlated judgments from one completion as two observations). Per body:

Truth indices, read off the materials before any call (materials/manifest.json):

arm breaks after segment segments
LEAD-FULL / LEAD-BRIDGE / LEAD-CUT 11, 33 35
BURTON 16, 44 46
LANE 13 (and one deleted boundary, not scored) 33

Exact scoring is primary; ±1 is reported as a secondary tolerance and cannot change a verdict.

LANE's second boundary is not scored. Lane abridges the third old man's tale to one paragraph, so the place where the other texts break falls in his last segment, with nothing after it for a break to separate. His deletion is a documentary fact about his text (collation.md §10) and needs no panel.

3a. Response coding, deterministic (critic MAJOR: no coding protocol)

The seat must answer in strict JSON {"breaks":[{"after_segment":int,"cue":str}]}. run.py's parse() strips a code fence, takes the outermost {...}, and json.loads it. Anything that does not parse is a dead body: reported as dead, counted in no rate, never re-rolled. A repeated index collapses to one. A non-integer after_segment is dropped and the body keeps its other entries. A cue that does not occur in that arm's text — checked mechanically, after normalising whitespace and quotation marks — makes that entry a false alarm regardless of its index, because a break supported by words that are not there is not a reading.

4. Predictions, registered

The primary is a comparison with an exact test, not a pair of thresholds (critic MAJOR: v1's thresholds left most of the outcome space uninterpreted).

Every outcome region is reported. A result that satisfies no prediction is written down as a result that satisfies no prediction, with the numbers and the permutation P, and the causal reading is withheld — not the data (critic MINOR on asymmetric suppression, accepted).

5. Controls, and what qualifies the reading

6. Cost, built from the caps the requests permit (note abc)

Worst case: 8,500 input tokens × 32 bodies, max_tokens 1,500 each, plus 4 short memory calls.

seat bodies worst cost
P1 $1.00/$6.00 8 $0.140
P2 $0.75/$3.75 8 $0.096
P3 $2.00/$6.00 8 $0.226
QR $1.475/$4.425 8 $0.166
memory assay (4 calls) 4 $0.020
pre-run critic (P1), already spent 1 $0.043
total worst 37 $0.691

Stop-loss $0.75 (enforced in run.py, which aborts mid-run) · declared ceiling $0.90 · headroom today $1.191 after the critic. Worst < stop-loss < ceiling, as note (bpq) requires.

max_tokens must exceed the reasoning cap, and that is a constraint of the panel and not a preference: QR returns HTTP 400 — "max_completion_tokens must be greater than thinking_budget" — whenever it does not, and P1 returned finish_reason: length with an empty body on a 600/1500 pair in a first attempt at the memory assay. Both were caught before any scored body was bought and are recorded here rather than in the result page's limits. Caps are 1,500 / 900 for the passage bodies and 1,400 / 700 for the memory assay.

7. What each outcome would mean

8. The critic's findings: what was taken, and the three refusals

Taken and rebuilt (BLOCKING unless marked): recall-only primary replaced by D (1); seams moved between segments and scored exactly (2); truth windows abolished (3); LEAD-BRIDGE added (4); D-per-body as the unit, no pseudoreplication (6); memory assay added (8); comparative estimand with an exact permutation test (MAJOR); coding protocol frozen in §3a (MAJOR); Burton and Lane demoted to descriptive (MAJOR); dawn removed from the instruction (MAJOR); failures reported not suppressed (MINOR); "two occurrences of one recurring sentence", not "one sentence" (MINOR).

Refused, with reasons:

  1. "Counterbalance the manipulation across B1 and B2, or one seam per body" (MAJOR). Refused. Every body sees both seams in order, and the second judgment is not independent of the first — which is true, and is why the unit is the body and the score is one number per body. Splitting the seams across bodies would halve the passage and destroy the thing being measured, which is whether a reader following a long stretch of prose can place the seams in it.
  2. "Use held-out or purpose-built texts whose seam structures have not been publicly exposed" (BLOCKING 8's remedy). Refused as out of scope. The question is about this work, whose English is public by definition; a purpose-built cycle would answer a question about segmentation in general and not about the Nights. The threat is instead measured (stage 0) and the result is qualified by what that measurement shows.
  3. "Retain Lane's editorial bracket, or run both Lane conditions" (MAJOR). Refused in part, and the object of study is declared instead: LANE here is Lane's narrative text, not Lane as-published. The bracket is a translator speaking in his own voice about his policy — "as this is expressed in the original work in nearly the same words at the close of every night, such repetitions will in the present translation be omitted" — and it is quoted in full here and in the result page, so nothing is hidden. Running both conditions would cost 8 more bodies to answer a question about paratext that this arm is not asking.