Repository path: workshop/experiments/E-20260816f-night-seam/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260816f-night-seam |
| status | frozen |
| created | 2026-08-16 |
| updated | 2026-08-16 |
| senses | — |
| links | wiki/arms/ARM-alf-layla.md, workshop/translations/alf-layla/R05-v1/translation.md, workshop/translations/alf-layla/register.md, workshop/translations/alf-layla/collation.md, wiki/findings/results/RS-20260815f-night-formula.md |
Does the night formula carry the work's segmentation into English, or does the frame narrative do it?
Frozen 2026-08-16, after the translation limb it hangs on was frozen and committed — span E,
the Tale of the Merchant and the Ifrit complete, log D56–D69, commit e8900b74. The comparators
were opened only after that commit.
1. The question, and where it came from
Span E turned up a fact about the copy-text that RS-20260815f could not have seen from one night.
The copy-text marks its night boundaries in two different ways and never both at once
(D56, V15, collation.md §10):
| boundary | what the copy-text prints |
|---|---|
| night 1 → 2 (p. 22/23) | dawn formula and frame passage and a heading |
| night 2 → 3 (p. 27) | a heading and nothing else — no formula, no frame passage |
| night 3 → 4 (p. 28) | dawn formula and frame passage and no heading |
Witness 2 and Burton's Calcutta II both put one boundary across this stretch, at the p. 28 place, and neither has anything at the p. 27 place. So the copy-text's middle boundary is a typographic boundary with no textual existence: strip the heading and it is gone.
That forces a question the translator cannot answer from the desk. V12 fixes an invariant English
sentence for the dawn formula and register.md open question 9 worries about whether its
compensating figure wears. Both presuppose that the formula is doing the work of telling a reader
where a night ends. It may not be. At every boundary that has a formula, the formula is one
sentence inside a seven-sentence frame passage — the sister's praise, the King's resolve, the night
in each other's arms, the court convening and breaking up, Dunyazad's request — and any of that would
tell a reader that a night had ended.
The question: is the dawn formula load-bearing for segmentation, or is the frame narrative around it sufficient?
What this unit teaches about translating literature (subject rule, wiki/tracks.md): whether a
work's structural device reaches its English readers, and which sentence of the seam is carrying it.
That is a fact about translating the Nights, not about this project's instruments. It bears directly
on a live translation decision — V12, V15, and open question 9 — and on what the framework can
say about compensating a device at a seam.
2. Design, as amended by the pre-run critic
The design below is v2. v1 was frozen, sent to an independent pre-run adversarial critic
(P1, critic.json, $0.043261) before any body was dispatched, and came back
NEEDS REDESIGN with fourteen findings, seven of them BLOCKING. Eleven were accepted and rebuilt
into what follows; three are refused in part, with the refusals on the record in §8. v1 was never
run. What v1 said, and what changed, is in §8 so that the amendments are auditable rather than
invisible.
Five arms, each the same stretch of story: from the opening of the Tale of the Merchant and the Ifrit to the dismissal of the three old men. Two published hands, and one lead rendering in three versions differing by one sentence.
| arm | text | the seam sentence |
|---|---|---|
LEAD-FULL |
T-alf-layla-R05-v1 spans D+E, 4,267 words, 35 segments |
V12 as written: And the morning broke upon Shahrazad, and she broke off the talk that was allowed. |
LEAD-BRIDGE |
the same, 4,245 words, 35 segments | replaced by And she stopped there. — the cessation kept, the dawn and the figure gone, coherence preserved |
LEAD-CUT |
the same, 4,237 words, 35 segments | deleted, which is Lane's published policy applied to this text |
BURTON |
Burton 1885 (PG #3435), 5,292 words, 46 segments | his own: And Shahrazad perceived the dawn of day and ceased to say her permitted say. |
LANE |
Lane 1839 (PG #34206), 3,961 words, 33 segments | his own at the first boundary; nothing at the second, which he deletes |
The three-way lead contrast is the critic's BLOCKING 4 rebuilt. v1 had only FULL and CUT, and
the critic was right that deleting the sentence removes a marker and breaks the seam's coherence at
once, so a CUT failure would not tell the two apart. BRIDGE separates them: FULL − BRIDGE
isolates what the dawn sentence says, holding coherence fixed; BRIDGE − CUT isolates
coherence.
Preparation, identical in kind across arms, and logged in full (materials/, manifest.json).
Headings, night rubrics and section breaks removed; footnote and note markers stripped ([FN#n],
[I_n], [Illustration]); Lane's bracketed editorial paragraph announcing that he will omit the
night formulae from here on is removed (§8, refusal 3). The text is then split at sentence
boundaries into segments of ~120 words, with a break forced exactly at each printed night
boundary — after …the court broke up, and King Shahriyar went into his palace and its equivalents
— so that the tested seam falls between segments and never inside one, which is the critic's
BLOCKING 2. Segment lengths are held even (medians 113–120 words, range 83–163) so that no truth
segment is an outlier in length; v1's forced breaks produced 12-word segments at the seams and
would have leaked the answer.
Task, identical across arms. The seat is told the passage is from a story-cycle in which a woman tells a king a story over several nights — she tells for a while, she stops, and later she resumes — and is asked to name every segment after which one night's telling ends and the next begins, quoting the words that told it. The number of boundaries is not given, and the word dawn and the word morning do not appear in the instruction: v1's instruction said she "breaks off at dawn", which told the seat what to hunt for and would have inflated exactly the arms that print it (the critic's MAJOR on demand effects).
Seats: P1 P2 P3 QR. Two replicates on each of the three lead arms (24 bodies), where
the primary comparison lives, and one on each published arm (8 bodies), which are descriptive
only — 32 bodies, plus 4 memory-assay calls. P4 (note bps), P5 (note bne) and GL (out on any long prompt, NEXT.md) are
excluded; these prompts are ~8k tokens, the shape GL died on.
3. Scoring, fixed here before dispatch
The experimental unit is the body, not the body-by-boundary (critic BLOCKING 6: v1's "12 of 16" counted two correlated judgments from one completion as two observations). Per body:
- hit — a reported break index exactly equal to a truth index for that arm;
- false alarm — any other reported index;
D= hits − false alarms, the primary per-body score. This is the critic's BLOCKING 1: a recall-only score can be won by naming every segment, andDcannot.
Truth indices, read off the materials before any call (materials/manifest.json):
| arm | breaks after segment | segments |
|---|---|---|
LEAD-FULL / LEAD-BRIDGE / LEAD-CUT |
11, 33 | 35 |
BURTON |
16, 44 | 46 |
LANE |
13 (and one deleted boundary, not scored) | 33 |
Exact scoring is primary; ±1 is reported as a secondary tolerance and cannot change a verdict.
LANE's second boundary is not scored. Lane abridges the third old man's tale to one paragraph,
so the place where the other texts break falls in his last segment, with nothing after it for a break
to separate. His deletion is a documentary fact about his text (collation.md §10) and needs no
panel.
3a. Response coding, deterministic (critic MAJOR: no coding protocol)
The seat must answer in strict JSON {"breaks":[{"after_segment":int,"cue":str}]}. run.py's
parse() strips a code fence, takes the outermost {...}, and json.loads it. Anything that does
not parse is a dead body: reported as dead, counted in no rate, never re-rolled. A repeated index
collapses to one. A non-integer after_segment is dropped and the body keeps its other entries. A
cue that does not occur in that arm's text — checked mechanically, after normalising whitespace and
quotation marks — makes that entry a false alarm regardless of its index, because a break
supported by words that are not there is not a reading.
4. Predictions, registered
The primary is a comparison with an exact test, not a pair of thresholds (critic MAJOR: v1's thresholds left most of the outcome space uninterpreted).
P1(primary). meanD(LEAD-FULL) > meanD(LEAD-CUT), tested by exact permutation over the 16 bodies (8 vs 8, 12,870 relabelings), one-sided, α = 0.05. The dawn sentence is doing work the rest of the seam does not do.P2(the decomposition).D(FULL) >D(BRIDGE) — the content of the dawn sentence matters, not merely that some sentence occupies the slot. If insteadD(FULL) ≈D(BRIDGE) >D(CUT), what the seam needs is any sentence that says the telling stopped, andV12's wording is free to be anything.P3(descriptive, published hands).BURTONmeanD≥ 1.0 over 4 bodies andLANEmeanD≥ 0.5 at its one scored boundary, over 4 bodies. Descriptive only: Burton and Lane differ from the lead in wording, length and training exposure, so neither can validate the instrument (critic MAJOR, accepted).P4(cue check). Of hits inLEAD-FULLandBURTON, ≥ 0.75 quote the dawn sentence. Reported as a fraction with its denominator; an arm with no hits is reported as no hits, not as 0.P5(the invisible boundary). No more than 1 of 24 lead-arm bodies reports a break at segment 26 or 27, where the copy-text's p. 27 heading falls and where the English, followingV15, prints nothing.
Every outcome region is reported. A result that satisfies no prediction is written down as a result that satisfies no prediction, with the numbers and the permutation P, and the causal reading is withheld — not the data (critic MINOR on asymmetric suppression, accepted).
5. Controls, and what qualifies the reading
- Prior knowledge, measured and not assumed (critic BLOCKING 8). v1 claimed
LEAD-CUTmeasured the memory threat; the critic was right that it does not. Stage 0 is a direct assay: each seat is asked, with no text at all, to name from memory the events after which Shahrazad breaks off in this tale. If a seat names B1 and B2 correctly from memory, every hit that seat scores is ambiguous between reading and recall, and the design says so on the face of the result rather than claiming a clean measurement. The assay bounds the threat; it does not remove it, and no sentence in the result page may say it does. - Cue-presence check. §3a's rule — a quotation absent from the text makes the entry a false alarm — catches the recalled hit that dresses itself in invented words. It does not catch a recalled hit dressed in real ones, and the design does not pretend otherwise.
- Failure criterion, reported not suppressed. If mean
D(LEAD-FULL) < 1.0 — a text that marks its seams with a dawn sentence and a seven-sentence frame passage, and readers still cannot place them — the instrument has failed. The numbers are still published; the causal reading is not. - Note (bpu) does not bite. The arms differ by the presence and content of one sentence, not by
how the same content is marked. What the
CUTdeletion removes and does not remove: she stopped speaking survives nowhere else in the seam; the night passed survives in the frame's own passed that night in each other's arms till morning, which is a sentence about the King and not about the telling. That residue is whatBRIDGEexists to hold constant. - Note (bps). Replicate 1 on the longest arm (
BURTON, ~8k tokens) is dispatched first, one call per seat, andfinish_reasonis read on all four before the other 36 bodies are bought. Any seat returninglengthis dropped and the fact recorded. - The lead does not judge its own translation (charter §5). No arm asks for a quality judgment; the task is localisation, and no prompt states who wrote anything.
6. Cost, built from the caps the requests permit (note abc)
Worst case: 8,500 input tokens × 32 bodies, max_tokens 1,500 each, plus 4 short memory calls.
| seat | bodies | worst cost |
|---|---|---|
P1 $1.00/$6.00 |
8 | $0.140 |
P2 $0.75/$3.75 |
8 | $0.096 |
P3 $2.00/$6.00 |
8 | $0.226 |
QR $1.475/$4.425 |
8 | $0.166 |
| memory assay (4 calls) | 4 | $0.020 |
pre-run critic (P1), already spent |
1 | $0.043 |
| total worst | 37 | $0.691 |
Stop-loss $0.75 (enforced in run.py, which aborts mid-run) · declared ceiling $0.90 · headroom
today $1.191 after the critic. Worst < stop-loss < ceiling, as note (bpq) requires.
max_tokens must exceed the reasoning cap, and that is a constraint of the panel and not a
preference: QR returns HTTP 400 — "max_completion_tokens must be greater than thinking_budget" —
whenever it does not, and P1 returned finish_reason: length with an empty body on a 600/1500
pair in a first attempt at the memory assay. Both were caught before any scored body was bought and
are recorded here rather than in the result page's limits. Caps are 1,500 / 900 for the passage
bodies and 1,400 / 700 for the memory assay.
7. What each outcome would mean
P1andP2hold — the dawn sentence is load-bearing and its content matters.V12's wording is structural, and open question 9 is a question about a working device.P1holds,P2fails (FULL≈BRIDGE>CUT) — the slot is load-bearing and the wording is not. Any sentence saying the telling stopped would serve, and the elaborate compensationD43bought at that seam is paid for something the segmentation does not need. That would be the most useful outcome forframework/v0.2§7.P1fails becauseCUTis also found — the frame narrative carries the segmentation and the dawn sentence is ornament at a seam the reader finds anyway.- The failure criterion fires — numbers published, causal reading withheld.
8. The critic's findings: what was taken, and the three refusals
Taken and rebuilt (BLOCKING unless marked): recall-only primary replaced by D (1); seams moved
between segments and scored exactly (2); truth windows abolished (3); LEAD-BRIDGE added (4);
D-per-body as the unit, no pseudoreplication (6); memory assay added (8); comparative estimand with
an exact permutation test (MAJOR); coding protocol frozen in §3a (MAJOR); Burton and Lane demoted to
descriptive (MAJOR); dawn removed from the instruction (MAJOR); failures reported not suppressed
(MINOR); "two occurrences of one recurring sentence", not "one sentence" (MINOR).
Refused, with reasons:
- "Counterbalance the manipulation across B1 and B2, or one seam per body" (MAJOR). Refused. Every body sees both seams in order, and the second judgment is not independent of the first — which is true, and is why the unit is the body and the score is one number per body. Splitting the seams across bodies would halve the passage and destroy the thing being measured, which is whether a reader following a long stretch of prose can place the seams in it.
- "Use held-out or purpose-built texts whose seam structures have not been publicly exposed" (BLOCKING 8's remedy). Refused as out of scope. The question is about this work, whose English is public by definition; a purpose-built cycle would answer a question about segmentation in general and not about the Nights. The threat is instead measured (stage 0) and the result is qualified by what that measurement shows.
- "Retain Lane's editorial bracket, or run both Lane conditions" (MAJOR). Refused in part, and
the object of study is declared instead:
LANEhere is Lane's narrative text, not Lane as-published. The bracket is a translator speaking in his own voice about his policy — "as this is expressed in the original work in nearly the same words at the close of every night, such repetitions will in the present translation be omitted" — and it is quoted in full here and in the result page, so nothing is hidden. Running both conditions would cost 8 more bodies to answer a question about paratext that this arm is not asking.