Repository path: workshop/experiments/E-20260825-level-recovery/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260825-level-recovery |
| status | frozen |
| created | 2026-08-25 |
| updated | 2026-08-25 |
| senses | style-correspondence, consistency |
| links | wiki/arms/ARM-alf-layla.md, wiki/findings/results/RS-20260823b-embedding-carriage.md, workshop/translations/alf-layla/R05-v1/translation.md, workshop/translations/alf-layla/register.md, config/models.md, config/budget.md |
Can a reader recover who is speaking when the translation supplies nothing at the seam?
ARM-alf-layla step 9, study limb of span I. v2, frozen before any scored call is dispatched.
v1 was put through an independent adversarial pre-run critic (P1, P2) and came back
NEEDS-REDESIGN with 11 findings, 4 of them BLOCKING. Every finding is taken; §11 lists them and
says what each changed. The two most consequential changes: a fourth arm, because v1's
manipulation confounded four things at once, and a different primary test, because v1's was
underpowered and two-sided against a directional prediction. Nothing was dispatched under v1.
1. The question, and where it comes from
RS-20260823b-embedding-carriage §9, limit 3, states it and refuses to answer it:
No reader is measured. Whether an English reader can in fact hold three levels of telling without apparatus is a different question and needs readers this project does not have. It is named here and not answered, and nothing above may be read as an answer to it. The indirect design that could ask it — a content question whose right answer depends on knowing which level one is on — goes to the arm as a named step.
That census established what each of four English books puts on the page at the five seams of
the Nights' fifth night. Three of the four supply apparatus at the two seams the Arabic leaves bare;
the fourth is this project's own rendering, which supplies nothing anywhere, because V19 of the
binding register forbids it.
This experiment measures what that decision costs a reader. Not whether the rendering is good — the lead never judges its own translation — but whether the referring expressions in it can be resolved by someone reading it once.
What this unit teaches about translating literature (subject rule, wiki/tracks.md). It
measures whether a reader of an unmarked nested tale can recover who is speaking to whom, and how
much of any loss is bought back by each of the three interventions a translator actually has. The
instrument is a means; the finding is about a choice every translator of framed narrative has to
make.
2. Materials, and why they are already frozen
The passage is the lead's own English of the copy-text's fifth night, T-alf-layla-R05-v1
span H, frozen 2026-08-23 at 93f40ad1, from She said: It has come to me, O fortunate King,
that King Yunan said to his vizier to And King Yunan said: You have spoken truly, vizier. —
1,637 words, 10 paragraphs as published, stored at materials/A0_bare.txt.
The stimulus was written and frozen two sessions before this design existed, for reasons that had nothing to do with it. Nothing in the passage was chosen to make this experiment come out.
Its structure, which is what the task is about and which the prompt never mentions: Shahrazad tells the King; inside that the fisherman tells the ifrit the Tale of King Yunan and the Sage; inside that King Yunan tells his vizier the Tale of King Sindbad and the Falcon, and the vizier tells King Yunan the tale of the vizier who contrived against a king's son. Four levels; and the copy-text prints one heading in six pages, and it is a night's.
Two of the four seams in this passage are bare in the Arabic — the entry to the ghoul tale and
the exit from it — and both fall inside a single و-joined sentence. Those two are the ones the
arms manipulate. The other two (entry to and exit from the falcon tale) are marked in the Arabic by
the vizier's And how was that? and by the named closing formula This is what there was of the
story of King Sindbad, and are left alone in every arm.
3. The four arms — a dose of information at the seam
Four displays of the same words, differing only at the two bare seams. Files
materials/A{0,1,3,2}_marked.txt; the dose runs A0 → A1 → A3 → A2.
| dose | arm | what is at each bare seam | words | paras | SHA-256 (first 16) |
|---|---|---|---|---|---|
| 0 | A0 BARE |
nothing — V19, exactly the published wording and paragraphing |
1,653 | 10 | 45885d154e631347 |
| 1 | A1 BREAK |
a paragraph break. No words added or removed | 1,653 | 12 | a8e6a23c58edaa2b |
| 2 | A3 DIVIDE |
a paragraph break, a heading that says a division happens, and an inserted — he went on to say —. Nobody is named: the added matter contains no King, no Yunan, no vizier, no sage, no son, no falcon | 1,687 (+34) | 14 | 4f6b20b81402e586 |
| 3 | A2 NAME |
a paragraph break, a heading that names the tale, and an inserted — continued the vizier of King Yunan —. This is the apparatus three of four published hands supply at exactly these two seams | 1,683 (+30) | 14 | 13412f1330800854 |
A0 and A1 differ by one character in the whole passage — and → And, because a sentence
that opens a paragraph is capitalised — established by word-level diff and recorded here rather than
claimed away. A3 and A2 are matched to within four words of added matter and are typographically
identical in shape; the difference between them is whether the added matter names anyone.
A2's apparatus is modelled on published practice, not quoted from it. Lane, Burton and Forster
each supply a heading at the ghoul-tale entry and something at the exit; Lane's exit apparatus is an
illustration, a chapter heading in capitals and the inserted continued the Wezeer of King Yoonán.
The wording used here is the lead's own and no comparator page was opened to build it.
What each contrast estimates, registered:
A1−A0— typography alone, the cheapest intervention a translator has and the only one that adds nothing a reader can mistake for the source.A3−A1— saying that a division happens, without saying whose speech resumes.A2−A3— naming the speaker and the tale, holding the typography and the quantity of added matter fixed. This is the contrast v1 could not make.A2−A0— the composite, i.e. what the published habit buys overV19. It is reported and it is explicitly not described as a seam-marking effect.
4. The task, and why it is indirect
Each body is shown one arm and asked, for each of sixteen numbered expressions marked in the
text as {1} … {16}, which person it refers to, choosing from a supplied roster of twelve.
The words level, nested, tale within a tale, frame, embedding, structure, narrator,
depth and speaker do not appear in the prompt. It is a reference-resolution task and nothing
else. Note (brb) requires this: a direct identification task on a formal property saturates at
1.000 (RS-20260824c, RS-20260823b's predecessor). A reader who has lost the level answers these
questions wrongly without ever being asked about levels.
The roster (identical in all four arms, so it cannot produce an arm difference):
A. King Yunan G. that king's son
B. King Yunan's vizier H. the vizier who went out hunting with the king's son
C. the sage I. the woman at the head of the road
D. the king of the Persians J. the gazelle
E. the falcon K. Shahrazad
F. the king whose son was given to the hunt and the chase
L. the King who is listening to Shahrazad
Chance is 1/12 = 0.083. Three roster entries (I, K, L) are the answer to no locus and are
pure distractors.
The exact request payload is frozen in run.py and is identical across arms but for the
passage: no system message; one user message; temperature: 1; provider defaults for top_p; no
seed; max_tokens 2000 with reasoning.max_tokens 900 (note (bnk): the cap covers the
hidden reasoning as well as the answer). Model versions are the OpenRouter slugs in
config/models.md and are not pinnable to a snapshot; the provider field is read off every
response and stored.
The unit of inference is a stochastic draw from a seat's decoding distribution at temperature 1,
not a reader. Six draws per (arm × seat) cell. Whether those draws are actually distinct is
checked and reported, not assumed: verify.py counts distinct 16-letter answer vectors per cell,
and if a cell returns fewer than three distinct vectors the seat's contribution to the permutation
test is reported separately with that fact stated.
5. The sixteen loci
Frozen at materials/loci.json with the answer key, which was written from the Arabic
(../../translations/alf-layla/source-hindawi-2022-spanH.txt) and the copy-text's own pronoun
morphology, not from the English. The pre-run critic audited all sixteen answers against the
passage and found no wrong answer and no genuine ambiguity (finding 1); the audit is on the record
at raw/C_P1.json and is the reason the key is not being re-derived here.
Eight controls (C1–C8) — in text that is byte-identical in all four arms, upstream of
both manipulated seams: six inside the falcon tale, two in the outer Yunan/vizier exchange
immediately before the first manipulation.
Eight targeted (T1–T8) — downstream of a manipulated seam: five inside the ghoul tale
(where a reader who missed the entry maps them onto King Yunan's court) and three after the exit
(where a reader who missed it is still inside the ghoul tale).
| # | id | class | marked expression | correct |
|---|---|---|---|---|
| 1 | C1 |
C | Vizier, envy has entered into you** | B |
| 2 | C2 |
C | So the King made ready to go out | D |
| 3 | C3 |
C | she reared upon her two legs | J |
| 4 | C4 |
C | strike at her eyes until it blinded her | E |
| 5 | C5 |
C | God disappoint you, most ill-omened of birds | E |
| 6 | C6 |
C | seeing that it saved him from destruction | D |
| 7 | C7 |
C | the vizier heard King Yunan's words he said to him** | A |
| 8 | C8 |
C | and if you accept it from me you are saved | A |
| 9 | T1 |
T | and the King ordered that vizier to be with his son | F |
| 10 | T2 |
T | and his father's vizier went out with him | H |
| 11 | T3 |
T | and the vizier said to the king's son: Yours is this beast | H |
| 12 | T4 |
T | I have an enemy and I am afraid of him | G |
| 13 | T5 |
T | turned away to his father and told him the story of the vizier | F |
| 14 | T6 |
T | and you, O King, whenever you trust this sage | A |
| 15 | T7 |
T | this sage he will kill you the ugliest of killings | C |
| 16 | T8 |
T | It may be as you have said, counselling vizier** | B |
T5 and T6 are eleven words apart and sit on either side of the exit seam. T6 is the sharpest
cell in the passage and is the one RS-20260823b §3 singled out: the Arabic joins the end of the
inner tale to the resumption of the outer one with a single و, and three of four published hands
break the sentence there.
Two limits declared here rather than repaired. (i) Class is confounded with position: both
manipulated seams are in the second half, so every T locus is downstream of every C locus. The
design therefore makes no comparison between C and T within an arm. (ii) C and T are not
matched on syntactic type, antecedent distance or number of competing referents (critic finding 6);
the C set is not a test of seam-specificity and is not used as one — see Q2 in §7.
6. Panel, dispatch, and the arithmetic
Seats: P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, QR
qwen/qwen3.7-max — the three that carried S219's 226 calls with 0 dead and 0 unparsed. P4 and
P5 are out (notes (bps), (bne)); GL is out on long prompts; P3 is a cost problem (NEXT.md).
4 arms × 3 seats × 6 bodies = 72 calls, each returning 16 answers → 1,152 item-judgments,
288 per arm, 96 per arm × seat cell, 18 per arm × locus. Dispatch order is randomised under the
published seed DISPATCH_SEED = 20260825 (critic finding 4); every call is an independent request
and no context is shared between calls. Concurrency capped at 6 in one foreground process, note
(brf); no background dispatcher.
Pre-flight estimate. Prompt ≈ 2,350 tokens worst case (A3, longest); output capped at 2,000.
At config/models.md prices — P1 $1.00/$6.00, P2 $0.75/$3.75, QR $1.475/$4.425 per M — 24
calls each gives $0.344 + $0.222 + $0.294 = $0.860 worst case. The pre-run critic has already
been billed at $0.12079 (P1 $0.08784, P2 $0.03295).
Declared ceiling $1.40, raised from v1's $1.10 because the critic added a fourth arm that was not budgeted — the same reason and the same mechanism as S218's raise. Against a fresh $5.00 for UTC day 2026-08-25 (no rows before this session). Key-usage snapshot before the run: 141.368450912.
De-scope if the ceiling is approached: drop QR and report on two seats, which costs the
item-set sensitivity check and nothing registered.
7. Analysis plan and registered predictions
Estimand. For a given pair of arms, the difference in the expected number of the eight targeted loci answered correctly by one call, on this passage, for these three seats at temperature 1.
Primary test. The unit is a call; its score is the count of T loci correct, 0–8. Arms are
compared by an exact stratified permutation test, strata = seat, statistic = the difference in
stratum-weighted mean call score. The null distribution is computed by exhaustive enumeration
within each stratum (C(12,6) = 924 arrangements per seat) and exact convolution across the three
strata — not sampled. All tests on Q1, Q3, Q4, Q5 are one-sided, in the direction
registered below, because the predictions are directional; the one-sided P is reported with the
observed difference and an exact 90% interval from the same enumeration.
Intention to treat. Every call that yields sixteen parsable answers is scored, in whichever arm it was dispatched to (critic finding 7). No call is excluded on its control score. Control accuracy is a diagnostic and the basis of one pre-specified sensitivity analysis, reported beside the primary and never in place of it.
| # | prediction | decided by |
|---|---|---|
Q1 |
A2 beats A0 at the targeted loci — the composite apparatus buys something |
primary test, one-sided, P ≤ 0.05 |
Q2 |
The arms do not differ at the control loci. Registered as equivalence, not as a null result: it holds if the observed A2 − A0 difference in mean C-locus score is within ±0.80 of eight (0.10 per locus) and the exact 90% interval lies inside that margin |
same permutation machinery on C scores |
Q3 |
A1 beats A0 at the targeted loci — a paragraph break alone does the work |
primary test, one-sided |
Q4 |
A2 beats A3 at the targeted loci — naming the speaker buys something beyond marking the division |
primary test, one-sided |
Q5 |
A3 beats A1 — saying a division happens buys something beyond the break |
primary test, one-sided |
Q1 and Q3–Q5 are four tests on one family. Q1 is the primary and carries no correction;
Q3, Q4 and Q5 are secondary and are reported with a Holm correction over the three, and with
their uncorrected P values printed beside so a reader can see both.
Descriptive, registered as exploratory and not tested (critic finding 9):
D1— per-locus hit rate for each arm, all 16 loci × 4 arms, printed in full.D2— whetherT6shows the largestA2−A0gap of the eight.D3— the modal wrong answer at eachTlocus, reported whatever it is. The design's expectation, written down so it can fail: the referent one level out —AforT1/T5,BforT2/T3,AorCforT4, andF/G/HforT6–T8.
The honest alternative, registered as an outcome and not as a failure. If A0 is at or near
ceiling on the T loci, then no decrement is detectable on this instrument, for these calls, on
this passage — and Q1, Q3, Q4, Q5 cannot fire. That is a result and the result page will
lead with it. It is not a demonstration that human readers hold the levels unaided and it is not
a general vindication of V19; §9 governs the wording (critic finding 10).
8. Failure criteria, written before the run
- Instrument void if mean accuracy on the
Cloci inA0, pooled over seats, is below 0.60. Below that the bodies are not resolving reference and noTfigure means anything. - Ceiling void for
Q1,Q3–Q5if every arm scores 1.000 at everyTlocus. §7's honest alternative governs. - A call is dead if it does not yield 16 parsable answers, each a single letter A–L. Dead calls are re-dispatched once, in the same foreground process (note (brf)), and the count is reported. A call still dead after one repeat is dropped and counted; if more than 6 of 72 are dropped the run is reported as unreliable rather than patched.
- Draw-independence check: distinct answer vectors per cell, reported. A cell with fewer than three distinct vectors is flagged in the result and its seat reported separately.
- No number in the result page is copied from a run log.
verify.pyrecomputes every figure from the stored raw JSON, including the permutation nulls by exhaustive enumeration, and is run by a path that imports nothing fromanalysis.
9. What this design cannot establish
- The bodies are language models, not human readers. Everything here is a fact about three model
seats reading English prose at temperature 1.
NEXT.mdhas carried independent human readers as named-and-not-built since S211 and this does not change it. No sentence of the result page may say that "a reader" can or cannot do anything. - The task is artificial, and the critic is right about how (finding 5). Inline
{n}markers tell a body where to look; the roster's twelve entries distinguish several kings and viziers with unusual sharpness; the questions cluster in the second half. What is measured may be entity matching over a passage held whole in context rather than the maintenance of narrative level by a reader moving through it once. The fix — an unannotated reading stage followed by separately presented questions — is a different and more expensive experiment, and is named here and not built. - The controls are not seam-specificity controls (finding 6). The whole document, apparatus
included, is in context when every answer is produced, so downstream apparatus can affect upstream
answers.
Q2therefore bounds a general "this arm is easier" effect and nothing narrower. - One passage, one work, one translator. The rendering under test is the lead's and obeys
V19by construction. A2's apparatus is a reconstruction, so anA2effect is a fact about this apparatus.- A pure lexical placebo was not built: an arm with
A2's thirty added words at a non-seam location would separate "adding words helps" from "adding words there helps".A3is the nearest thing this design has and it still marks the seam. Named and not built. - The lead is not judged. No body is asked whether the prose is good. All four arms are the lead's own words, so every comparison is internal to one rendering.
A0is not the published page: it carries sixteen{n}markers. It is the published wording and paragraphing plus the experiment's annotation layer, and is described that way throughout.
10. Provenance of the answer key and the stimulus
materials/A0_bare.txt— extracted mechanically fromT-alf-layla-R05-v1at commit6a79c72c, which contains span H exactly as frozen at93f40ad1.materials/loci.json— the sixteen anchors, targets and answers, each anchor verified unique inA0_bare.txtbefore marking.materials/A{0,1,3,2}_marked.txt— built bymaterials/build.pyfromA0_bare.txtandloci.json; hashes in §3; every arm verified to carry each of{1}–{16}exactly once.raw/C_P1.json,raw/C_P2.json— the pre-run critic calls, billed $0.12079, kept whole.
11. The pre-run critic's findings, and what each changed
P1 returned a complete reply; P2's reply hit the max_tokens cap and is truncated — the
part that survives duplicates P1's finding 3 (the sign test is underpowered and mismatched to a
directional prediction) with its own worked arithmetic, and no further P2 finding is available.
Both files are kept. No second critic round was bought (note (bqp)).
| # | grade | finding | taken? |
|---|---|---|---|
| 1 | BLOCKING | The answer key audits as correct at all sixteen loci; have it audited before dispatch | Taken — the audit is the critic's; §5 records it and cites the file |
| 2 | BLOCKING | A2 confounds paragraphing, headings, thirty lexical words and an explicit attribution, so A2 − A0 cannot mean "seam marking" |
Taken — arm A3 added, and §3 registers what each contrast estimates. The critic's six-arm factorial is beyond this budget; the pure lexical placebo is named in §9 as not built |
| 3 | BLOCKING | An eight-locus two-sided sign test has almost no power and discards magnitude; 864 judgments are not 864 tests | Taken in full — the primary is now an exact stratified permutation test on call-level scores, one-sided, with the null enumerated |
| 4 | BLOCKING | "Bodies" are not defined as sampling units; decoding parameters, allocation and order are unstated | Taken — §4 freezes the payload and names the unit of inference; dispatch order randomised under a published seed; a draw-distinctness check is registered as a failure criterion |
| 5 | MAJOR | The task is not cleanly indirect: markers, roster and question placement make it entity-matching | Taken as a limit, not repaired — §9 states it in the critic's own terms and names the two-stage design that would fix it |
| 6 | MAJOR | The controls are not controls: the whole document is in context, and C and T are not matched |
Taken — Q2 is re-registered as bounding a general arm effect only; §5 and §9 state both halves |
| 7 | MAJOR | Excluding calls on control score is post-treatment selection | Taken in full — the primary is intention-to-treat; the exclusion rule is demoted to a sensitivity analysis |
| 8 | MAJOR | A1 was not shown to the critic; the 1,637 / 1,653 word counts are unreconciled |
Taken — §2 and §3 distinguish the published passage (1,637) from the annotated stimulus (1,653, the sixteen markers); SHA-256 for every arm in §3; all four arms are in materials/ |
| 9 | MAJOR | Directional claims tested two-sided; P2 treats non-significance as equivalence; P4/P5 guaranteed to find something |
Taken in full — one-sided tests; Q2 is an explicit equivalence claim with a stated margin and an interval; the old P4/P5 are relabelled D2/D3, exploratory and untested |
| 10 | MINOR | A ceiling result would not show that "a reader can hold the levels unaided" | Taken — §7's honest alternative and §9's first bullet are reworded |
| 11 | MINOR | A0 is not "exactly as published"; it carries the markers |
Taken — §3 and §9 say so |