Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260825-level-recovery/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260825-level-recovery
statusfrozen
created2026-08-25
updated2026-08-25
sensesstyle-correspondence, consistency
linkswiki/arms/ARM-alf-layla.md, wiki/findings/results/RS-20260823b-embedding-carriage.md, workshop/translations/alf-layla/R05-v1/translation.md, workshop/translations/alf-layla/register.md, config/models.md, config/budget.md

Can a reader recover who is speaking when the translation supplies nothing at the seam?

ARM-alf-layla step 9, study limb of span I. v2, frozen before any scored call is dispatched.

v1 was put through an independent adversarial pre-run critic (P1, P2) and came back NEEDS-REDESIGN with 11 findings, 4 of them BLOCKING. Every finding is taken; §11 lists them and says what each changed. The two most consequential changes: a fourth arm, because v1's manipulation confounded four things at once, and a different primary test, because v1's was underpowered and two-sided against a directional prediction. Nothing was dispatched under v1.

1. The question, and where it comes from

RS-20260823b-embedding-carriage §9, limit 3, states it and refuses to answer it:

No reader is measured. Whether an English reader can in fact hold three levels of telling without apparatus is a different question and needs readers this project does not have. It is named here and not answered, and nothing above may be read as an answer to it. The indirect design that could ask it — a content question whose right answer depends on knowing which level one is on — goes to the arm as a named step.

That census established what each of four English books puts on the page at the five seams of the Nights' fifth night. Three of the four supply apparatus at the two seams the Arabic leaves bare; the fourth is this project's own rendering, which supplies nothing anywhere, because V19 of the binding register forbids it.

This experiment measures what that decision costs a reader. Not whether the rendering is good — the lead never judges its own translation — but whether the referring expressions in it can be resolved by someone reading it once.

What this unit teaches about translating literature (subject rule, wiki/tracks.md). It measures whether a reader of an unmarked nested tale can recover who is speaking to whom, and how much of any loss is bought back by each of the three interventions a translator actually has. The instrument is a means; the finding is about a choice every translator of framed narrative has to make.

2. Materials, and why they are already frozen

The passage is the lead's own English of the copy-text's fifth night, T-alf-layla-R05-v1 span H, frozen 2026-08-23 at 93f40ad1, from She said: It has come to me, O fortunate King, that King Yunan said to his vizier to And King Yunan said: You have spoken truly, vizier. — 1,637 words, 10 paragraphs as published, stored at materials/A0_bare.txt.

The stimulus was written and frozen two sessions before this design existed, for reasons that had nothing to do with it. Nothing in the passage was chosen to make this experiment come out.

Its structure, which is what the task is about and which the prompt never mentions: Shahrazad tells the King; inside that the fisherman tells the ifrit the Tale of King Yunan and the Sage; inside that King Yunan tells his vizier the Tale of King Sindbad and the Falcon, and the vizier tells King Yunan the tale of the vizier who contrived against a king's son. Four levels; and the copy-text prints one heading in six pages, and it is a night's.

Two of the four seams in this passage are bare in the Arabic — the entry to the ghoul tale and the exit from it — and both fall inside a single و-joined sentence. Those two are the ones the arms manipulate. The other two (entry to and exit from the falcon tale) are marked in the Arabic by the vizier's And how was that? and by the named closing formula This is what there was of the story of King Sindbad, and are left alone in every arm.

3. The four arms — a dose of information at the seam

Four displays of the same words, differing only at the two bare seams. Files materials/A{0,1,3,2}_marked.txt; the dose runs A0 → A1 → A3 → A2.

dose arm what is at each bare seam words paras SHA-256 (first 16)
0 A0 BARE nothing — V19, exactly the published wording and paragraphing 1,653 10 45885d154e631347
1 A1 BREAK a paragraph break. No words added or removed 1,653 12 a8e6a23c58edaa2b
2 A3 DIVIDE a paragraph break, a heading that says a division happens, and an inserted — he went on to say —. Nobody is named: the added matter contains no King, no Yunan, no vizier, no sage, no son, no falcon 1,687 (+34) 14 4f6b20b81402e586
3 A2 NAME a paragraph break, a heading that names the tale, and an inserted — continued the vizier of King Yunan —. This is the apparatus three of four published hands supply at exactly these two seams 1,683 (+30) 14 13412f1330800854

A0 and A1 differ by one character in the whole passage — and → And, because a sentence that opens a paragraph is capitalised — established by word-level diff and recorded here rather than claimed away. A3 and A2 are matched to within four words of added matter and are typographically identical in shape; the difference between them is whether the added matter names anyone.

A2's apparatus is modelled on published practice, not quoted from it. Lane, Burton and Forster each supply a heading at the ghoul-tale entry and something at the exit; Lane's exit apparatus is an illustration, a chapter heading in capitals and the inserted continued the Wezeer of King Yoonán. The wording used here is the lead's own and no comparator page was opened to build it.

What each contrast estimates, registered:

4. The task, and why it is indirect

Each body is shown one arm and asked, for each of sixteen numbered expressions marked in the text as {1} … {16}, which person it refers to, choosing from a supplied roster of twelve.

The words level, nested, tale within a tale, frame, embedding, structure, narrator, depth and speaker do not appear in the prompt. It is a reference-resolution task and nothing else. Note (brb) requires this: a direct identification task on a formal property saturates at 1.000 (RS-20260824c, RS-20260823b's predecessor). A reader who has lost the level answers these questions wrongly without ever being asked about levels.

The roster (identical in all four arms, so it cannot produce an arm difference):

A. King Yunan            G. that king's son
B. King Yunan's vizier   H. the vizier who went out hunting with the king's son
C. the sage              I. the woman at the head of the road
D. the king of the Persians   J. the gazelle
E. the falcon            K. Shahrazad
F. the king whose son was given to the hunt and the chase
                         L. the King who is listening to Shahrazad

Chance is 1/12 = 0.083. Three roster entries (I, K, L) are the answer to no locus and are pure distractors.

The exact request payload is frozen in run.py and is identical across arms but for the passage: no system message; one user message; temperature: 1; provider defaults for top_p; no seed; max_tokens 2000 with reasoning.max_tokens 900 (note (bnk): the cap covers the hidden reasoning as well as the answer). Model versions are the OpenRouter slugs in config/models.md and are not pinnable to a snapshot; the provider field is read off every response and stored.

The unit of inference is a stochastic draw from a seat's decoding distribution at temperature 1, not a reader. Six draws per (arm × seat) cell. Whether those draws are actually distinct is checked and reported, not assumed: verify.py counts distinct 16-letter answer vectors per cell, and if a cell returns fewer than three distinct vectors the seat's contribution to the permutation test is reported separately with that fact stated.

5. The sixteen loci

Frozen at materials/loci.json with the answer key, which was written from the Arabic (../../translations/alf-layla/source-hindawi-2022-spanH.txt) and the copy-text's own pronoun morphology, not from the English. The pre-run critic audited all sixteen answers against the passage and found no wrong answer and no genuine ambiguity (finding 1); the audit is on the record at raw/C_P1.json and is the reason the key is not being re-derived here.

Eight controls (C1–C8) — in text that is byte-identical in all four arms, upstream of both manipulated seams: six inside the falcon tale, two in the outer Yunan/vizier exchange immediately before the first manipulation.

Eight targeted (T1–T8) — downstream of a manipulated seam: five inside the ghoul tale (where a reader who missed the entry maps them onto King Yunan's court) and three after the exit (where a reader who missed it is still inside the ghoul tale).

# id class marked expression correct
1 C1 C Vizier, envy has entered into you** B
2 C2 C So the King made ready to go out D
3 C3 C she reared upon her two legs J
4 C4 C strike at her eyes until it blinded her E
5 C5 C God disappoint you, most ill-omened of birds E
6 C6 C seeing that it saved him from destruction D
7 C7 C the vizier heard King Yunan's words he said to him** A
8 C8 C and if you accept it from me you are saved A
9 T1 T and the King ordered that vizier to be with his son F
10 T2 T and his father's vizier went out with him H
11 T3 T and the vizier said to the king's son: Yours is this beast H
12 T4 T I have an enemy and I am afraid of him G
13 T5 T turned away to his father and told him the story of the vizier F
14 T6 T and you, O King, whenever you trust this sage A
15 T7 T this sage he will kill you the ugliest of killings C
16 T8 T It may be as you have said, counselling vizier** B

T5 and T6 are eleven words apart and sit on either side of the exit seam. T6 is the sharpest cell in the passage and is the one RS-20260823b §3 singled out: the Arabic joins the end of the inner tale to the resumption of the outer one with a single و, and three of four published hands break the sentence there.

Two limits declared here rather than repaired. (i) Class is confounded with position: both manipulated seams are in the second half, so every T locus is downstream of every C locus. The design therefore makes no comparison between C and T within an arm. (ii) C and T are not matched on syntactic type, antecedent distance or number of competing referents (critic finding 6); the C set is not a test of seam-specificity and is not used as one — see Q2 in §7.

6. Panel, dispatch, and the arithmetic

Seats: P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, QR qwen/qwen3.7-max — the three that carried S219's 226 calls with 0 dead and 0 unparsed. P4 and P5 are out (notes (bps), (bne)); GL is out on long prompts; P3 is a cost problem (NEXT.md).

4 arms × 3 seats × 6 bodies = 72 calls, each returning 16 answers → 1,152 item-judgments, 288 per arm, 96 per arm × seat cell, 18 per arm × locus. Dispatch order is randomised under the published seed DISPATCH_SEED = 20260825 (critic finding 4); every call is an independent request and no context is shared between calls. Concurrency capped at 6 in one foreground process, note (brf); no background dispatcher.

Pre-flight estimate. Prompt ≈ 2,350 tokens worst case (A3, longest); output capped at 2,000. At config/models.md prices — P1 $1.00/$6.00, P2 $0.75/$3.75, QR $1.475/$4.425 per M — 24 calls each gives $0.344 + $0.222 + $0.294 = $0.860 worst case. The pre-run critic has already been billed at $0.12079 (P1 $0.08784, P2 $0.03295).

Declared ceiling $1.40, raised from v1's $1.10 because the critic added a fourth arm that was not budgeted — the same reason and the same mechanism as S218's raise. Against a fresh $5.00 for UTC day 2026-08-25 (no rows before this session). Key-usage snapshot before the run: 141.368450912.

De-scope if the ceiling is approached: drop QR and report on two seats, which costs the item-set sensitivity check and nothing registered.

7. Analysis plan and registered predictions

Estimand. For a given pair of arms, the difference in the expected number of the eight targeted loci answered correctly by one call, on this passage, for these three seats at temperature 1.

Primary test. The unit is a call; its score is the count of T loci correct, 0–8. Arms are compared by an exact stratified permutation test, strata = seat, statistic = the difference in stratum-weighted mean call score. The null distribution is computed by exhaustive enumeration within each stratum (C(12,6) = 924 arrangements per seat) and exact convolution across the three strata — not sampled. All tests on Q1, Q3, Q4, Q5 are one-sided, in the direction registered below, because the predictions are directional; the one-sided P is reported with the observed difference and an exact 90% interval from the same enumeration.

Intention to treat. Every call that yields sixteen parsable answers is scored, in whichever arm it was dispatched to (critic finding 7). No call is excluded on its control score. Control accuracy is a diagnostic and the basis of one pre-specified sensitivity analysis, reported beside the primary and never in place of it.

# prediction decided by
Q1 A2 beats A0 at the targeted loci — the composite apparatus buys something primary test, one-sided, P ≤ 0.05
Q2 The arms do not differ at the control loci. Registered as equivalence, not as a null result: it holds if the observed A2 − A0 difference in mean C-locus score is within ±0.80 of eight (0.10 per locus) and the exact 90% interval lies inside that margin same permutation machinery on C scores
Q3 A1 beats A0 at the targeted loci — a paragraph break alone does the work primary test, one-sided
Q4 A2 beats A3 at the targeted loci — naming the speaker buys something beyond marking the division primary test, one-sided
Q5 A3 beats A1 — saying a division happens buys something beyond the break primary test, one-sided

Q1 and Q3–Q5 are four tests on one family. Q1 is the primary and carries no correction; Q3, Q4 and Q5 are secondary and are reported with a Holm correction over the three, and with their uncorrected P values printed beside so a reader can see both.

Descriptive, registered as exploratory and not tested (critic finding 9):

The honest alternative, registered as an outcome and not as a failure. If A0 is at or near ceiling on the T loci, then no decrement is detectable on this instrument, for these calls, on this passage — and Q1, Q3, Q4, Q5 cannot fire. That is a result and the result page will lead with it. It is not a demonstration that human readers hold the levels unaided and it is not a general vindication of V19; §9 governs the wording (critic finding 10).

8. Failure criteria, written before the run

  1. Instrument void if mean accuracy on the C loci in A0, pooled over seats, is below 0.60. Below that the bodies are not resolving reference and no T figure means anything.
  2. Ceiling void for Q1, Q3–Q5 if every arm scores 1.000 at every T locus. §7's honest alternative governs.
  3. A call is dead if it does not yield 16 parsable answers, each a single letter A–L. Dead calls are re-dispatched once, in the same foreground process (note (brf)), and the count is reported. A call still dead after one repeat is dropped and counted; if more than 6 of 72 are dropped the run is reported as unreliable rather than patched.
  4. Draw-independence check: distinct answer vectors per cell, reported. A cell with fewer than three distinct vectors is flagged in the result and its seat reported separately.
  5. No number in the result page is copied from a run log. verify.py recomputes every figure from the stored raw JSON, including the permutation nulls by exhaustive enumeration, and is run by a path that imports nothing from analysis.

9. What this design cannot establish

10. Provenance of the answer key and the stimulus

11. The pre-run critic's findings, and what each changed

P1 returned a complete reply; P2's reply hit the max_tokens cap and is truncated — the part that survives duplicates P1's finding 3 (the sign test is underpowered and mismatched to a directional prediction) with its own worked arithmetic, and no further P2 finding is available. Both files are kept. No second critic round was bought (note (bqp)).

# grade finding taken?
1 BLOCKING The answer key audits as correct at all sixteen loci; have it audited before dispatch Taken — the audit is the critic's; §5 records it and cites the file
2 BLOCKING A2 confounds paragraphing, headings, thirty lexical words and an explicit attribution, so A2 − A0 cannot mean "seam marking" Taken — arm A3 added, and §3 registers what each contrast estimates. The critic's six-arm factorial is beyond this budget; the pure lexical placebo is named in §9 as not built
3 BLOCKING An eight-locus two-sided sign test has almost no power and discards magnitude; 864 judgments are not 864 tests Taken in full — the primary is now an exact stratified permutation test on call-level scores, one-sided, with the null enumerated
4 BLOCKING "Bodies" are not defined as sampling units; decoding parameters, allocation and order are unstated Taken — §4 freezes the payload and names the unit of inference; dispatch order randomised under a published seed; a draw-distinctness check is registered as a failure criterion
5 MAJOR The task is not cleanly indirect: markers, roster and question placement make it entity-matching Taken as a limit, not repaired — §9 states it in the critic's own terms and names the two-stage design that would fix it
6 MAJOR The controls are not controls: the whole document is in context, and C and T are not matched Taken — Q2 is re-registered as bounding a general arm effect only; §5 and §9 state both halves
7 MAJOR Excluding calls on control score is post-treatment selection Taken in full — the primary is intention-to-treat; the exclusion rule is demoted to a sensitivity analysis
8 MAJOR A1 was not shown to the critic; the 1,637 / 1,653 word counts are unreconciled Taken — §2 and §3 distinguish the published passage (1,637) from the annotated stimulus (1,653, the sixteen markers); SHA-256 for every arm in §3; all four arms are in materials/
9 MAJOR Directional claims tested two-sided; P2 treats non-significance as equivalence; P4/P5 guaranteed to find something Taken in full — one-sided tests; Q2 is an explicit equivalence claim with a stated margin and an interval; the old P4/P5 are relabelled D2/D3, exploratory and untested
10 MINOR A ceiling result would not show that "a reader can hold the levels unaided" Taken — §7's honest alternative and §9's first bullet are reworded
11 MINOR A0 is not "exactly as published"; it carries the markers Taken — §3 and §9 say so