Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260814g-supplied-sound/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260814g-supplied-sound
statusfrozen
created2026-08-14
updated2026-08-14
sensesstyle-correspondence
internal-judgment-onlytrue
linksworkshop/translations/alf-layla/R05-v1/translation.md, workshop/translations/alf-layla/collation.md, workshop/experiments/E-20260814-saj-carriage/design.md, wiki/findings/results/RS-20260814-saj-carriage.md, wiki/arms/ARM-alf-layla.md, config/models.md

E-20260814g — where the Arabic is plain, does the translator supply the sound anyway?

AMENDED 2026-08-14 on a pre-run adversarial critic pass that returned NEEDS REDESIGN with three BLOCKING findings, before any cell was dispatched. §10 lists all eleven findings and what was done with each. The largest consequence: P1–P5 are no longer registered predictions. They are pre-specified descriptive comparisons, explicitly non-confirmatory, and the confirmatory version of this question is handed to ARM-alf-layla step 3 with its loci to be frozen before the comparators are opened.

ARM-alf-layla step 2, study limb. The translation limb — span B, 582 Arabic words, log D15–D28 — was frozen at commit 4d4e6d0 before either comparator was opened for span B.

1. Question

RS-20260814 §6b handed this arm one unregistered observation and told it not to treat it as a finding. At the single locus in span A where the Arabic is plain and coarse — three neutral verbal nouns that do not rhyme — Burton wrote "kissing and clipping, coupling and carousing", an alliterating quadruplet. At five of seven loci where the Arabic is rhymed, he supplied nothing. One instance, one direction, unregistered.

Registered here: is a published translator's English sound-patterning a response to the source's sound-patterning, or a property of his own hand that runs regardless of what the Arabic is doing?

This is a question about translating literature, not about the project's instruments. If the second answer is right, then "carries the source's form" and "sounds ornate in English" are two different things that a reader cannot tell apart from the English alone — which is the seam four of this project's perceived-source-carriage figures already sit on (NEXT.md, S184), reached here from a completely different direction: a real published pair, a real formal device, and a source whose patterning is locatable to the word.

2. Design in one line

Sixteen loci in the Arabic, eight where the copy-text is patterned and eight where it is plain, classified before any comparator was opened; each hand's English at each locus extracted by anchor from a stored file; every passage put to three panel seats blind to the Arabic, to the hand, and to the domain, with one closed question about sound.

3. Materials

copy-text Hindawi 2022 vol. 1 pp. 11–15, workshop/translations/alf-layla/source-hindawi-2022-span{A,B}.txt
hand 1 Lane 1839, PG #34206 — materials/lane1839_spanB.txt and E-20260814/materials/lane1839_spanA.txt
hand 2 Burton 1885, PG #3435 — materials/burton1885_spanB.txt and E-20260814/materials/burton1885_spanA.txt
hand 3 the lead, T-alf-layla-R05-v1 — materials/lead_spanB.txt and E-20260814/materials/lead_spanA.txt. contamination: high; printed in italics; never evidence (charter §3, A4)
seats P1, P2, P3 per config/models.md. P5 is out on any task shape, note (bne). No Anthropic model sits on the panel and the lead never judges its own translation

4. The loci, and how they were fixed

PATTERNED (8). Loci where the copy-text repeats a sound.

PLAIN (8). Span B's prose was segmented mechanically on the copy-text's sentence punctuation (., ؟, !), giving 19 segments. A segment is eligible if it is 9–24 Arabic words, contains no PT locus, and contains no verse. The first eight eligible segments in copy-text order are PL1–PL8. No discretion is exercised at any point; cells.py recomputes the segmentation and the eligibility test from the stored source and asserts the eight identifiers, exiting non-zero if the rule and the page disagree.

Why a band and not a floor. The upper bound is the fix for a defect found while building the extraction and recorded here rather than quietly repaired: with a bare floor of 9 words the rule admits three segments of 41–132 Arabic words, and their English runs to 284 words in Burton against a PT mean of 33. That would have compared phrase-sized patterned passages against paragraph-sized plain ones, and a longer passage contains conspicuous sound more often for reasons that have nothing to do with translating. The band is set to the grain of the PT loci's containing clauses. It is stated before any passage was rated and it is outcome-independent: it is a fact about the copy-text, not about any hand's English.

The classification predates the reading of the comparators, and the selection is mechanical. Both facts are load-bearing, and one of them is a repair — see §9.1.

5. Extraction of the passages

For each (hand, locus) the passage is the contiguous English stretch that renders that Arabic segment, bounded by a start/end anchor pair, extracted by cells.py from the stored file. Nothing is transcribed by hand, so a mis-quotation is not possible; cells.py exits non-zero if any anchor fails to match or the anchors are out of order.

Where a hand interpolates matter of his own inside a segment, the interpolation is inside the passage. Burton's expansions are Burton's rendering of that segment and are not excised — excising them would be choosing the evidence. This makes passage length unequal between hands and the inequality is reported, not corrected (§7, §9.2).

absent is a fourth verdict, used where a hand does not render the segment at all. It is never counted as no patterning.

6. The instrument

Each passage goes to each seat as its own request. Judgment is not parallelised: bodies are dispatched one at a time. Each seat sees the passage and nothing else — no Arabic, no translator's name, no domain label, no neighbouring passage — and answers:

Below is a short passage of English prose. Answer one question about its sound only, not its meaning, accuracy or quality.

Does the passage use conspicuous sound-patterning — alliteration, internal or end rhyme, assonance used as a device, or jingling repetition of sound — such that a reader would notice the sound as a deliberate effect?

Reply with exactly one line: Y or N, then a semicolon, then in no more than 12 words name the pattern, or write none.

temperature: 0. Caps: P1 400, P2 1400, P3 800 completion tokens — set from the measured figures in NEXT.md (P1 80, P2 594, P3 397 on the last closed-form jury task, P2's hidden reasoning included). A body that does not return stop, or whose first character is not Y or N, is re-dispatched once and then recorded unparsed.

Ballot rules (critic finding 4). A ballot is valid if it returned stop and its first character is Y or N. A passage needs ≥ 2 valid ballots; with fewer it is invalid and is excluded from every rate and from the agreement denominator. Majority is taken over valid ballots only; a 1–1 split among two valid ballots is invalid. If more than 3 of the 48 dispatched passages are invalid, the instrument has failed.

All counts, once, so no figure drifts (critic finding 11). 16 loci × 3 hands = 48 cells, of which 3 are absent (Lane at PT8, PL5, PL6) → 45 live passages; + 3 controls = 48 dispatched passages; × 3 seats = 144 bodies. The agreement floor in §7.2 is over all 48 dispatched passages, controls included.

7. Pre-specified comparisons (NOT registered predictions — see §10.1)

Domain rates are per hand, over the 8 loci of each domain, absent excluded from both numerator and denominator. Every one of P1–P5 was formulated after the English it concerns had been read (§10.1). They are stated in advance of the rating and not in advance of the material; a comparison that comes out as stated is therefore not evidence that it was predicted, and is reported as a description. Proportions are reported with Wilson 95% intervals and no significance test.

Three analyses are pre-specified, and where they disagree the third governs: (a) all locatable cells; (b) dropping, per hand per domain, the single longest passage (ties by locus id, ascending); (c) restricted to passages of 13–45 words, the band in which both domains have cells for all three hands.

prediction resolves
power ≥ 14 of the 16 loci are locatable in both published hands MISSED BEFORE THE RUN, at 13 of 16, and recorded here rather than adjusted. Lane does not render PT8, PL6 or PL7 at all — the three loci around the maiden's demand and the brothers' hesitation, which is his expurgation and is discussed as such. Per-hand n: Burton 8 PATTERNED and 8 PLAIN; Lane 7 and 6. The primary is on Burton, whose cells are complete; P3 and P4 are reported on Lane's locatable subset and called underpowered
P1 (primary) Burton's PLAIN rate ≥ 0.50 — he supplies patterning where the Arabic has none, in at least half the plain loci
P2 (primary) Burton's PLAIN rate is not below his PATTERNED rate by more than 0.25 — his patterning does not track the source's this is the within-hand comparison and is the one immune to the between-hand length gap
P3 Lane's PLAIN rate < Burton's PLAIN rate
P4 Lane's PATTERNED rate ≤ 0.25 — RS-20260814 measured him at 0 of 7 carried, and this is the same claim through a different instrument, so a large disagreement is evidence the instrument is measuring something else
P5 (italic, not evidence) the lead's PATTERNED rate > the lead's PLAIN rate — a hand that was trying at the patterned loci and not at the plain ones
length Burton's mean PLAIN passage length is within 1.5× his mean PATTERNED passage length MISSED BEFORE THE RUN, at 54.0 against 33.2 = 1.63×, driven by a single passage (PL2, 143 words, where Burton interpolates a speech of his own). Lane is at 1.15× and the lead at 1.36×. Registered remedy, fixed before any rating: P2 is reported twice — over all loci, and again with each hand's single longest passage in each domain dropped. If the two disagree, the second governs

Failure criteria, registered.

  1. Instrument control. Two anchor passages are dispatched with the census and are not part of it: a POS passage written to be heavily alliterative and a NEG passage written to be flat, both authored by the lead for this purpose and both exhibited in cells.py. If POS is not coded Y unanimously or NEG is not coded N unanimously, the instrument has failed and no rate on this page is reported as a finding.
  2. Seat agreement. If the three seats are unanimous on fewer than 60% of the 50 bodies' passages, the coding is too noisy and P1–P4 are withheld with the figure stated.
  3. Recension. Any locus absent in a hand is reported as absent and reasoned about only as RS-20260814 §6a reasons — never read as an omission by the translator.

8. Cost

Pre-flight, built from the caps and not from expected lengths (note (abc)): 45 live passages (48 minus Lane's three absences) + 2 controls = 47 passages × 3 seats = 141 bodies. Prompt ≈ 320 tokens each. Worst case at the caps: P1 47 × (320 × $1 + 400 × $6)/1e6 = $0.128; P2 47 × (320 × $0.75 + 1400 × $3.75)/1e6 = $0.258; P3 47 × (320 × $2 + 800 × $6)/1e6 = $0.256. Plus the pre-run critic and one budgeted continuation (note (bnr)) at ≤ $0.12. Declared ceiling $0.77.

Today's UTC ledger stands at $2.920146 of $5.00 across five sessions; headroom $2.079854. The run fits with room. Lead translation is $0 and is not ledgered.

9. Limits, declared before the run

  1. This design was written after the comparators' span B had been read, and E-20260814's was not. That is a real weakening of the ordering discipline and it is named here rather than buried. The compensating controls are two: the patterned/plain classification of span B was frozen in the translator's log at 4d4e6d0, in a commit made before either comparator file was opened; and the plain loci are selected by a mechanical rule recomputed by cells.py, so no discretion survives between the reading and the census. What cannot be ruled out is that the criteria themselves were framed by what had been read — the criteria are D19 and D20, both pre-freeze, but the decision to use them as the classifier was made after. A later span that fixes the loci before opening the hands would close this, and step 3 of the arm should.
  2. Passage length differs systematically between hands, and within Burton between domains — Burton's span B is 1,747 words against Lane's 848 for the same 582 Arabic words. A longer passage has more room to contain patterning by chance. This is why P2, the within-hand comparison, is the primary and P3 is secondary; why the band in §4 exists; and why the length check fired before the run and bought a drop-the-longest re-analysis rather than a shrug.
  3. "Conspicuous sound-patterning" is a judgement, made here by three seats on a closed question, against RS-20260814's single-coder census. The two instruments overlap at Lane's patterned loci (P4) and that overlap is the only calibration available.
  4. n = 8 per domain per hand. A rate of 0.50 has a wide interval on eight trials and no significance test is registered or will be reported.
  5. The jury is not calibrated (config/models.md: Tier D not passed). Everything here is provisional and internal-judgment-only; no anchor is cited and none is claimed.
  6. Nothing here licenses a claim about good translation. senses: [style-correspondence].
  7. PATTERNED is confounded with register and speech-act type (critic finding 5). The PT cells are liturgical saj', epithet chains and marked rhetorical figures; the PL cells are narrative and dialogue. A hand that ornaments elevated passages rather than rhymed ones produces the same numbers. The analysis therefore reports three strata separately — saj' chains (PT1–PT6), cognate accusatives (PT7–PT8), and plain narrative/dialogue — and the page may not say "responds to the source's sound" where "responds to the source's elevation" would fit the same data.
  8. PT2 and PT3 are adjacent cola of one stretch of moralising in span A and may be one rhetorical unit split across two cells (critic finding 8). They are reported as two and the dependence is stated.
  9. The controls are lead-authored (critic finding 9), so a shared style between control and the lead's own cells cannot be ruled out. Third-party controls would be better and are owed if this instrument is used again.

10. The pre-run critic pass, and what was done with it

One non-Anthropic seat (P3, x-ai/grok-4.5), given the frozen design and the full passage table, asked to attack it. Raw output: critic.json. Verdict: NEEDS REDESIGN, 3 BLOCKING, 5 MAJOR, 3 MINOR. Cost $0.048448. Every finding and its disposition:

# severity finding disposition
1 BLOCKING The predictions were written after the English they concern had been read, so "registered prediction" is theatre ACCEPTED in full. P1–P5 demoted to pre-specified descriptive comparisons, labelled non-confirmatory throughout; the confirmatory design goes to arm step 3 with loci frozen before the comparators are opened
2 BLOCKING PLAIN was defined as the absence of a PT locus, which never establishes that the Arabic is plain ACCEPTED. An affirmative plainness audit now applies PT's own three criteria (saj', cognate accusative, near-synonym doublet) to every candidate segment. It excludes segment 2 (الصيد والقنص, a doublet), which is replaced by segment 15 under the same mechanical order rule. In cells.py as NOT_PLAIN, asserted
3 BLOCKING Footnote markers ([i_21], [FN#12]) were inside the rated strings, differ by hand, and are not prose ACCEPTED. Stripped in norm() before extraction; [Illustration] plates likewise
4 MAJOR Unparsed/partial ballots undefined; the agreement denominator said 50 where the census has 47 ACCEPTED. Ballot validity, the ≥2 rule, 1–1 ties and the invalid cap are now in §6; all counts consolidated in one place
5 MAJOR PATTERNED confounded with register and speech-act type ACCEPTED. Three strata reported separately; causal wording restricted (§9.7)
6 MAJOR The length confound survives the drop-longest remedy ACCEPTED. A third, length-band-restricted analysis (13–45 words) is pre-specified and governs where the three disagree
7 MAJOR The instrument was never calibrated against archaic-but-flat English, which is the register both hands write in ACCEPTED. Third control ARCH added, expected N unanimously, and a failure of it fails the instrument. The prompt now requires the pattern to recur and says ornate wording is not sound-patterning
8 MAJOR Passage grains are non-comparable; PT2/PT3 may be one unit split PARTLY ACCEPTED. The 9–24-word band already addresses the grain; the PT2/PT3 dependence is recorded as a limit (§9.8) rather than merged, because merging would silently revise a locus set frozen at 2c5b612
9 MINOR Lead-authored controls sit in the same pool as lead-translated cells ACCEPTED as a limit (§9.9); third-party controls owed if the instrument is reused
10 MINOR Thresholds are brittle and framed as decisive ACCEPTED. Wilson intervals; directional language only; replication required before the claim enters wiki/findings as more than a census
11 MINOR Cost and agreement arithmetic drift ACCEPTED, see §6

Nothing was overruled. The critic's finding 1 is the one that governs the result page: this run can describe what two published hands do, and it cannot claim to have predicted it.