Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260820-rhyme-bearer/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260820-rhyme-bearer
statusfrozen
created2026-08-20
updated2026-08-20
sensesstyle-correspondence, accuracy
linksworkshop/experiments/E-20260820-rhyme-bearer/predictions-frozen.md, workshop/experiments/E-20260820-rhyme-bearer/critic-response.md, workshop/translations/alf-layla/R05-v1/translation.md, workshop/translations/alf-layla/register.md, wiki/arms/ARM-alf-layla.md, wiki/findings/results/RS-20260814-saj-carriage.md

E-20260820-rhyme-bearer v3 — a descriptive audit of four English sentences

This is version 3. Versions 1 and 2 were killed by two independent pre-run critic passes before a single item call was made — both NEEDS-REDESIGN, eleven BLOCKING findings between them, all answered in critic-response.md. Two earlier designs were killed before those, and §0 says why. The second pass took the design's inference away and left a description, and §1 and §5 are what is left. Nothing in this file had been dispatched at the time of writing.

0. Three designs killed before dispatch, and the reasons

Design 1 — the fifteen-locus bearer census, registered in predictions-frozen.md at 1a6a2d3 before either comparator was opened. Abandoned on first sight of the comparators: Burton sets every quoted poem as rhymed English verse as a matter of format, so the verse cell is saturated by construction and measures his prosody rather than his answer to any Arabic rhyme; and neither published hand rhymes at any of span F's prose loci, so the prose cell has no variance. A design with one saturated cell and one empty cell cannot fail, so it cannot succeed. Registered predictions P1–P4 are reported as not run; P5 survives and binds this design.

Design 2 — a large blind rating of coordinated pairs across spans A–F. A mechanical extractor pulled 278 adjacent coordinated pairs and coded 17 as rhymed; inspection showed the 17 are mostly morphological accidents — verbs sharing an object suffix, not saj'. The fault is conceptual: Arabic prose rhyme falls at the ends of cola, which are phrases, and no adjacent-word rule can see a colon. A colon segmenter is instrument work (T6) and is never a session's principal unit; it is named in NEXT.md as the thing a larger version of this question needs first.

Design 3 — v1 of this file. Killed by the critic; see §7 and critic-response.md.

1. What v3 claims, after two critics took the rest away

The second critic pass (critic-v2.json) established that this sentence cannot support a controlled inference, and it is right. Its BLOCKING 4 is decisive: the three Arabic-rhymed pairs are couplet-internal (A–B, D–E, F–G), while the five pairs v2 called controls are transitions between couplets. They are not matched, and a NO at a transition would be unsurprising whatever the seats can hear. v2's claim that the five are "adjacent pairs of the same syntactic shape" was false on its own materials, and is withdrawn here rather than defended.

So v3 makes one claim and no inference:

At this catalogue, the assertion that Lane 1839, Burton 1885 and the lead's R05 produce no English rhyme at the two places where the lead found a rhyme available only by renaming is currently the lead's own ear, and nothing else. This run puts those two places, in those three sentences, to three model seats that are never told what is being tested, with a constructed rhyming version alongside to show whether the seats can hear a rhyme when one is present.

That is a small claim. It is worth the ten cents because it is the one sentence of the span-F log that a reader has no way to check and that the lead is least entitled to assert: S1's whole justification for refusing flails / sails and cave / grave rests on the published hands having refused them too, and the lead's ear is the only witness to that.

Everything else the run collects — the five transition pairs, the same-thing scores, the recognition line — is reported as an inventory of what the seats said. None of it is a control and none of it supports an inference.

2. The material

Hindawi 2022 vol. 1 p. 29, the ifrit's body, in nine parts. Three pairs rhyme in Arabic:

part Arabic literal gloss rhymes with
A رأسه في السحاب his head is in the cloud(s) B (saḥāb / turāb)
B ورجلاه في التراب and his two feet are in the dust A
C برأس كالقبة with a head like a dome —
D وأيدٍ كالمداري and hands like winnowing-forks E (al-madārī / al-ṣawārī)
E ورجلين كالصواري and two legs like ships' masts D
F وفَمٍ كالمغارة and a mouth like a cave G (al-maghāra / al-ḥijāra)
G وأسنان كالحجارة and teeth like stones F
H ومناخير كالإبريق and nostrils like a long-spouted ewer —
I وعينين كالسراجين and two eyes like two lamps —

Every one of the eight consecutive pairs is scored, uniformly — A–B, B–C, C–D, D–E, E–F, F–G, G–H, H–I. The item set is the whole sentence, in order, so nothing is selected by the lead. Three of the eight rhyme in Arabic. The other five are not controls — they are transitions between couplets, not couplet-internal pairs, and §1 withdraws v2's claim that they are matched. They are collected because collecting all eight costs nothing and leaving five out would be a selection.

3. The four renderings

Blind labels; the mapping is in materials/items.json and in no prompt.

label hand why it is in
H1 the lead's R05, span F, frozen at c2dd9bc the subject whose claim is being checked
H2 Burton 1885 (PG #3435), one sentence a published hand who reaches for sound elsewhere in this work
H4 Lane 1839 (PG #34206), one sentence a published hand who does not
H3 a decoy written by the lead for this run the coarse instrument control

The decoy carries an exact rhyme at each of the three Arabic-rhymed pairs, each bought by renaming the thing compared — sky / sty, flails / sails, cave / grave — and leaves the other five pairs plain.

4. The instrument check, and what it does and does not license

H3 carries an exact rhyme at each of the three Arabic-rhymed pairs, bought by renaming. If the seats do not score all three EXACT, the run says nothing at all and is reported as an instrument failure.

What passing it licenses, stated narrowly because the critic was right to press here. It shows the seats can detect a conspicuous monosyllabic exact rhyme between two final words. It does not show they can detect a weak or polysyllabic near-rhyme, so a NO from them is evidence about exact rhyme and is weak evidence about anything fainter. The one data point at the fainter end is H1's own clouds / clods, and it is not an independent calibration item — it is part of the thing in dispute. It is registered below as a prediction and reported as a single observation, not as a calibration.

A properly calibrated instrument would need a separate blinded set of pre-adjudicated EXACT, NEAR and NO word-pairs at several difficulties (critic v2, BLOCKING 1). That is not built here and is named in NEXT.md as part of what a real version of this question needs.

5. Registered predictions, v3, with the aggregation rule fixed first

Aggregation (critic v2, BLOCKING 3). A cell's value is the value two or more of the three seats give it. EXACT and NEAR are distinct values and are never merged for the purpose of finding a majority; where the question is only was any rhyme detected, EXACT and NEAR both count as detected and that is stated each time. A three-way split, or any cell missing a seat, is INDETERMINATE and is never read as NO. Every seat's raw answer is printed in the result.

Q1 — instrument. H3 is EXACT at A–B, D–E and F–G. Failure voids everything below.

Q2 — the primary, and it is two cells wide per hand. In H1, H2 and H4, the pairs D–E and F–G are NO — six cells. These are the two places where a rhyme was available only by renaming. This is the whole of the primary. It is a claim about final-word rhyme at two named places in three named sentences, under the prompt's stated definitions. It is not a claim that any hand fails to reproduce saj', nor that these sentences contain no sound figure of any kind (critic v2, BLOCKING 5).

Q3 — the one faint case, registered so it cannot be reinterpreted afterwards. H1's A–B (clouds / clods) is NEAR. If it comes back NO, the seats are less sensitive than the lead's log assumed and Q2's NOs mean correspondingly less; that consequence is registered now. If it comes back EXACT, the scale is being used more loosely than intended and Q2's NOs mean correspondingly more.

Q4 — inventory, not control. The five transition pairs and the two published hands' A–B are reported as observed. No prediction is registered on them and nothing is inferred from them.

Q5 — agreement with the supplied gloss, honestly named (critic v2, BLOCKING 2). SAME is not a test of fidelity to the Arabic; it is agreement with the lead's English gloss, which embeds contestable choices the critic listed (dust/earth, cloud(s), winnowing-forks against pitchforks, ships' masts against masts, long-spouted ewer against ewer). It is reported under that name and supports no fidelity claim. The lead's reading is that 20 of the 21 real-hand cells at parts C–I agree, the exception being H4's trumpets; the seats' figure is reported beside it as a second reading of the same three sentences.

Q6 — the decoy's price. H3 is SAME=NO at A, B, D, E, G and YES at C, F, H, I: five of nine parts renamed to buy three rhymes. Same caveat as Q5 — agreement with the gloss, not fidelity to the Arabic.

Q7 — recognition. Reported, never used to exclude. It is a weak measure and is labelled so: a seat that recognised a hand while scoring can still answer NO at the end, so a low recognition count is not evidence of blindness (critic v2, BLOCKING 3).

6. Procedure

One call per (rendering × seat), twelve in all; seats P1, P2, P3 (config/models.md), temperature 0, max_tokens 1500 against a reasoning cap of 250 (note bpv). Each seat answers eight pair-rhyme lines (EXACT / NEAR / NO, with the two final words named), nine same-thing lines, and one recognition line.

The rhyme question is directed at the last word of each part only, with alliteration, repetition of and / his / like, and any resemblance outside the pair explicitly excluded — the critic's BLOCKING 2. The seats are never told which Arabic parts rhyme, that Arabic rhyme is the subject, who wrote any rendering, or that one rendering was built.

7. What this is, and what it is not

It is a descriptive audit of four English sentences by three model seats, and nothing more.

What would make this a real experiment, written down because the next unit should not have to rediscover it: an Arabic colon segmenter, an independently authored semantic key from a reader of Arabic the project does not have, a blinded calibration set of pre-adjudicated rhyme pairs at several difficulties, and matched couplet-internal pairs from many sentences rather than three from one. Three of those four are instrument work.

8. Budget

Pre-flight from the caps the requests permit (note abc). Twelve calls, input ~1,100 tokens, output capped at 1,500: worst case ≈ $0.10, plus four worst-case re-dispatches at double cap ≈ $0.07. Dispatch worst case $0.17 < stop-loss $0.24 < ceiling $0.30, both enforced in run.py (note bpq, two-sided). The two critic passes cost $0.0749395 and $0.0829490 and sit outside those thresholds and inside the session's declared ceiling of $0.45. UTC day 2026-08-20 opened at $0.00 of $5.00.