Repository path: workshop/experiments/E-20260820-rhyme-bearer/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260820-rhyme-bearer |
| status | frozen |
| created | 2026-08-20 |
| updated | 2026-08-20 |
| senses | style-correspondence, accuracy |
| links | workshop/experiments/E-20260820-rhyme-bearer/predictions-frozen.md, workshop/experiments/E-20260820-rhyme-bearer/critic-response.md, workshop/translations/alf-layla/R05-v1/translation.md, workshop/translations/alf-layla/register.md, wiki/arms/ARM-alf-layla.md, wiki/findings/results/RS-20260814-saj-carriage.md |
E-20260820-rhyme-bearer v3 — a descriptive audit of four English sentences
This is version 3. Versions 1 and 2 were killed by two independent pre-run critic passes before a
single item call was made — both NEEDS-REDESIGN, eleven BLOCKING findings between them, all answered in
critic-response.md. Two earlier designs were killed before those, and §0 says why. The second
pass took the design's inference away and left a description, and §1 and §5 are what is left.
Nothing in this file had been dispatched at the time of writing.
0. Three designs killed before dispatch, and the reasons
Design 1 — the fifteen-locus bearer census, registered in predictions-frozen.md at 1a6a2d3
before either comparator was opened. Abandoned on first sight of the comparators: Burton sets every
quoted poem as rhymed English verse as a matter of format, so the verse cell is saturated by
construction and measures his prosody rather than his answer to any Arabic rhyme; and neither
published hand rhymes at any of span F's prose loci, so the prose cell has no variance. A design with
one saturated cell and one empty cell cannot fail, so it cannot succeed. Registered predictions
P1–P4 are reported as not run; P5 survives and binds this design.
Design 2 — a large blind rating of coordinated pairs across spans A–F. A mechanical extractor
pulled 278 adjacent coordinated pairs and coded 17 as rhymed; inspection showed the 17 are mostly
morphological accidents — verbs sharing an object suffix, not saj'. The fault is conceptual:
Arabic prose rhyme falls at the ends of cola, which are phrases, and no adjacent-word rule can see
a colon. A colon segmenter is instrument work (T6) and is never a session's principal unit; it is
named in NEXT.md as the thing a larger version of this question needs first.
Design 3 — v1 of this file. Killed by the critic; see §7 and critic-response.md.
1. What v3 claims, after two critics took the rest away
The second critic pass (critic-v2.json) established that this sentence cannot support a
controlled inference, and it is right. Its BLOCKING 4 is decisive: the three Arabic-rhymed pairs
are couplet-internal (A–B, D–E, F–G), while the five pairs v2 called controls are
transitions between couplets. They are not matched, and a NO at a transition would be
unsurprising whatever the seats can hear. v2's claim that the five are "adjacent pairs of the same
syntactic shape" was false on its own materials, and is withdrawn here rather than defended.
So v3 makes one claim and no inference:
At this catalogue, the assertion that Lane 1839, Burton 1885 and the lead's
R05produce no English rhyme at the two places where the lead found a rhyme available only by renaming is currently the lead's own ear, and nothing else. This run puts those two places, in those three sentences, to three model seats that are never told what is being tested, with a constructed rhyming version alongside to show whether the seats can hear a rhyme when one is present.
That is a small claim. It is worth the ten cents because it is the one sentence of the span-F log
that a reader has no way to check and that the lead is least entitled to assert: S1's whole
justification for refusing flails / sails and cave / grave rests on the published hands
having refused them too, and the lead's ear is the only witness to that.
Everything else the run collects — the five transition pairs, the same-thing scores, the recognition line — is reported as an inventory of what the seats said. None of it is a control and none of it supports an inference.
2. The material
Hindawi 2022 vol. 1 p. 29, the ifrit's body, in nine parts. Three pairs rhyme in Arabic:
| part | Arabic | literal gloss | rhymes with |
|---|---|---|---|
| A | رأسه في السحاب |
his head is in the cloud(s) | B (saḥāb / turāb) |
| B | ورجلاه في التراب |
and his two feet are in the dust | A |
| C | برأس كالقبة |
with a head like a dome | — |
| D | وأيدٍ كالمداري |
and hands like winnowing-forks | E (al-madārī / al-ṣawārī) |
| E | ورجلين كالصواري |
and two legs like ships' masts | D |
| F | وفَمٍ كالمغارة |
and a mouth like a cave | G (al-maghāra / al-ḥijāra) |
| G | وأسنان كالحجارة |
and teeth like stones | F |
| H | ومناخير كالإبريق |
and nostrils like a long-spouted ewer | — |
| I | وعينين كالسراجين |
and two eyes like two lamps | — |
Every one of the eight consecutive pairs is scored, uniformly — A–B, B–C, C–D, D–E,
E–F, F–G, G–H, H–I. The item set is the whole sentence, in order, so nothing is selected by
the lead. Three of the eight rhyme in Arabic. The other five are not controls — they are
transitions between couplets, not couplet-internal pairs, and §1 withdraws v2's claim that they are
matched. They are collected because collecting all eight costs nothing and leaving five out would be
a selection.
3. The four renderings
Blind labels; the mapping is in materials/items.json and in no prompt.
| label | hand | why it is in |
|---|---|---|
H1 |
the lead's R05, span F, frozen at c2dd9bc |
the subject whose claim is being checked |
H2 |
Burton 1885 (PG #3435), one sentence | a published hand who reaches for sound elsewhere in this work |
H4 |
Lane 1839 (PG #34206), one sentence | a published hand who does not |
H3 |
a decoy written by the lead for this run | the coarse instrument control |
The decoy carries an exact rhyme at each of the three Arabic-rhymed pairs, each bought by renaming the thing compared — sky / sty, flails / sails, cave / grave — and leaves the other five pairs plain.
4. The instrument check, and what it does and does not license
H3 carries an exact rhyme at each of the three Arabic-rhymed pairs, bought by renaming.
If the seats do not score all three EXACT, the run says nothing at all and is reported as an
instrument failure.
What passing it licenses, stated narrowly because the critic was right to press here. It shows
the seats can detect a conspicuous monosyllabic exact rhyme between two final words. It does not
show they can detect a weak or polysyllabic near-rhyme, so a NO from them is evidence about
exact rhyme and is weak evidence about anything fainter. The one data point at the fainter end is
H1's own clouds / clods, and it is not an independent calibration item — it is part of the
thing in dispute. It is registered below as a prediction and reported as a single observation, not
as a calibration.
A properly calibrated instrument would need a separate blinded set of pre-adjudicated EXACT,
NEAR and NO word-pairs at several difficulties (critic v2, BLOCKING 1). That is not built here
and is named in NEXT.md as part of what a real version of this question needs.
5. Registered predictions, v3, with the aggregation rule fixed first
Aggregation (critic v2, BLOCKING 3). A cell's value is the value two or more of the three
seats give it. EXACT and NEAR are distinct values and are never merged for the purpose of
finding a majority; where the question is only was any rhyme detected, EXACT and NEAR both
count as detected and that is stated each time. A three-way split, or any cell missing a seat, is
INDETERMINATE and is never read as NO. Every seat's raw answer is printed in the result.
Q1 — instrument. H3 is EXACT at A–B, D–E and F–G. Failure voids everything below.
Q2 — the primary, and it is two cells wide per hand. In H1, H2 and H4, the pairs D–E and
F–G are NO — six cells. These are the two places where a rhyme was available only by
renaming. This is the whole of the primary. It is a claim about final-word rhyme at two named
places in three named sentences, under the prompt's stated definitions. It is not a claim that any
hand fails to reproduce saj', nor that these sentences contain no sound figure of any kind
(critic v2, BLOCKING 5).
Q3 — the one faint case, registered so it cannot be reinterpreted afterwards. H1's A–B
(clouds / clods) is NEAR. If it comes back NO, the seats are less sensitive than the lead's
log assumed and Q2's NOs mean correspondingly less; that consequence is registered now. If it
comes back EXACT, the scale is being used more loosely than intended and Q2's NOs mean
correspondingly more.
Q4 — inventory, not control. The five transition pairs and the two published hands' A–B are
reported as observed. No prediction is registered on them and nothing is inferred from them.
Q5 — agreement with the supplied gloss, honestly named (critic v2, BLOCKING 2). SAME is
not a test of fidelity to the Arabic; it is agreement with the lead's English gloss, which
embeds contestable choices the critic listed (dust/earth, cloud(s), winnowing-forks against
pitchforks, ships' masts against masts, long-spouted ewer against ewer). It is reported
under that name and supports no fidelity claim. The lead's reading is that 20 of the 21 real-hand
cells at parts C–I agree, the exception being H4's trumpets; the seats' figure is reported
beside it as a second reading of the same three sentences.
Q6 — the decoy's price. H3 is SAME=NO at A, B, D, E, G and YES at C, F, H,
I: five of nine parts renamed to buy three rhymes. Same caveat as Q5 — agreement with the
gloss, not fidelity to the Arabic.
Q7 — recognition. Reported, never used to exclude. It is a weak measure and is labelled so:
a seat that recognised a hand while scoring can still answer NO at the end, so a low recognition
count is not evidence of blindness (critic v2, BLOCKING 3).
6. Procedure
One call per (rendering × seat), twelve in all; seats P1, P2, P3 (config/models.md),
temperature 0, max_tokens 1500 against a reasoning cap of 250 (note bpv). Each seat answers
eight pair-rhyme lines (EXACT / NEAR / NO, with the two final words named), nine same-thing
lines, and one recognition line.
The rhyme question is directed at the last word of each part only, with alliteration, repetition of and / his / like, and any resemblance outside the pair explicitly excluded — the critic's BLOCKING 2. The seats are never told which Arabic parts rhyme, that Arabic rhyme is the subject, who wrote any rendering, or that one rendering was built.
7. What this is, and what it is not
It is a descriptive audit of four English sentences by three model seats, and nothing more.
- Not blind readers. Three temperature-0 outputs from proprietary systems with heavily overlapping training data, prompted identically. A majority among them is not replication and is never reported as corroboration (critic v2, BLOCKING 3).
- Not a control design. The five transition pairs are an inventory. v2's control claim is withdrawn (BLOCKING 4).
- Not a fidelity measure.
SAMEis agreement with the lead's gloss (BLOCKING 2). - Not a test of saj' carriage. Final-word rhyme at two named places (BLOCKING 5).
- Two relations per hand, in one sentence. No significance test is reported and none would be honest.
P5binds:H1is the lead's own text, written underS1; its scores are reported and are never evidence for the lead's claim.- Missing data: any cell short of a two-seat majority is
INDETERMINATE, neverNO. A partial run is descriptive only and is labelled so.
What would make this a real experiment, written down because the next unit should not have to rediscover it: an Arabic colon segmenter, an independently authored semantic key from a reader of Arabic the project does not have, a blinded calibration set of pre-adjudicated rhyme pairs at several difficulties, and matched couplet-internal pairs from many sentences rather than three from one. Three of those four are instrument work.
8. Budget
Pre-flight from the caps the requests permit (note abc). Twelve calls, input ~1,100 tokens,
output capped at 1,500: worst case ≈ $0.10, plus four worst-case re-dispatches at double cap
≈ $0.07. Dispatch worst case $0.17 < stop-loss $0.24 < ceiling $0.30, both enforced in
run.py (note bpq, two-sided). The two critic passes cost $0.0749395 and $0.0829490 and sit
outside those thresholds and inside the session's declared ceiling of $0.45. UTC day 2026-08-20 opened at $0.00 of $5.00.