Repository path: workshop/experiments/E-20260823-run-depth/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260823-run-depth |
| status | frozen |
| created | 2026-08-23 |
| updated | 2026-08-23 |
| senses | style-correspondence |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-run-depth.md, workshop/regimes/R44-run-depth.md, workshop/translations/gulistan-bab1b/passages-frozen.md, workshop/translations/gulistan-bab1b/collation.md, wiki/findings/results/RS-20260822b-echo-threshold.md, wiki/findings/results/RS-20260822c-persian-hands.md, config/models.md, config/budget.md, tools/rhyme_pairs.py |
E-20260823-run-depth — the same rhyme at two distances
Frozen before any English rendering exists, before any published hand is opened for this span, and
before any API call is dispatched. ARM-run-depth step 1, track T2.
1. The question
Sa'di uses two verse forms inside the same tale. The مثنوی bayt rhymes with itself: in English, with one line to a hemistich, that is a rhyme at adjacent line-ends. The قطعه holds one rhyme across the ends of three or four consecutive bayts and leaves the lines between them unrhymed: in English that is a rhyme at every second line-end. The published English tradition replaces the second with the first — it rhymes each bayt within itself — which supplies more rhyme than the Persian has and puts none of it where the Persian put it.
Is a rhyme held at every second line-end registered by a reader at the rate an adjacent rhyme is? And at the places where Sa'di holds one rhyme three and four bayts deep, does any published English hand hold it at all?
What this is not. It is not a judgment of any rendering: no seat is asked whether anything is
good. It is not a claim about human readers; the seats are models and every figure is a figure about
models reading. It is not a test of tools/rhyme_pairs.py, which supplies the ground truth by a rule
fixed on 2026-08-22 and is not under test here.
2. Why it is worth a session
wiki/goodness-senses.md §style-correspondence is evidenced on device presence (+3.524 for
sixteen formal devices carried rather than flattened) and has no evidence about device extent.
framework/v0.2 §7.26 measured that a chime is registered at +0.542 at a verse line-end and
+0.222 at a prose colon-end — a difference of position. Nothing in the project has measured a
difference of distance, and run length is the property R42's own log flagged and R42 rule 4
forbade the hand from attempting.
3. Materials
Source. «گلستان» باب اول حکایات ۶–۱۳, workshop/translations/gulistan-bab1b/source-ganjoor-bab1-h6-13.txt,
59 blocks, 19 prose and 40 bayts. Copy-text and its single-witness limit:
workshop/translations/gulistan-bab1b/collation.md.
Structure, frozen from the Persian before any English existed.
workshop/translations/gulistan-bab1b/passages-frozen.md — 21 passages, 17 adjacent rhymes and
14 distance links across 9 قطعه runs, two of them four deep. Exclusions E1–E3 are on that
page and are not revisited here.
Contamination, declared in advance and high. All four published hands are public domain and
certainly in the lead's training data; the Gulistan is among the most translated books in the
language. The primary of this study does not depend on the lead's independence: it is a
within-lead, within-passage contrast between two renderings the lead makes himself, one under R44
and one under the tradition's rule, and the seats never see a published hand. The published census
(§5, P5) is a count of what four printed books do and the lead's column is not pooled with them.
Overlap against all four hands is measured with tools/dependence_check.py after the log is frozen
and is reported on the translation page whatever it says. The one direction in which recall could
manufacture a result is ruled out by the prediction itself: P4 predicts that no published hand
carries the distance rhyme, so there is no published deep-run rendering for the lead to recall.
One priming is declared on the passages page: the lead has read حکایت ۱۰'s بنی آدم bayts (V12)
in English many times. V12 is a MATHNAWI passage contributing three adjacent rhymes and no
distance link — it falls entirely on the control class.
4. Procedure
Stage A — the lead's rendering, free, no API
حکایات ۶–۱۳ whole under R44 (workshop/regimes/R44-run-depth.md): prose as prose, verse as
verse, and the English rhyme placed where the Persian rhyme is, at the distance the Persian put it.
Every passage graded HELD / SHORT / REFUSED / UNREACHABLE. The log is frozen and committed
before stage B exists and before any hand is opened.
Stage B — the two counterfactual arms, free, no API
NEAR-PAIR / FAR-PAIR — the matched minimal pair, and the arm the primary is measured on
(amendments A1 and A6). A1's whole-passage permutation was itself confounded — the round-2 critic's
BLOCKING 1: putting every rhyming line together turns an alternating pattern into a contiguous block,
which changes cluster structure and coherence as well as distance. A6 replaces it with a stimulus in
which those cannot vary.
A6's four-line pair was itself confounded — the round-3 critic's BLOCKING 1: in FAR the lines
y₁ x₁ are the two hemistichs of one bayt and stay semantically continuous, and in NEAR neither
bayt survives. That difference biases towards the registered prediction, so it could not be
declared and kept. Amendment A9 removes it by giving the target pair no bayt-mate at all.
Each of the 14 distance links is one item, and each is presented as a six-line stimulus built from the same six lines in both conditions:
| condition | order presented | rhyme partners at |
|---|---|---|
FAR |
f₁ y₁ f₂ y₂ f₃ f₄ |
positions 2 and 4 — distance 2 |
NEAR |
f₁ y₁ y₂ f₂ f₃ f₄ |
positions 2 and 3 — distance 1 |
y₁ and y₂ are the two rhyme-bearing English lines of the link — the lines that render the two
Persian bayt-ends the قطعه holds one rhyme across. f₁–f₄ are filler lines drawn from elsewhere
in this same span's R44 rendering, so they are the same hand, the same register and the same
source, and they belong to neither of the item's two bayts.
The two conditions differ by a single transposition of two adjacent lines — positions 3 and 4
swap. Held identical: the six lines and their words, the line count, y₁'s serial position, the
first and last two positions, the rhyme density (one pair), the cluster size (2), and — the round-3
finding — the number of within-bayt adjacencies, which is zero in both, because no filler is a
hemistich of either of the item's bayts. Neither target is the final line, so the recency asymmetry
A6 carried is gone too.
Filler rule, fixed here and executed by the script, not chosen by hand. The filler pool is every
unrhymed first-hemistich line (x) of the span's QITA passages. For item i, fillers are taken
from the pool in order starting at offset 4i (mod pool size), skipping any line that (a) belongs to
either of the item's own bayts, (b) is graded above NONE against y₁, y₂ or an already-chosen
filler by tools/rhyme_pairs.py, or (c) is already used in this item. So each stimulus contains
exactly one rhyming pair and the tool says so before dispatch.
What A9 costs, declared. A six-line stimulus of one couplet's worth of sense and four unrelated
lines is not verse anybody would print. It is a psychophysical stimulus and the result page calls
it one. The literary reading is carried by S2, where the passages are whole and in order.
COUPLET — the ecological arm, secondary. For each of the 9 QITA passages, a second rendering
of the same bayts under the tradition's rule: each bayt rhymed within itself, couplets, and no
rhyme at any distance-2 position. Same content, same line count, same hand — but different words
and twice the rhyme density, which is exactly why it cannot carry the primary (round-1 critic,
BLOCKING 1).
Order and what it costs, declared. Stage A is written first and frozen first, so the R44
rendering cannot be anchored on the couplet one. The couplet rendering is therefore a re-rendering of
content the hand has already put into English, and may inherit wording — a limit the result page
states. It is the conservative direction: shared wording makes the two arms more alike, not less.
Every couplet passage is graded mechanically before dispatch and any accidental distance-2 rhyme is
re-rendered, because a couplet arm that also carries the run is not the tradition's move.
Stage S — the registration measurement (paid)
Two sub-stages, one instrument, one prompt:
S1, the primary — 28 bodies: 14 links × {NEAR,FAR}, each to three seats — 84 calls.S2, the ecological arms — 30 bodies: the 21R44passages whole + the 9COUPLETpassages, each to three seats — 90 calls.
The seat sees numbered English lines and nothing else. No mention of Persian, of Sa'di, of translation, of rhyme placement, of this study, and no indication that any two bodies are related. The instruction is fixed:
Below are numbered lines of English verse. Look at the last word of each line. Give every line a letter. Two lines get the same letter when their last words rhyme — judge by how the words sound, not by how they are spelled. A line whose last word rhymes with no other line in the passage gets a letter of its own. Answer with one line per input line, in the form
3: b, and nothing else.
(Amendment A2: the round-1 critic's MAJOR 2 — the instruction originally read "chime, fully or
nearly" while the scoring kept STRICT only. A7, round-2 MAJOR 2: the instruction still does not
operationalise STRICT, and cannot without teaching the seat the scoring rule. What answers the
objection is that NEAR and FAR contain the identical words, so whatever private threshold a
seat uses is the same in both conditions and cancels in the paired difference. It does not cancel in
S2, which is one more reason S2 is secondary. The clause judge by how the words sound, not by
how they are spelled is added on note (bqy).)
Scoring. For every unordered pair of lines at distance 1 or 2 as presented, the seat links
the pair if it gave both lines the same letter. Ground truth for the pair is tools/rhyme_pairs.py
on the two line-final expressions, computed before dispatch and not revised after:
| class | definition |
|---|---|
A+ |
distance 1 as presented, tool says STRICT |
D+ |
distance 2 as presented, tool says STRICT |
A− |
distance 1 as presented, tool says NONE |
D− |
distance 2 as presented, tool says NONE |
Pairs the tool grades NEAR-*, IDENTICAL or UNKNOWN are excluded from S2's classes and
reported separately; the NEAR-inclusive recomputation is reported as a sensitivity analysis on
every figure (amendment A3). In S1 the target pair is the same words in both conditions whatever
the tool grades it, so S1's primary keeps every link whose tool grade is STRICT and reports
the NEAR and UNKNOWN links separately. Dispatch order is a single shuffle of all 174 calls from
seed 20260823.
On transitivity (amendment A3, round-1 critic MAJOR 2, second half). A letter is an equivalence
class and a phonetic relation is not, so a NEAR pair inside a passage can pull a STRICT pair's
letters around. Between NEAR and FAR the words are identical, so any such artefact is present
in both conditions in the same amount and cancels in the primary difference. It does not cancel in
S2, which is a further reason S2 is secondary.
The primary statistic, and why it is not a pooled rate (amendment A8, round-2 critic MAJOR 3). Links inside one run are not independent — a seat that gives one letter to a four-member set produces several pairwise hits mechanically — and three model seats are three calls to mutable services, not a sample of readers. So:
- The unit is the (link, condition, seat) cell, and the paired comparison is within each
(link, seat): linked in
NEARbut notFAR, or the reverse. - The inferential unit is the RUN, not the link (amendment A10, round-3 critic MAJOR 2). Links inside one run share a bayt, a rhyme-bearing line and one translated rhyme set, so a link-level exact test has no valid null. Each of the 9 runs gets one preregistered outcome — the sign of its links × seats majority — and the registered test is an exact one-sided sign test over the 9 runs, on the discordant ones. The 14-link and 84-cell tables are printed whole and reported descriptively.
P1is reported only with three seats retained (amendment A11, round-3 critic MAJOR 3). IfP3is dropped for cost, or any seat for unparsed bodies, the majority rule is undefined:P1's pooled statistic is then withheld and the run reports each retained seat's own paired difference separately. The rule is fixed here so the denominator cannot be chosen after the outcomes are seen.- Every claim is stated as what these seats did on these calls on this date, with slugs logged as provenance. No rate is offered as an estimate of anything beyond them.
Seats: P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5. P4 is
out on note (bps), P5 on note (bne), GL on LONG prompts, and QR is excluded by name on note
(bqy) — a seat that reads spelling rather than sound must not sit on a design whose whole question
is where a sound falls. P3 is a cost problem and is priced from its own first twelve calls before
the rest of the stage is dispatched (note (bqk)); if its measured rate breaks the arithmetic it is
dropped and the result reports two seats. There is no fourth reading seat to buy and the result
will say so rather than implying a panel.
Stage F — the addition check on the forced renderings (paid)
A four-member rhyme set is where a translator pads. 18 items — the 9 QITA passages in both arms
— plus 6 planted items, go to two seats (P1, P2) with the Persian blocks themselves as
the standard and the question: does the English state anything the Persian does not? 48 calls.
The six plants are the same passage with one added modifier, intensifier or evaluation of at most
three words — the shape an addition takes when a translator buys a rhyme — constructed and frozen
before dispatch, never a new event or character. If the seats do not catch at least 4 of the 6
plants pooled, stage F clears nothing, the R44 grades are reported as self-certified, and the
result page says so.
What a stage-F clearance means, at its true strength: no addition was detected by two seats
reading the Persian, on an instrument shown to catch four of six three-word plants. It is not a
certification by an established Persian reader; the project has none, and NEXT.md has carried
independent human readers as named, not built for weeks.
Stage P — the published census (free, mechanical, no API)
After stages A and B are frozen, the four hands are extracted for this span exactly as S213 extracted
them (materials/hands.md, to be written then). At each of the three deep runs (V07 depth 3,
V09 depth 4, V11 depth 4) and, if the extraction is clean, at the six depth-2 runs, for each hand:
the English word rendering each Persian rhyme-bearer is identified — bearers, not clause-ends,
which is note (bra)'s requirement and the reason S213's S11 did not become a false four-for-four.
Amendment A5 (round-1 critic, MAJOR 4). A source rhyme position is counted only when the
English word rendering its bearer stands at an English line-end, and the hand's depth at that
run is the largest mutually-STRICT set among those positioned ends. Three codes, not one scale:
- depth n — n of the source's rhyme positions are rendered at English line-ends and rhyme;
MOVED— the bearer is rendered but not at a line-end; the English is quoted, andMOVEDis reported as its own count, because a relocated bearer is neither a held rhyme nor a failure to rhyme;n/a — prose setting— the hand sets the passage as prose, so no English line-end exists. This is the honest code for Gladwin and Ross, whose policy S213 established, and it is not depth 0.
Every extracted cell's raw text is preserved in raw/.
5. Predictions, registered
| id | statement | bar |
|---|---|---|
F1 |
gate — hit rate on the target pair in the NEAR condition, pooled over 14 links × 3 seats (the matched instrument floor) |
≥ 0.80 |
F2 |
gate — link rate on A− and D− items (non-rhyming pairs), pooled over S1 and S2 |
≤ 0.20 |
P1 |
primary, matched, identical words — the paired NEAR − FAR target-pair hit rate over the 84 S1 cells, with the registered test an exact one-sided sign test over the 9 runs (A10) and reporting conditional on three retained seats (A11) |
rate difference ≥ 0.25 |
P2 |
secondary, ecological — hit(A+, COUPLET arm) − hit(D+, R44 arm) in S2, same nine passages, different words and twice the rhyme density |
same sign as P1 |
P3 |
secondary, natural — hit(A+, the 17 MATHNAWI adjacent rhymes) − hit(D+, R44) |
same sign as P1 |
P4 |
descriptive — hit(D+) by position in the run (link 1 vs links 2–3 of the deep runs) |
no bar |
P5 |
the published census — no published hand holds the source's rhyme at English line-ends to the source's depth at any run of depth ≥ 3 | 0 of 12 cells |
P1's direction is registered: the adjacent placement is predicted to be heard more often.
What a miss means, in the critic's own words (amendment A4, round-1 MAJOR 3). With 14 distance links, nine passages and links correlated inside a run, a difference below 0.25 licenses exactly one sentence — "the preregistered ≥ 0.25 contrast was not observed" — and the result page is forbidden the sentence distance is not the obstacle. A real adverse effect smaller than 0.25 is fully compatible with a miss.
6. Failure criteria, registered
F1below 0.80 → the task does not detect an adjacent rhyme on the same words that the distance arm carries; everything is withheld and the session reports an instrument that did not work.F2above 0.20 → the seats are filling in a scheme rather than reading line-ends; the primary is withheld and only the false-alarm figure is published.- Fewer than 8 of the 14 links carrying a tool-
STRICTEnglish rhyme — i.e.R44failed to place the rhyme often enough to measure — →P1is withheld; the reachability grades from stage A are reported alone, as a translator's finding about what could not be done. - Stage F catches fewer than 4 of 6 plants → the
R44grades are reported self-certified andP1is reported with a stated fidelity limit, not withheld (the seats' rhyme reading does not depend on the renderings being faithful). - Any seat returning fewer than 49 of its 58 stage-S bodies parsed → that seat is dropped whole and the result reports the remaining seats.
- Saturation — if the target-pair hit rate exceeds 0.95 in both
NEARandFAR, the task is too easy to discriminate andP1is reported as uninformative rather than as a null.
7. What the result may say, and what it may not
- It may say what three model seats do with a rhyme at two distances on one hand's English.
- It may not say that human readers hear or fail to hear a قطعه.
- It may say what four printed books do at nine places in one author.
- It may not say that English cannot hold a rhyme at distance; stage A measures one hand's reach, under a rule that forces the attempt, and that hand is contaminated.
- Tier D is NOT PASSED; nothing here is a quality claim;
internal-judgment-onlyandprovisionalstand on every page this run produces.
8. Budget
Declared experiment ceiling $1.60. Runner ceiling $1.35, stop-loss $1.20. UTC day 2026-08-23 opens with the full $5.00 unspent.
The ceiling rose from $1.20 to $1.50 and then to $1.60 during design, before any data call, both
times on a critic finding: A1 added an arm, and A6 replaced that arm with a 28-body matched-pair
stage while S2 kept the whole-passage arms. The reasons are in critic-response.md and the
arithmetic is here.
| stage | calls | worst case | basis |
|---|---|---|---|
| pre-run critic, up to 3 rounds | 3 | $0.30 | rounds 1 and 2 billed $0.054955 and $0.079273 |
| stage S1 — matched pairs | 84 | $0.40 | max_tokens 400; stimuli are 4 lines; P3 priced from its own first 12 calls |
| stage S2 — whole passages | 90 | $0.45 | max_tokens 400; passages ≤ 8 lines |
| stage F | 48 | $0.35 | max_tokens 500; prompts carry Persian and English |
| headroom | — | $0.10 | re-dispatches |
| total | 225 | $1.60 |
The worst case is built from max_tokens, not from an assumed output length — note (abc). Costs
are the API-returned billed figures with "usage": {"include": true}; the key-usage delta is not
a cross-check and is not reported as one, note (bof).