Repository path: workshop/experiments/E-20260824c-run-placement/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260824c-run-placement |
| status | frozen |
| created | 2026-08-24 |
| updated | 2026-08-24 |
| senses | style-correspondence |
| provisional | true |
| internal-judgment-only | true |
| links | wiki/arms/ARM-run-depth.md, workshop/translations/gulistan-bab1e/R47-v1/translation.md, workshop/regimes/R47-three-placements.md, workshop/experiments/E-20260824c-run-placement/critic-response.md, wiki/findings/results/RS-20260823-run-depth.md, wiki/findings/results/RS-20260822b-echo-threshold.md, wiki/findings/results/RS-20260824-eye-or-ear.md, wiki/method-notes.md, config/models.md, tools/rhyme_pairs.py |
E-20260824c — the same chime at two distances: ARM-run-depth step 2, rebuilt as an indirect measure
Design v2, frozen before dispatch. v1 was frozen, sent to one adversarial critic round, and
rebuilt on its findings; v1 is in git history and every finding is answered in
critic-response.md. The translator's log this is built on was frozen and committed at
ef22a3df, before v1 existed.
1. The question, and why this is not E-20260823 again
A Persian قطعه holds one rhyme across the ends of consecutive bayts and leaves the hemistichs
between them bare: in English, one line to a hemistich, a chime at every second line-end. The
published tradition replaces it with couplets. ARM-run-depth step 1 (RS-20260823-run-depth)
established that the source's placement can be carried, that carrying it costs repetition rather
than invention, and that no published hand has ever held one (0 of 12 cells). It could not
establish whether the placement is registered by a reader: the task it bought asked three seats to
name the rhyme scheme, and they named it at 1.000 in both conditions. Method note (brb):
a direct identification task on model seats is a ceiling, not a measurement, and the remedy it
prescribes is RS-20260822b's shape — never ask the seat to identify a chime; ask which of two
variants reads better, and infer registration from the difference.
This design is that rebuild. The seat is never asked about rhyme, sound, endings, or form. It is asked which of two passages reads better. The passages differ in one word.
2. The design
For each locus, one four-line English passage in five placements, and within each placement two variants differing only in the last word of line 4:
| placement | line 2's end | line 3's end | chime |
|---|---|---|---|
B0 |
as written | as written | none |
D2 |
rewritten to chime with w+ |
as written | at line-ends 2 and 4 — the Persian's placement |
D2c |
rewritten, not chiming | as written | none — the control for D2 |
D1 |
as written | rewritten to chime with w+ |
at line-ends 3 and 4 — adjacent |
D1c |
as written | rewritten, not chiming | none — the control for D1 |
w+completes the placement's chime;w−does not. The same word pair is used in all five placements.D1andD2are matched on the number of chimes (one), on the rhyme family, on the target word, and on the fact that one carrier line has been rewritten. The only difference between them is the distance between the chime's partners.D1is not the tradition's couplet form, which rhymes both bayts and so supplies two chimes. It is the right control for distance, which is this arm's question, and the design claims nothing about the tradition's form.
Outcome per cell: which variant the seat says reads better. hit = it chose +.
Registered quantities, each the mean over loci of the per-locus rate (equal weight per locus),
pooled over both orders and the two reporting seats P1 and P2:
- Δ_D1 = hit(
D1) − hit(D1c) — the power gate: is a chime at adjacent line-ends registered at all, over and above the effect of having rewritten that line? - Δ_D2 = hit(
D2) − hit(D2c) — the primary: the same question at the Persian's distance. - Δ_D1 − Δ_D2 — the extent statement, and the number
ARM-run-depthwas constituted to get. - Printed alongside, not primary: the carrier effects hit(
D1c) − hit(B0) and hit(D2c) − hit(B0), which say how much of aB0-referenced difference is not the chime at all.
3. Materials
Seven loci, every one of them four lines of the lead's own English rendering «گلستان» باب اول
۲۶–۳۱, written under R47 at one sitting and frozen at ef22a3df before this design existed. The
Persian rhyme structure was frozen earlier still, from the Persian alone
(workshop/translations/gulistan-bab1e/passages-frozen.md).
| locus | Persian passage | source depth | w+ / w− |
D2 partner |
D1 partner |
|---|---|---|---|---|---|
L1 |
V2, p003–p004 |
2 | go / rise | so | below |
L2 |
V3, p006–p007 |
2 | disarray / confusion | day | lay |
L3 |
V4, p009–p010 |
2 | go / pass | below | so |
L4 |
V6, p017–p018 |
2 | tend / serve | depend | end |
L5 |
V7 bayts 1–2 |
4 | head / skull | bled | fed |
L6 |
V7 bayts 3–4 |
4 | told / said | hold | mould |
L7 |
V10, p029–p030 |
2 (+مطلع) | passed / went | last | cast |
Two of the nine four-line units the span supplies were declared NO-ALTERNATIVE by the regime
before this design existed (V5, V9) — the hand could not write two equally faithful final words
in a family the neighbouring lines could reach. They supply nothing here and are not replaced.
F5, the mechanical check, is discharged before the run and is a precondition of it. All 70
variants were graded over all six line-end pairs with tools/rhyme_pairs.py: every + variant of
D1 and D2 carries exactly its intended STRICT pair and no other relation of any kind; every
+ variant of B0, D1c and D2c carries none; and no − variant carries any relation at all.
Recomputed by verify.py.
One instrument limit, found by the critic and not by the tool. v1's L6 chimed the written
decree is read with said. tools/rhyme_pairs.py graded it STRICT because CMUdict carries
both /rɛd/ and /riːd/ and the rule relates two bearers if any pair of pronunciations relates — but
a reader reading a passive present verb does not hear /rɛd/. L6 was rebuilt on hold · mould ·
told · said, which is unambiguous. A mechanical grader that accepts any dictionary pronunciation
can certify a chime a reader will not get**, and that goes on the result page.
Contamination. Measured after the log froze and before v1 was written:
tools/dependence_check.py, the lead's whole record English of the span against Eastwick 1852 —
5 shared 7-grams, 0 twelve-grams, longest common run 8 tokens, verdict clean. The declaration
on the artifact stays suspected (a measurement locates a run; it does not license an independence
claim). No quantity in this design turns on the lead's independence from any published hand: the
carriers are the lead's own English and the manipulation is applied to it.
4. Procedure
Stage P — the primary. Prompt, fixed, naming neither rhyme nor sound nor line-endings:
Below are two versions of the same short passage of verse. They differ in one word. Which version reads better? Answer with a single letter, A or B, and nothing else.
Both orders are run for every cell (+ first and − first). Cells: 7 loci × 5 placements ×
2 orders × 3 seats = 210 — of which P1 70, P2 70, QR 70.
Stage A — the faithfulness check on the word pair, descriptive, prunes nothing. For each
locus: the Persian bayts, the B0 English quatrain with the final word blanked, and three
candidates — w+, w−, and a planted word that introduces something the Persian has not got
(scream, flames, sell, shear, coffin, bought, rained) — in an order randomised per (locus, seat)
with seed 20260824, the map recorded per cell. The seat returns, per candidate, one of ADDS /
LOSES / FINE. Cells: 7 loci × 2 seats (P1, P2) = 14.
Seats (config/models.md; slugs are provenance, roles are configuration):
| role | slug | use |
|---|---|---|
P1 |
openai/gpt-5.6-terra |
reporting |
P2 |
google/gemini-3.6-flash |
reporting |
QR |
qwen/qwen3.7-max |
declared item-set sensitivity check, 70 stage-P calls, reported under its own heading, never pooled with P1/P2, no bar attached |
QR's role is exactly what RS-20260824 §4 used it for: if the reporting seats show nothing and
QR shows something, the null is a fact about those seats and not about the materials. It is
outcome-conditioned selection, declared before dispatch.
Temperature 0, one call per cell, max_tokens 1500 on stage P and 2500 on stage A
(note (bnk): the cap covers hidden reasoning as well as the answer). All calls in one
foreground process with a thread pool capped at 6 concurrent (note (brf); S210), so at most six
calls are ever in flight; the stop-loss is checked before every dispatch and again between placement
batches. Raw responses preserved per cell in run.jsonl, which is only ever appended to.
Critic. One round, P1 and P2, on v1 — done, $0.118432, fifteen findings, all accepted,
critic-response.md. No second round (note (bqp)).
5. Predictions, registered
R1Δ_D1 ≥ +0.30. A chime at adjacent line-ends is registered.R2Δ_D2 ≥ +0.15, and Δ_D2 < Δ_D1. The Persian's placement is registered, but less.R3hit(B0), hit(D1c) and hit(D2c) each lie in [0.20, 0.80] — no reference cell is at floor or ceiling.
The lead's ground for R2 rather than a flat null: §7.26 measured a chime worth +0.542 at a
verse line-end, and RS-20260823 showed the seats can identify a distance-2 rhyme perfectly. What
is unknown is whether identification and preference come apart.
6. Failure criteria, registered
F1— the power gate. If Δ_D1 < +0.25, the task has no demonstrated power to register a chime at all, and Δ_D2 is reported as uninformative, not as a null.ARM-run-depththen closesretiredand the result page says the project could not measure this with the instruments it has.F2— floor/ceiling. If hit(D1c) or hit(D2c) is outside [0.20, 0.80], the Δ that subtracts it is reported as uninformative; the same bound on hit(B0) voids the carrier effects only.F3— order. If |hit(order A) − hit(order B)| ≥ 0.30 pooled, the primary is reported as order-confounded and per-order figures are printed instead of the pooled one.F4— faithfulness, descriptive only. Nothing is dropped on stageA. Every locus where eitherw+orw−is returnedADDSorLOSESby both reporting seats is flagged, and a sensitivity analysis excluding all flagged loci is printed beside the primary. The planted word's catch rate is printed as the check on stageA's own power; if it is caught at fewer than 5 of 7 loci by both seats, stageAis reported as carrying no weight.F5— mechanical. Discharged before dispatch (§3) and recomputed byverify.py. A single unintended relation in any variant voids that locus.F6— instability contingency. If Δ_D1 lands within ±0.05 of +0.25, or Δ_D2 within ±0.05 of +0.15, a full repeat of the affected placement pair is bought before any verdict is written (note (bre): temperature 0 is not determinism).F7— per-locus reference balance. Any locus whose twelve reference observations (B0,D1c,D2c× 2 orders × 2 seats) are unanimous is flagged, and a sensitivity analysis excluding it is printed.
7. Analysis, fixed before the run
Per-locus rate → mean over loci → Δ by subtraction. Intervals are percentile bootstrap over loci,
10,000 resamples, seed 20260824, and are dispersion statistics, not inferences (A13): seven
hand-built loci are not a sample of English quatrains, and every criterion above is a decision rule
about this item set. Per-seat and per-locus figures printed. QR printed separately and never
pooled with the reporting seats.
Verification. verify.py recomputes every reported number by a path that imports nothing from
analyse.py, re-grades all 70 variants from the raw item file, and runs mutation tests that must
each be caught.
8. Budget
Declared ceiling $1.20, of which $0.118432 is already spent on the critic round. Pre-flight
for what remains: stage P 210 calls ≈ $0.53; stage A 14 calls ≈ $0.08; the F6 repeat allowance
≈ $0.30. Expected total ≈ $0.75, worst planned ≈ $1.03.
Worst case is built from max_tokens, not from an assumed output length (note (abc)), and now
includes input tokens (round-1 finding P1-11): 210 × (400 in + 1500 out) + 14 × (700 in + 2500
out), priced at the dearest seat in each stage, is $1.98 — above the ceiling, so the ceiling is
the binding constraint: the dispatcher stops and reports at $1.20 − $0.118432 of run spend, and
with concurrency capped at 6 the overshoot cannot exceed six calls. Today's UTC headroom at the time
of writing: $1.777 after the critic round.
9. What this cannot show
- The seats are models. There are no human readers;
NEXT.mdhas carried independent human readers as named-not-built for weeks. Tier D is NOT PASSED and every figure isprovisionalandinternal-judgment-only. - The construct is preference under explicit comparison, not registration in ordinary reading
(round-1 finding
P1-10). A side-by-side forced choice between two passages differing in one word makes that word salient and may invite deliberate inspection of the line-ends. Salience is identical across all five placements, so it cannot manufacture a difference between them — andD1andD2are matched on it exactly, which is why the extent statement is the quantity least exposed. But no sentence anywhere may say a reader hears anything. D1is not the tradition's couplet. Nothing here compares the Persian's placement with what Gladwin, Ross, Eastwick or Arnold actually do —RS-20260823§6 is where that lives.- Seven loci, one hand, one book, one language pair, and two of the span's nine units were dropped by the regime before the design existed.
- The hand wrote all five placements and knew the question. The controls remove the carrier
rewrite as an explanation; they do not remove the possibility that the hand's
D2rewrites are systematically limper than itsD1rewrites. That is now the leading alternative explanation for any Δ_D1 − Δ_D2 gap, and it is smaller than what it replaced. - One call per cell at temperature 0, which is not determinism — note (bre);
F6is the registered contingency, not a general remedy. w+andw−are not perfectly matched words and cannot be: they are what a translator would actually write.F4andF7measure the exposure rather than removing it.