Repository path: workshop/experiments/E-20260822c-persian-hands/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260822c-persian-hands |
| status | frozen |
| created | 2026-08-22 |
| updated | 2026-08-22 |
| senses | style-correspondence |
| links | wiki/arms/ARM-persian-hands.md, workshop/translations/gulistan-bab1/loci-frozen.md, workshop/translations/gulistan-bab1/collation.md, workshop/regimes/R43-prose-holds.md, framework/v0.2/README.md, tools/rhyme_pairs.py, config/models.md, config/budget.md |
E-20260822c — where four published hands put the chime, and whether the prose one was reachable
Frozen before any published hand was opened beyond the priming declared on
workshop/translations/gulistan-bab1/loci-frozen.md, and before a word of the lead's rendering
existed.
1. Question
framework/v0.2 §7.26, written 2026-08-22, tells a practitioner that a chime is registered at
+0.542 at a verse line-end and +0.222 at a prose colon-end — that these are two moves, not one.
That is a reading measurement on constructed variants. Nothing in this project says what
translators do about it.
Sa'di's Gulistan puts sound figures in both places, alternating inside a single paragraph. At a prose one a translator has exactly three answers:
- drop it — render the sense, let the figure go;
- chime in prose — find an English pair that echoes at the same colon-ends;
- change genre — lift the passage into English verse, where a rhyme is conventional and needs no defence.
Which did four published English hands take, over ninety-three years; and is answer 2 rare because translators decline it or because English will not supply it?
The second half is what the translation limb is for. The wire between the limbs, in one sentence: the lead renders the same five tales under a regime that forces an attempt at a prose chime at every one of Sa'di's 32 rhymed-prose loci and forbids the escape into verse, so that the study limb's count of what published hands did can be read against a measurement of what was reachable.
2. Materials
Source. «گلستان» باب اول, حکایات ۱–۵ — 68 blocks, 24 prose and 44 bayts, 1,598 normalised
Persian tokens. Copy-text and its single-witness declaration: ../../translations/gulistan-bab1/collation.md.
Loci, frozen from the Persian before any English: ../../translations/gulistan-bab1/loci-frozen.md.
42 VERSE bayts (2 Arabic excluded), 32 PROSE-SAJ loci (15 STEM, 17 AFFIX), 16
CONTROL-PLAIN narrative stretches.
The four hands, all public domain, all full text from the Internet Archive, identifiers recorded
in materials/hands.md with the fetch date:
| code | hand | edition used |
|---|---|---|
GLA |
Francis Gladwin, 1806 | Boston 1865 reprint, gulistan00unkngoog |
ROS |
James Ross, 1823 | gulistanorflowe00rossgoog |
EAS |
Edward B. Eastwick, 1852 | 2nd ed. 1880, gulistanorrosega00sadiuoft |
ARN |
Sir Edwin Arnold, 1899 | frompersianguli00arnogoog |
The lead's rendering. T-gulistan-bab1-R43-v1, under R43 (../../regimes/R43-prose-holds.md),
written after this page is frozen and before any hand beyond the priming is opened. Contamination
high and declared in advance — all four of these hands are public-domain and certainly in the
lead's training data, and the lead has read حکایت ۱ in all four. The primary of this study does not
depend on the lead's independence: the primary is a count of what four printed books do, and the
lead's column is reported separately, labelled, and never pooled with them. This is the standing rule
in CLAUDE.md §Contamination applied rather than waived.
3. Procedure
Stage A — the lead's rendering. حکایات ۱–۵ whole under R43. Log frozen and committed before
stage B opens any hand.
Stage B — the census. For each of the 90 admitted units × 4 hands = 360 cells, two codes:
FORM ∈ {VERSE, PROSE, OMITTED} and CHIME ∈ {STRICT, NEAR, NONE, n/a}. Raw text of
every cell is preserved in raw/. How each is obtained is fixed in §4a below and is not the coder's
to vary.
Stage L — the availability check on the lead's own chimes, paid. (Amended after the round-2
critic's MAJOR 3: four historical translations are not semantic ground truth — they can omit in
common, depend on each other, or misread the Persian together.) Every locus the lead codes TAKEN
under R43 goes to two seats (P1, P2) with the Persian block itself as the standard, the
four published renderings supplied as aids and not as the criterion, and the lead's English
clause. The question is source-faithfulness: does the English clause state anything the Persian
does not?
The seats' Persian is unprobed, so stage L carries its own positive control — and the plants are
built to be the kind of addition this study is looking for. (Round-3 critic, MAJOR 4: catching
eight conspicuous insertions would not show that a subtle one is caught.) Eight items are
planted: the same lead clause with one added modifier, intensifier or evaluation of at most
three words — the shape an addition takes when a translator buys a rhyme — constructed and frozen
before dispatch, never a new event or a new character. If the seats do not catch at least 6 of the
8 plants pooled, stage L clears nothing, P4's TAKEN count is reported as self-certified, and
the result page says so. A TAKEN locus that both seats mark as adding material is reclassified
AVAILABLE-REFUSED and does not count towards P4.
What a stage-L clearance means, stated at its true strength. It means no addition was detected
by two seats reading the Persian with four historical renderings to hand, on an instrument shown to
catch six of eight three-word plants. It is not a certification of source-faithfulness by
established Persian readers; the project has none, and NEXT.md has carried independent human
readers as named, not built for weeks. The result page states the clearance in these words and
not in stronger ones.
Stage C — the copy-text substitute check. Every locus is checked against the four renderings for
a reading implying a different Persian word (collation.md §The substitute check). Descriptive;
findings go to §Limits.
3a. How FORM and CHIME are obtained (amended; critic findings 1, 2, 5)
FORM is read off the page scan, not the OCR. The Internet Archive serves the page images of all
four books (archive.org/download/<id>/page/n<N>_w800.jpg); the study span is 13–15 pages per hand
and every one of them is read. This is a fact about the printed book. The OCR is used only to locate
the cell.
The geometry classifier is the independent check. layout.py reads the per-word bounding boxes
in each book's _djvu.xml and classifies every printed line VERSE-SET / PROSE-SET / LABEL by a
rule frozen in its docstring before the span was coded. The agreement between the image coding and
the geometry is the reported verification of FORM — two independent readings of the same pages,
one by eye and one by coordinates, neither of them the OCR text.
CHIME is graded at the locus's own terminal positions, and the alignment that fixes them is
written down before any rhyme is graded. (Round-2 critic, two BLOCKING findings: an
all-clause-ends maximum credits incidental rhymes, double-counts a cell shared by two loci, and never
looks at a printed verse line-end at all. Both are accepted; the rule below replaces it.)
Per locus and hand, in this order and no other:
- Alignment, recorded first. For each Persian rhyme-bearer of the locus, the coder writes
down the English word or phrase that renders it, quoted verbatim, into
raw/alignment-<hand>.json. This is a judgment about sense, made bearer by bearer, and the file is committed before the grading script is run. A bearer with no English rendering isUNRENDERED; a locus with two bearers rendering to one word isCOLLAPSED; a locus the hand does not render at all isOMITTED. All three are counted and reported, never silently dropped. - The graded words are the aligned renderings themselves. (Round-3 critic, BLOCKING 2: substituting the segment-final word for a bearer rendered mid-clause imports a semantically unrelated word, which is the incidental-rhyme fault the alignment was introduced to remove.) The pair sent to the tool is the last word of each aligned rendering — never a word the alignment did not name. No segment substitution is performed anywhere.
- Multi-bearer loci have one frozen aggregation rule. (Round-3 critic, BLOCKING 1:
S10,S25andS26carry three or four bearers.) For a locus with n bearers in Persian order, grade the n−1 adjacent pairs, andCHIMEis the best verdict among them, with the winning pair recorded. For the 29 two-bearer loci this is exactly one pair. Saj' chains are sequential, so adjacency is the source's own grouping and not a choice made for this study. TERMINALis a separate flag, not a substitution. A bearer isTERMINALif its aligned rendering ends its segment; segments are bounded by, ; : . ? ! —, clause-leveland/or/but/nor, and — mandatory in every cell the hand prints as verse — the printed verse-line boundaries taken from the page scan. A locus isANSWERED-TERMINALwhen it isANSWEREDand both members of its winning pair areTERMINAL. Both rates are reported everywhere.- Grade.
tools/rhyme_pairs.pyon that pair and no other. No pair is credited to more than one locus.
The coder's own reading is logged beside the tool's and where they differ both stand. The
identical procedure runs on the CONTROL-PLAIN units, whose "bearers" are the two parallel elements
named in loci_control.json.
4. Registered predictions
All rates are over admitted units. حکایت ۱ is PRIMED; every figure below is computed twice, with
and without it, and where the two disagree the smaller effect governs.
-
P1— the prose chime is not in the record. Pooled over the four hands and the 32PROSE-SAJloci (128 cells), the rate ofANSWEREDwithFORM = PROSEis below 0.15. Fails if ≥ 0.15. -
P2— the answer that is taken is the move into verse. DESCRIPTIVE, not causal (critic finding 3). Over the same 128 cells, theSET-AS-VERSErate exceeds the prose-chime rate ofP1, and at least one hand sets ≥ 0.30 of its 32 saj' loci as verse. Fails if either clause fails. No claim is made that the verse-setting is a response to Sa'di's sound figure: the saj' loci are aphorisms and the controls are narration, so discourse function is an uneliminated rival explanation and is named as such wherever this number is reported. -
P3— the source's form, not the English position. (Round-3 critic, MAJOR 3: the verse side mixes cells the hand printed as verse with cells it printed as prose, so the difference cannot be read as an effect of English terminal position.) Among hands that print any English verse at all, theANSWEREDrate at the 42 source-verse loci exceeds theANSWEREDrate at the 32 source-prose-saj loci by ≥ 0.30, for every such hand. Fails if any such hand falls short. This is a comparison of what Sa'di wrote in verse against what he wrote in rhymed prose, and it is labelled that way wherever it appears. The English-side split — source-verse loci the hand set as verse against those it set as prose — is reported as a stratification and is not the prediction. -
F1— the control, reported as a difference and no longer licensing a causal claim. For each hand,SET-AS-VERSEatPROSE-SAJminusSET-AS-VERSEat the 16CONTROL-PLAINunits, with the controls split intoCONTROL-SPEECH(7) andCONTROL-NARRATION(9) and both reported. A hand whose difference is at or below zero is setting verse by its own policy; the study says so about that hand. A positive difference is not evidence that the sound figure caused it — seeP2. -
P4— was the prose answer reachable, by this hand, under this rule? UnderR43, which requires the attempt and forbids the escape into verse, the lead reachesTAKENat ≥ 0.40 of the 32 saj' loci, counting only loci that survive stage L. This is not a blind test of the lead: the lead writes the rendering knowing this prediction, and the rendering is contaminatedhigh. And the inference is one-directional and narrow (critic finding 4). If the hands sit near zero and the lead reaches 0.40 on stage-L-cleared loci, then a prose chime was available to this hand under this rule, and the hands' near-zero is not forced by the absence of candidates. If the lead also falls short, that shows only that this hand, enumerating this way, found no pair — it does not show that English is the constraint, and the result page may not say that it does. -
S1— secondary,STEMvsAFFIX. The same rates split by locus class, on the reasoning imported from the دیباچه loci page. No threshold; reported.
Failure criteria are these five statements and nothing else. No rate is redefined after the coding; no locus is added, dropped or reclassified after this page is committed.
5. What this design cannot establish
- It is four hands, one work, one language pair. Nothing here generalises to rhymed prose in Arabic, or to any author but Sa'di, and the anchor will say so on its face.
- The four hands are not independent. Eastwick and Arnold had their predecessors on the desk; Arnold's title page names his. Agreement among them is weak evidence and is reported as such.
ANSWEREDisrhyme_pairs.py's verdict, not a reader's.RS-20260822bmeasured what a reader registers; this study measures what a dictionary grades. The two are different objects and the result page does not slide between them.FORMis what the book prints, which is partly the publisher's doing and not only the translator's. The study says what four books set as verse; it does not say what four men intended. Where a hand's policy is stated on its own title page (EAS: "translated for the first time into prose and verse") that is quoted rather than inferred.- The copy-text is single-witness (
collation.md), so a locus could in principle rhyme in Foroughi and not in the recension a hand used. Stage C is the substitute and its limits are stated there.
6. Cost
Declared experiment ceiling $1.30, raised from $0.75 in two steps during design and before any data
call — to $1.10 because round 1's finding 4 turned P4's self-certified success into a bought
check (stage L) while its findings 2 and 5 deleted stage V, and to $1.30 because round 2's finding 3
put the Persian source into every stage-L prompt and added eight planted controls. Today's UTC headroom before this session: $2.736569250.
| stage | seats | calls | worst case |
|---|---|---|---|
| pre-run critic, round 1 — spent, $0.0551315 | P1 |
1 | $0.12 |
| pre-run critic, round 2, on the amended design | P1 |
1 | $0.15 |
| pre-run critic, round 3, on the twice-amended design | P1 |
1 | $0.18 |
| stage L — up to 32 loci + 8 planted, × 2 seats, long prompts (Persian + four aids) | P1, P2 |
80 + re-dispatches | $0.70 |
| ~~stage V~~ — deleted, critic findings 2 and 5; replaced by the free image/geometry check | — | 0 | — |
| headroom | $0.15 | ||
| ceiling | $1.30 |
Worst case is built from max_tokens and the measured per-call rates of 2026-08-22 (P1
$0.001488, P2 $0.000458), not from config/models.md's list prices — note (abc), and the
S212 precedent. P3, P4, P5 and GL are not used; the reasons are in NEXT.md's blocked block.
Runner stop-loss $1.05. If billed cost passes it before stage L completes, the run stops and the partial sample is reported with its size.
- The saj' loci are aphorisms and the controls are narration (critic finding 3). Discourse
function is an uneliminated rival explanation for every
SET-AS-VERSEdifference, and function-matched controls are not buildable in this author, note (bqr).
4a. Amendments made after the pre-run critic, before any data call
Three rounds, all seat P1, $0.209577500 in total, all three returning NEEDS-REDESIGN: 5 findings
then 3 then 5, 4 of them BLOCKING. Every finding was accepted; one remedy was declined in part with
the reason written. Full record: critic-response.md.
What changed. Round 1: FORM moves from the OCR to the page scans and gains a free mechanical
check (layout.py); P2 loses its causal reading and is renamed SET-AS-VERSE; F1 stops
licensing a carriage claim and the controls split speech from narration; P4 is restricted to this
hand under this rule and its success case is bought (stage L); stage V is deleted.
Round 2: the rhyme pair is fixed by a per-bearer alignment committed before grading, not by a
cell-wide maximum; printed verse-line boundaries become mandatory segment boundaries; stage L's
criterion becomes the Persian itself and gains eight planted controls.
Round 3: the graded words become the aligned renderings themselves rather than segment-final
substitutes; multi-bearer loci get one frozen adjacent-pair aggregation rule; P3 is relabelled a
comparison of source forms; the plants are built subtle and the clearance is stated at its true
strength; §7's dangling stage-V reference is replaced.
The ceiling rose $0.75 → $1.10 → $1.30, in both cases before any data call.
Why the loop stops at three rounds, written rather than left to be noticed. The round-3 findings are refinements of a coding rule, not new ways for the primary to be false, and the primary is a count of what four printed books do — an object no further critic round makes more or less true. Three rounds have cost $0.21 and bought four BLOCKING repairs; a fourth would trade session depth for diminishing precision on a definition that is now written down in five numbered steps. The decision is the lead's and is recorded here so that a reader can disagree with it.
7. Verification
analysis/verify.py, importing nothing from analysis/analyse.py, recomputes every number this
design reports from raw/ — the rates, the with/without-حکایت ۱ pair, the F1 differences, the
ANSWERED / ANSWERED-TERMINAL pair, the stage-L plant-detection count — and re-derives the pooled
counts by direct enumeration. The FORM verification it recomputes is the image-versus-geometry
comparison (round-3 critic, MINOR 5: §7 still named the deleted stage V): layout.py's
classification against the lead's page-scan coding, cell by cell, with the agreement rate printed and
every disagreement listed by cell id. Disagreements are resolved in favour of the page scan,
which is the printed book, and the count of them is reported in §Limits. It must exit 0 with zero
failures before the result page is written.