Repository path: workshop/experiments/E-20260822-synonym-reach/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260822-synonym-reach |
| status | frozen |
| created | 2026-08-22 |
| updated | 2026-08-22 |
| links | wiki/arms/ARM-synonym-reach.md, workshop/translations/gulistan/loci-frozen.md, workshop/translations/gulistan/R41-v1/translation.md, workshop/translations/gulistan/collation.md, workshop/regimes/R41-synonym-enumeration.md, tools/rhyme_pairs.py, framework/v0.2/README.md, config/models.md, wiki/goodness-senses.md |
| senses | style-correspondence |
E-20260822 — is a rhyme actually available among the ordinary synonyms, and is it available only where the source rhymed?
Track T5, ARM-synonym-reach step 1. Study limb of the Gulistan دیباچه, whose prose was
rendered whole under R41 and whose translator's log was frozen at 04f457ad before this design
was written. The locus inventory was frozen one commit earlier, before any English for the span
existed.
1. The claim under test, quoted
framework/v0.2 §7.24 item 2, published 2026-08-21 with no test behind it:
at a locus where the source rhymes, do not ask whether the first English word chimes; enumerate the ordinary synonyms for each member and look for a pair that does. The available rhymes live there.
It came out of one measurement on one Arabic proverb — العواقب is consequences to two blind
seats, and both Burton and this project's own hand nevertheless wrote ends, to chime with
friends. That is one locus, and it licensed a general instruction. This design asks the two
questions the instruction needs answered before a practitioner should spend time on it:
(a) What is the yield? At a rhymed locus, how often does the enumeration actually put a chiming pair within reach — compared with the first, most ordinary word, which is what §7.24 tells the translator not to rely on?
(b) Is the yield about the source at all? English has a large lexicon. If enumerating six
synonyms per member turns up a chime as readily at places the source left plain, then what §7.24
describes is a property of English and following it would be inventing ornament, which
framework/v0.2 §7.14 and §7.20 already warn against in the strongest terms the handbook uses.
2. Why a new source language
§7.24's limits say Arabic→English only, one night of one work. Sa'di's دیباچه is the densest saj' in Persian and the passage the form's reputation rests on. Persian is the project's sixth source language for a long work and its first here, so a result on it is a second language family for the claim, not a replication of the same one.
3. Materials
workshop/translations/gulistan/loci-frozen.md, in full. 36 loci, 76 members, drawn by a salted
hash from a pool of 76 before any English existed:
| class | n | what it is |
|---|---|---|
| STEM | 12 | rhyme not generable by a shared productive ending — جهل/سهل, انیس/جلیس |
| AFFIX | 12 | rhyme carried by a class-general ending — نشینم/چینم, نُزهت/فُسحت |
| CONTROL | 12 | parallel cola of the same text that do not rhyme — موجود/واجب, کجاوه/حجره |
RHYMED = STEM ∪ AFFIX, n = 24.
4. The instrument that decides "chime"
Not a model. tools/rhyme_pairs.py, built this session as a declared gate, decides every chime
mechanically from CMUdict under a rule fixed before any list was drawn, and reports STRICT and
NEAR separately and never summed. Its fixtures (tools/tests/test_rhyme_pairs.py, 49 checks,
0 failures) were written before it scored anything and include §7.24's own paradigm case both ways
round: ends / friends must come back STRICT and consequences / ends must come back NONE.
This removes the defect RS-20260821c §7 declares in its own limits — the instrument counts a
near-echo — and it removes it in the only direction that matters here, since the quantity being
measured is a count of chiming pairs in a list.
5. Procedure
Seats are config/models.md roles; slugs are logged as provenance. P4, P5 and GL are out
(notes (bne), (bps)); P1 P2 P3 QR are what is left and there is no fourth reading seat to
buy, which is stated in the limits rather than implied away.
Stage G — the source-side gate. 36 loci × 3 seats (P1 P2 P3), one call each. The seat is
shown the Persian cola of one locus with the member words marked, and asked whether the marked words
rhyme in Persian: RHYME / NO / UNSURE. No hypothesis, no counts, no English, no mention of
translation, and the RHYMED and CONTROL loci are shuffled together into one indistinguishable
stream.
Stage A — the first ordinary word. 76 members × 2 seats (P2 P3), one call each. The seat is
shown one member at a time, in its own colon, with the sibling members and their cola removed,
and asked for the single most ordinary English rendering. This is RS-20260821c's stage-A
protocol unchanged, so the two runs' first-word figures are comparable. The masking is the whole
validity of the stage: a seat that can see both members can chime on purpose.
Stage B — the enumeration. The same 76 members × the same 2 seats, same masking, a separate call: up to six ordinary English renderings, ranked, each of which the seat would accept as an unmarked rendering of that word in that colon. This is §7.24 item 2's procedure, executed by a hand that cannot see what it would need to rhyme with.
Scoring, mechanical. For each locus:
reach_first— does any cross-member pair of the stage-A first words chime?reach_syn— does any cross-member pair of the union of stage A and stage B chime?
Both computed by rhyme_pairs.py. The scored quantity is seat-specific (§6a.2): a locus counts
as reached for a seat only if that seat's own lists contain the chiming pair. Pooled figures are
reported and none is scored. STRICT and STRICT+NEAR are reported separately and never summed.
Stage C — the usability screen. For each (seat, locus) that is reach_syn and not
reach_first, exactly one chiming pair is screened — the pair whose two words have the lowest
summed rank in that seat's own lists — and both of its members are screened (§6a.4). A blind
seat is shown the Persian colon and the candidate word and asked whether the candidate is an
acceptable ordinary rendering of that word here: SAME / DIFFERENT. 2 seats. A pair counts as
usable only if both members come back SAME from both seats. Four planted-error items
are mixed in, each substituting a word that plainly changes the sense; a screen that does not catch
them is not a screen.
Cross-call repeat control. 8 stage-B members are dispatched a second time to the same seat.
Measures how much of reach_syn is a stable property of the lists and how much is sampling.
6. Registered predictions, bars and failure criteria
Written before dispatch. Bars are not moved after the fact.
Amended before dispatch on the pre-run critic's three BLOCKING and two MAJOR findings; the
amendments are §6a and the response is critic-response.md.
| statement | bar | |
|---|---|---|
| G | the gate: the seats can tell the rhymed loci from the controls in the Persian | majority-RHYME at ≥ 20 of 24 RHYMED and majority-NO at ≥ 10 of 12 CONTROL |
| P1 | the yield: enumeration reaches a chime more often than the first word does | reach_syn(**STEM**) − reach_first(**STEM**) ≥ **+0.25**, STRICT, seat-specific |
| P2′ | the source's own pairings beat the same lists re-paired at random — and nothing wider, per round 2 | reach_syn(**STEM**) − reach_shuffled(**STEM**) ≥ **+0.25**, STRICT, seat-specific |
| P2 | the pair-level-matched half of the same question: parallel cola of this text that do not rhyme | reach_syn(STEM) − reach_syn(CONTROL) ≥ **+0.25**, STRICT, seat-specific |
| P2-joint | the source-side claim itself, and it needs both | made only if P2 and P2′ both clear +0.25; if they disagree, the disagreement is the result and no source-side claim is made |
| P3 | usability: the found chimes are words a translator could actually use | ≥ 0.50 of screened pairs called SAME on both members by both seats |
| P4 | the class split, secondary | reach_syn(STEM) > reach_syn(AFFIX) |
6a. The five amendments, and what each one gives up
BLOCKING 1accepted in full. The primary condition is STEM, not STEM ∪ AFFIX. The manifest itself says the instruction can only be about STEM loci, and scoring the union would let a pass or a fail be driven by loci that do not instantiate the mechanism. The cost is power: the primary now rests on n = 12 loci × 2 seats. That is stated in the limits and not hidden. RHYMED figures are still reported, as secondary.BLOCKING 2accepted in full. The aggregation is declared here, before dispatch:reachis SEAT-SPECIFIC. A locus counts as reached for a seat only if that seat's own lists contain a chiming cross-member pair; the rate is over 24 seat-loci per class. Pooling across seats would let a pair be assembled from one hand's word for one member and another hand's word for the other, which is available to no translator. Every pooled figure in the result is descriptive and none is scored.BLOCKING 3accepted, and remedied differently and better. The critic is right that the CONTROL loci differ from the rhymed loci in more than rhyme — part of speech, frequency, polysemy and synonym-set size all move together with saj' selection, and all of them drive English synonym reach. Its own remedy, controls matched member-by-member on five variables, is not buildable here. What replaces it holds every one of those variables fixed exactly:reach_shuffledre-pairs the members across loci within a class — member 1 of locus X against member 2 of locus Y — and recomputes reach on the same seats' same lists. The only thing destroyed is the pairing the source made. 2,000 re-pairings, seeded, no API cost.P2′is now the primary source-side test andP2is secondary. In addition, mean synonym-set size and part-of-speech mix are reported per class, so the confound the critic names is visible rather than argued away.MAJOR 4accepted in full. Stage C screens BOTH members of every scored pair, andP3counts a pair as usable only if both members pass with both seats. To remove the selection the critic identified, the pair screened is fixed by rule: for each (seat, locus) newly reached, the single chiming pair whose two words have the lowest summed rank in the seat's own lists. No cap, no discretion.MAJOR 5accepted in full. The estimand is named: General American rhyme availability as CMUdict assigns it. Rhoticity and vowel mergers make some pairs dialect-dependent, and the synonym prompts do not name a variety. The result claims availability under this dictionary's convention and nothing wider.
A second critic round WAS bought, on the amendments, precisely because reach_shuffled was a
component the critic had not seen — and it earned its $0.03: round 2's BLOCKING 1 established that
reach_shuffled cannot carry the "source rhyme matters" reading at all, only the narrower one now
written into P2′. Round 2 returned 1 BLOCKING and 2 MAJOR; all three are disposed of in
critic-response.md, one of them accepted-in-part with the refusal written out.
6b. reach_shuffled, specified to the level round 2's MAJOR 2 requires
- Unit of randomisation: one class, one seat. Loci of the class are held in their manifest order.
- Permutation: a derangement — a bijection on the loci with no fixed point, so no locus is ever paired with itself and no member list is used twice in a draw. Locus i keeps its member-1 list and takes locus σ(i)'s member-2 list.
- Loci with more than two members contribute members 1 and 2 only, in the shuffled arm and in
the true-pair arm it is compared against, so the two arms are computed over identical material.
Among the drawn STEM loci — the primary — every locus has exactly two members, so this rule
changes nothing there; it bites only on
A12,C08andC14. - Draws: 2,000. Seed: 20260822.
reach_shuffledis the mean per-draw reach rate over the 2,000 draws, computed per class per seat and then averaged over the two seats, exactly as the true rate is.- The empirical p-value is also reported — the share of draws whose reach rate is at least the true rate — because it is the more informative number and because it does not depend on the bar.
F1 — the withdrawal criterion. If reach_syn(RHYMED) − reach_first(RHYMED) < 0.10, §7.24 item 2
is withdrawn as an instruction and the section says so.
F2 — the manipulation check. If the stage-B lists add fewer than 2 new words per member on average over stage A, the two stages are not different treatments and the P1 comparison is void and withheld.
F3 — the instrument check. If more than 15% of stage-A ∪ stage-B expressions have a rhyme-bearer absent from the dictionary after the two declared repairs, the mechanical rule is not deciding the question; every figure is then descriptive and no prediction is scored.
F4 — the recall screen, and it replaces a screen that could not have fired. The obvious leak
screen — does the answer quote the masked sibling — cannot fire here: the prompt is Persian and the
answers are English, so there is nothing to match. The real exposure is different and larger. The
دیباچه is one of the most famous pages in Persian, and a seat that recognises it could supply a
remembered published translator's chime rather than an available synonym, which would inflate
reach_syn(RHYMED) with no lexical availability behind it. Two things are therefore done, both
declared here:
- A recognition probe. 8 member-cola × 2 seats, asked plainly what work the phrase is from. The rate is reported whatever it is and does not gate anything by itself — the point is that a reader knows it.
- A mechanical recall check, free. Every expression in every stage-A and stage-B list is tested
for presence in the stored Eastwick 1852 text of this same دیباچه. If the chiming pairs that
make loci reachable are drawn from Eastwick's vocabulary at a higher rate than the non-chiming
expressions are, the result is reported as confounded by recall and no prediction is scored.
Round 2's MAJOR 3 struck the bar this line originally carried. Exact vocabulary overlap with
one hand is not a reliable marker of remembered translation — recall can come from another hand,
in another morphological form, or in words common enough to match by chance — and a +0.20
allowance would have let a real association stand while the run was called clean. The check is
descriptive and clears nothing. Matching is normalised (lowercased, and stemmed through
rhyme_pairs's own suffix stripper) and the rate is reported beside the base rate. Recall therefore stands as an unexcluded alternative explanation and is named in the limits. Making recognition an exclusion criterion was proposed and is declined in writing (critic-response.md, round 2 §3): no seat this project can buy will fail to recognise this page, so that criterion excludes every possible run rather than a confounded one.
If G fails, every primary is withheld and the run reports what it saw with no claim attached.
7. What this design cannot do
- It measures availability, not use. That a chiming pair exists among the ordinary synonyms does
not say a translator would find it, want it, or be right to take it. The lead's own
R41log is the only use-side record here, it is one hand, and that hand holds the hypothesis. - No published hand is scored per locus. Eastwick 1852 is the only free English of this دیباچه the project could reach whole, and he converts long stretches of the saj' into metrical rhymed English verse. "Did he echo at this locus" is not well posed against a hand that has changed the form; scoring it would have manufactured a number. Declared NOT RUN, with the reason, rather than run badly.
- Model seats, not human readers, at every stage; the project has none and says so.
- Tier D NOT PASSED. No jury, no quality claim about any rendering,
internal-judgment-only.
8. Budget
Today's ledger (UTC 2026-08-22) is $5.00 unspent. Worst case built from the caps the requests permit, per note (abc), and assuming every call needs the doubled-cap re-dispatch, per notes (bqs) and (bqk):
| stage | seats | calls | first pass | every call re-dispatched |
|---|---|---|---|---|
| critic | P1 |
1 | 0.040 | 0.080 |
| G | P1 P2 P3 | 108 | 0.113 | 0.170 |
| A | P2 P3 | 152 | 0.122 | 0.184 |
| B | P2 P3 | 152 | 0.209 | 0.314 |
| C | P2 P3 | ~72 | 0.058 | 0.087 |
| repeat + recognition | P2 P3 | 24 | 0.024 | 0.036 |
| total | ~509 | 0.566 | 0.871 |
Experiment ceiling $1.60. Runner ceiling $1.40, stop-loss $1.20. Both guards sit above the absolute worst case, which is the point: if one fires, the arithmetic was wrong.