Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260822-synonym-reach/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260822-synonym-reach
statusfrozen
created2026-08-22
updated2026-08-22
linkswiki/arms/ARM-synonym-reach.md, workshop/translations/gulistan/loci-frozen.md, workshop/translations/gulistan/R41-v1/translation.md, workshop/translations/gulistan/collation.md, workshop/regimes/R41-synonym-enumeration.md, tools/rhyme_pairs.py, framework/v0.2/README.md, config/models.md, wiki/goodness-senses.md
sensesstyle-correspondence

E-20260822 — is a rhyme actually available among the ordinary synonyms, and is it available only where the source rhymed?

Track T5, ARM-synonym-reach step 1. Study limb of the Gulistan دیباچه, whose prose was rendered whole under R41 and whose translator's log was frozen at 04f457ad before this design was written. The locus inventory was frozen one commit earlier, before any English for the span existed.

1. The claim under test, quoted

framework/v0.2 §7.24 item 2, published 2026-08-21 with no test behind it:

at a locus where the source rhymes, do not ask whether the first English word chimes; enumerate the ordinary synonyms for each member and look for a pair that does. The available rhymes live there.

It came out of one measurement on one Arabic proverb — العواقب is consequences to two blind seats, and both Burton and this project's own hand nevertheless wrote ends, to chime with friends. That is one locus, and it licensed a general instruction. This design asks the two questions the instruction needs answered before a practitioner should spend time on it:

(a) What is the yield? At a rhymed locus, how often does the enumeration actually put a chiming pair within reach — compared with the first, most ordinary word, which is what §7.24 tells the translator not to rely on?

(b) Is the yield about the source at all? English has a large lexicon. If enumerating six synonyms per member turns up a chime as readily at places the source left plain, then what §7.24 describes is a property of English and following it would be inventing ornament, which framework/v0.2 §7.14 and §7.20 already warn against in the strongest terms the handbook uses.

2. Why a new source language

§7.24's limits say Arabic→English only, one night of one work. Sa'di's دیباچه is the densest saj' in Persian and the passage the form's reputation rests on. Persian is the project's sixth source language for a long work and its first here, so a result on it is a second language family for the claim, not a replication of the same one.

3. Materials

workshop/translations/gulistan/loci-frozen.md, in full. 36 loci, 76 members, drawn by a salted hash from a pool of 76 before any English existed:

class n what it is
STEM 12 rhyme not generable by a shared productive ending — جهل/سهل, انیس/جلیس
AFFIX 12 rhyme carried by a class-general ending — نشینم/چینم, نُزهت/فُسحت
CONTROL 12 parallel cola of the same text that do not rhyme — موجود/واجب, کجاوه/حجره

RHYMED = STEM ∪ AFFIX, n = 24.

4. The instrument that decides "chime"

Not a model. tools/rhyme_pairs.py, built this session as a declared gate, decides every chime mechanically from CMUdict under a rule fixed before any list was drawn, and reports STRICT and NEAR separately and never summed. Its fixtures (tools/tests/test_rhyme_pairs.py, 49 checks, 0 failures) were written before it scored anything and include §7.24's own paradigm case both ways round: ends / friends must come back STRICT and consequences / ends must come back NONE.

This removes the defect RS-20260821c §7 declares in its own limits — the instrument counts a near-echo — and it removes it in the only direction that matters here, since the quantity being measured is a count of chiming pairs in a list.

5. Procedure

Seats are config/models.md roles; slugs are logged as provenance. P4, P5 and GL are out (notes (bne), (bps)); P1 P2 P3 QR are what is left and there is no fourth reading seat to buy, which is stated in the limits rather than implied away.

Stage G — the source-side gate. 36 loci × 3 seats (P1 P2 P3), one call each. The seat is shown the Persian cola of one locus with the member words marked, and asked whether the marked words rhyme in Persian: RHYME / NO / UNSURE. No hypothesis, no counts, no English, no mention of translation, and the RHYMED and CONTROL loci are shuffled together into one indistinguishable stream.

Stage A — the first ordinary word. 76 members × 2 seats (P2 P3), one call each. The seat is shown one member at a time, in its own colon, with the sibling members and their cola removed, and asked for the single most ordinary English rendering. This is RS-20260821c's stage-A protocol unchanged, so the two runs' first-word figures are comparable. The masking is the whole validity of the stage: a seat that can see both members can chime on purpose.

Stage B — the enumeration. The same 76 members × the same 2 seats, same masking, a separate call: up to six ordinary English renderings, ranked, each of which the seat would accept as an unmarked rendering of that word in that colon. This is §7.24 item 2's procedure, executed by a hand that cannot see what it would need to rhyme with.

Scoring, mechanical. For each locus:

Both computed by rhyme_pairs.py. The scored quantity is seat-specific (§6a.2): a locus counts as reached for a seat only if that seat's own lists contain the chiming pair. Pooled figures are reported and none is scored. STRICT and STRICT+NEAR are reported separately and never summed.

Stage C — the usability screen. For each (seat, locus) that is reach_syn and not reach_first, exactly one chiming pair is screened — the pair whose two words have the lowest summed rank in that seat's own lists — and both of its members are screened (§6a.4). A blind seat is shown the Persian colon and the candidate word and asked whether the candidate is an acceptable ordinary rendering of that word here: SAME / DIFFERENT. 2 seats. A pair counts as usable only if both members come back SAME from both seats. Four planted-error items are mixed in, each substituting a word that plainly changes the sense; a screen that does not catch them is not a screen.

Cross-call repeat control. 8 stage-B members are dispatched a second time to the same seat. Measures how much of reach_syn is a stable property of the lists and how much is sampling.

6. Registered predictions, bars and failure criteria

Written before dispatch. Bars are not moved after the fact.

Amended before dispatch on the pre-run critic's three BLOCKING and two MAJOR findings; the amendments are §6a and the response is critic-response.md.

statement bar
G the gate: the seats can tell the rhymed loci from the controls in the Persian majority-RHYME at ≥ 20 of 24 RHYMED and majority-NO at ≥ 10 of 12 CONTROL
P1 the yield: enumeration reaches a chime more often than the first word does reach_syn(**STEM**) − reach_first(**STEM**) ≥ **+0.25**, STRICT, seat-specific
P2′ the source's own pairings beat the same lists re-paired at random — and nothing wider, per round 2 reach_syn(**STEM**) − reach_shuffled(**STEM**) ≥ **+0.25**, STRICT, seat-specific
P2 the pair-level-matched half of the same question: parallel cola of this text that do not rhyme reach_syn(STEM) − reach_syn(CONTROL) ≥ **+0.25**, STRICT, seat-specific
P2-joint the source-side claim itself, and it needs both made only if P2 and P2′ both clear +0.25; if they disagree, the disagreement is the result and no source-side claim is made
P3 usability: the found chimes are words a translator could actually use ≥ 0.50 of screened pairs called SAME on both members by both seats
P4 the class split, secondary reach_syn(STEM) > reach_syn(AFFIX)

6a. The five amendments, and what each one gives up

  1. BLOCKING 1 accepted in full. The primary condition is STEM, not STEM ∪ AFFIX. The manifest itself says the instruction can only be about STEM loci, and scoring the union would let a pass or a fail be driven by loci that do not instantiate the mechanism. The cost is power: the primary now rests on n = 12 loci × 2 seats. That is stated in the limits and not hidden. RHYMED figures are still reported, as secondary.
  2. BLOCKING 2 accepted in full. The aggregation is declared here, before dispatch: reach is SEAT-SPECIFIC. A locus counts as reached for a seat only if that seat's own lists contain a chiming cross-member pair; the rate is over 24 seat-loci per class. Pooling across seats would let a pair be assembled from one hand's word for one member and another hand's word for the other, which is available to no translator. Every pooled figure in the result is descriptive and none is scored.
  3. BLOCKING 3 accepted, and remedied differently and better. The critic is right that the CONTROL loci differ from the rhymed loci in more than rhyme — part of speech, frequency, polysemy and synonym-set size all move together with saj' selection, and all of them drive English synonym reach. Its own remedy, controls matched member-by-member on five variables, is not buildable here. What replaces it holds every one of those variables fixed exactly: reach_shuffled re-pairs the members across loci within a class — member 1 of locus X against member 2 of locus Y — and recomputes reach on the same seats' same lists. The only thing destroyed is the pairing the source made. 2,000 re-pairings, seeded, no API cost. P2′ is now the primary source-side test and P2 is secondary. In addition, mean synonym-set size and part-of-speech mix are reported per class, so the confound the critic names is visible rather than argued away.
  4. MAJOR 4 accepted in full. Stage C screens BOTH members of every scored pair, and P3 counts a pair as usable only if both members pass with both seats. To remove the selection the critic identified, the pair screened is fixed by rule: for each (seat, locus) newly reached, the single chiming pair whose two words have the lowest summed rank in the seat's own lists. No cap, no discretion.
  5. MAJOR 5 accepted in full. The estimand is named: General American rhyme availability as CMUdict assigns it. Rhoticity and vowel mergers make some pairs dialect-dependent, and the synonym prompts do not name a variety. The result claims availability under this dictionary's convention and nothing wider.

A second critic round WAS bought, on the amendments, precisely because reach_shuffled was a component the critic had not seen — and it earned its $0.03: round 2's BLOCKING 1 established that reach_shuffled cannot carry the "source rhyme matters" reading at all, only the narrower one now written into P2′. Round 2 returned 1 BLOCKING and 2 MAJOR; all three are disposed of in critic-response.md, one of them accepted-in-part with the refusal written out.

6b. reach_shuffled, specified to the level round 2's MAJOR 2 requires

F1 — the withdrawal criterion. If reach_syn(RHYMED) − reach_first(RHYMED) < 0.10, §7.24 item 2 is withdrawn as an instruction and the section says so.

F2 — the manipulation check. If the stage-B lists add fewer than 2 new words per member on average over stage A, the two stages are not different treatments and the P1 comparison is void and withheld.

F3 — the instrument check. If more than 15% of stage-A ∪ stage-B expressions have a rhyme-bearer absent from the dictionary after the two declared repairs, the mechanical rule is not deciding the question; every figure is then descriptive and no prediction is scored.

F4 — the recall screen, and it replaces a screen that could not have fired. The obvious leak screen — does the answer quote the masked sibling — cannot fire here: the prompt is Persian and the answers are English, so there is nothing to match. The real exposure is different and larger. The دیباچه is one of the most famous pages in Persian, and a seat that recognises it could supply a remembered published translator's chime rather than an available synonym, which would inflate reach_syn(RHYMED) with no lexical availability behind it. Two things are therefore done, both declared here:

If G fails, every primary is withheld and the run reports what it saw with no claim attached.

7. What this design cannot do

8. Budget

Today's ledger (UTC 2026-08-22) is $5.00 unspent. Worst case built from the caps the requests permit, per note (abc), and assuming every call needs the doubled-cap re-dispatch, per notes (bqs) and (bqk):

stage seats calls first pass every call re-dispatched
critic P1 1 0.040 0.080
G P1 P2 P3 108 0.113 0.170
A P2 P3 152 0.122 0.184
B P2 P3 152 0.209 0.314
C P2 P3 ~72 0.058 0.087
repeat + recognition P2 P3 24 0.024 0.036
total ~509 0.566 0.871

Experiment ceiling $1.60. Runner ceiling $1.40, stop-loss $1.20. Both guards sit above the absolute worst case, which is the point: if one fires, the arithmetic was wrong.