Repository path: workshop/experiments/E-20260822b-echo-threshold/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260822b-echo-threshold |
| status | frozen |
| created | 2026-08-22 |
| updated | 2026-08-22 |
| senses | style-correspondence |
| provisional | true |
| internal-judgment-only | true |
| links | wiki/arms/ARM-echo-threshold.md, wiki/findings/results/RS-20260822-synonym-reach.md, workshop/translations/gulistan/R41-v1/translation.md, workshop/translations/gulistan/R42-v1/translation.md, tools/rhyme_pairs.py, config/models.md, config/budget.md, framework/v0.2/README.md |
E-20260822b — how much sound-relation does an English pair need before it is heard as an echo
ARM-echo-threshold step 1, track T3. Frozen before dispatch. Nothing below is revised after the
first call; amendments accepted from the pre-run critic are recorded in critic-response.md with the
commit that carried them, and the frozen text stands as written.
1. The question, and why it is not a question about our instruments
RS-20260822-synonym-reach §4 printed seven English pairs that a mechanical near-echo rule counts as
chimes — spread / tend, creatures / entities, scattered / suspended, put on / set,
eat / are bought, does not persist / is inappropriate, delight / comfort — and asserted, in the
lead's voice, with nothing behind it:
No reader hears any of them as a chime.
That sentence is load-bearing. framework/v0.2 §7.25 item 3 tells a practitioner, on the strength of
it, that the relaxed form of the enumeration instruction points at a property of English rather than
at his author. The question here is whether it is true, and the unit of analysis is a reader's
report about a passage of English, not a property of tools/rhyme_pairs.py, whose arithmetic is not
in doubt and which is used here to define the independent variable, not to measure anything.
What this teaches about translating literature: at what degree of phonetic relation an English pair standing at two phrase-ends is recoverable as an echo by a competent reading instrument that is looking for one — which is what decides whether a translator answering a source rhyme with a slant is carrying his author's music across or decorating.
2. The estimand, stated so it cannot be overclaimed later
Detectability of a planted phrase-end sound relation by three non-Anthropic reading seats whose attention has been directed at sound.
The word reader does not appear in any prediction, criterion or consequence below (critic BLOCKING 1, amendment A1). The seats are models; every figure is a figure about models reading English, and the handbook consequence in §6 carries its warrant on its face.
Three things this is not, said now rather than in the limits:
- It is not what an unprompted seat picks up. Stage U measures that rather than assuming it (critic MAJOR 4, amendment A4): the original text of this section claimed the prompted figures were an "upper bound" on ordinary reading, which was an assumption dressed as a measurement. The claim is withdrawn and the quantity is bought.
- It is not a human measurement. No human readers exist to buy (
NEXT.md, named-not-built) and the three seats are models. Every item and every seat verdict is printed in the result page's appendix so that a human reader can check them, which is the most the project can do here. - It is not a claim about any translation's quality. No seat is asked to judge anything.
3. Materials
138 items over 25 loci, in a 2 × 3 crossover (critic BLOCKING 2, amendment A2). Each locus is one short English passage with two parallel phrase-ends. One end takes the fixed word, the other the filler; the variants of a locus differ in exactly one word from their row- and column-neighbours.
Each locus supplies two fixed words and two or three fillers:
| filler X | filler Y | filler Z | |
|---|---|---|---|
| fixed a | NEAR — the only slant cell |
NONE |
STRICT — the only full cell |
| fixed b | NONE |
NONE |
NONE |
b is a second ordinary word for the same slot, chosen so that tools/rhyme_pairs.py scores it
NONE against every filler at that locus. All levels are assigned by the tool and verified by
the verifier before dispatch.
Why the crossover is the design and not a refinement. In a plain three-level comparison the
phonetic level travels with the lexical identity of the filler — its meaning, frequency, register,
collocation, morphology and syntactic fit — and no naturalness control can separate them. Here
filler X stands once in a NEAR cell and once in a NONE cell; filler Y stands in two NONE
cells; fixed words a and b each stand in every column. Any main effect of either word cancels
in the interaction, and the phonetic relation is present in exactly one of the four core cells.
19 loci carry all six cells; six loci (PN1–PN6) carry four, because STRICT is unreachable
there from any ordinary rendering — which is S211's own headline — and their X/Y fillers are
drawn whole from the S211 blind-seat pool, including the three pairs RS-20260822 §4 named
(scattered / suspended, creatures / entities, spread / tend).
Counts: NEAR 25 cells, STRICT 19, NONE 94. By class: N1 ×5, N2 ×17, N3 ×3.
3.1 Strata
- PROSE, 14 loci (
PR*,PN*) — sentences carrying the sense of the Gulistan دیباچه's prose, in the English ofT-gulistan-R41-v1where the constraint allows, minimally rearranged so that each member stands at its phrase's end. The rearrangement is identical across every cell of a locus and therefore cannot differ by level or by fixed word. - VERSE, 11 loci (
VE*) — couplets and quatrains fromT-gulistan-R42-v1, frozen at3108ee96before this design was written.
3.2 The items
The full manifest is materials/items.json — every carrier, both fixed words, every filler,
every rule verdict, the bearers the rule used, the phrase-end word list, the unplanted-echo screen
and the pool provenance flag — and it is printed whole in the result page's appendix so that a human
can check any cell. Two loci, given here so the shape is legible:
PR4(PROSE) — "The one comes {fixed} to artifice; the other stands {filler}." fixed a = near, b = close; fillers X = far (N2), Y = remote, Z = clear (STRICT).
VE10(VERSE) — "What is the tongue in the mouth, O wise man? / The key to the treasure-door of a man of {fixed}; / and while the door stays shut, who is to know / whether he trades in jewels or in {filler}?" fixed a = skill, b = craft; fillers X = wool (N2), Y = silk, Z = twill (STRICT).
VE2 is the one locus whose filler slot is at the first phrase-end rather than the second;
nothing depends on which end varies, and it is named so the asymmetry is on the record.
3.3 The NEAR classes are not balanced, and that is declared here
N1 ×5, N2 ×17, N3 ×3. The imbalance is structural, not a choice: N1 arises almost only between
polysyllables, and STRICT between polysyllables is nearly unreachable — S211's headline again. The
run measures NEAR as a class, with a per-class breakdown reported and no class carrying a claim
on its own.
3.4 Unplanted phrase-end relations, screened mechanically
Critic MAJOR 5, amendment A5. Every passage is segmented at line breaks and at ; : , . ? !, the
last word of each segment taken, and every pair of those words scored by the rule.
14 of 138 items, spanning 5 of 25 loci —
PR6,PN3,VE1,VE8,VE11— carry an unplanted phrase-end relation. Every one isNEAR; not one isSTRICT.
A registered sensitivity analysis recomputes every criterion with those five loci dropped, and both figures are reported side by side.
4. Procedure
Five stages. All calls temperature 0, one item per call, isolated context, dispatched in one shuffled stream per stage with seed 20260822.
Seats: P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, QR qwen/qwen3.7-max.
P3 is excluded on the cost finding of S211 (NEXT.md), P4 on note (bps), P5 on note (bne),
GL on LONG prompts. There is no fourth reading seat to buy and the result will say so rather than
implying a panel.
Stage D — prompted detection (138 items × 3 seats = 414 calls)
The seat sees the passage and nothing else. No mention of Persian, of Sa'di, of translation, of rhyme regimes, or that variants exist.
Read this passage.
<PASSAGE>
Does the passage contain a sound echo -- two words that chime with each other in
sound -- standing at the ends of its phrases or lines?
Answer in exactly this form and add nothing else.
ECHO: yes
PAIR: <word> / <word>
STRENGTH: full
-- or --
ECHO: no
If you hear more than one such pair, give at most two, separated by a semicolon.
STRENGTH is "full" or "half".
max_tokens 300, reasoning cap 100. The yes branch of the template is shown first, which can
only raise the yes rate; it is identical for every item, so it moves the floor and the cells
together and cancels in the interaction. Failure criterion 4 guards the case where it saturates.
Stage U — unprompted, added on critic MAJOR 4 (57 items × 1 seat = 57 calls)
The aZ, aX and aY cells of the 19 three-level loci, to P1. The word sound does not occur
in this prompt, and neither does echo, rhyme or chime.
Here is a short passage.
<PASSAGE>
Is there anything you notice about how it is written? Answer in one sentence,
then stop.
Coded mechanically: does the sentence name both planted words? Secondary coding: does it use any
of rhyme / chime / echo / assonance / alliteration / sound? max_tokens 400, reasoning cap 120.
Stage R — naturalness ranking (25 loci × 2 seats = 50 calls)
The four core cells of each locus (aX, aY, bX, bY) shown together, letters assigned by a
per-(locus, seat) shuffle from the seed.
Below are 4 versions of the same short passage.
A) ...
B) ...
C) ...
D) ...
Rank them by how natural and ordinary the English reads, best first.
Answer in exactly this form and add nothing else, using each letter once.
RANK: <letter> > <letter> > <letter> > <letter>
Seats P1 and P2. The answer template names no concrete order, and the letter→cell map is shuffled
per (locus, seat), so any residual letter preference spreads over cells rather than onto one of them.
The Z cells are not ranked; M1 is a gate, not a finding, and that is declared.
Stage S — fluency screen (138 variants + 4 planted breaks = 142 calls, seat P2)
Is the following a fluent, ordinary piece of English that makes sense?
<PASSAGE>
Answer in exactly this form and add nothing else.
FLUENT: yes
-- or --
FLUENT: no
REASON: <at most twelve words>
The 4 planted items replace the filler with a word that cannot stand in the slot — a syntactic break, not a semantic oddity.
5. Scoring, fixed before dispatch
Per (item, seat):
hit—ECHO: yesand the two planted words (the locus's fixed word and that item's filler) both appear inside one named pair, matched case-insensitively on the last token of each side after stripping punctuation. For a multi-word filler (relied on) the bearer token governs, asrhyme_pairs.pydefines it.yes_any—ECHO: yes, whatever pair was named.other—yes_anyand nothit: the seat heard an echo somewhere else in the passage.half—hitandSTRENGTH: half.
Unparseable or empty → one re-dispatch, then dropped and counted as dead. Dead cells are reported and are not imputed.
6. Predictions and criteria, registered
Write hit(f g) for the mean hit over (locus, seat) cells with fixed word f and filler g. The
primary quantities are interactions, so that any main effect of either word cancels (amendment
A2):
Δ_N = [hit(aX) − hit(aY)] − [hit(bX) − hit(bY)] Δ_F = [hit(aZ) − hit(aY)] − [hit(bZ) − hit(bY)]
Both are computed on the 20 loci that survive the unplanted-echo screen (amendment A9, critic
round 2 MAJOR 3 — the round-1 refusal gave a false reason and the critic caught it). Δ_F uses the 15
of those that carry a Z cell. The 25-locus and 19-locus figures are reported beside them, and
if the two disagree about which criterion fires, the 20-locus figure governs and the disagreement
is the headline.
M1— manipulation check, and a gate on everything else. Δ_F ≥ 0.40, with the bootstrap interval excluding 0. IfM1fails, the run is reported as uninterpretable and no claim whatever is made aboutNEAR.P1— the primary. Δ_N ≥ 0.25, bootstrap interval excluding 0. Holding means the slant is recovered, andRS-20260822§4's sentence andframework/v0.2§7.25 item 3 are too strong and must be softened by name.F1— the registered strengthening criterion. Δ_N < 0.10 and the bootstrap interval's upper end below 0.25. Then §7 gains a positive instruction with a measurement under it: at a phrase-end, a slant is not recovered as an answer to the source's rhyme; write a full rhyme or write none — carrying, on its face, the warrant three model seats, no human reader, one work, one language pair.- Between 0.10 and 0.25, or an interval spanning both bars, the run is indeterminate and step 2 of the arm is designed to resolve it. The band is registered so a middling result cannot be written up as either outcome.
P2— position, secondary. Δ_N in VERSE minus Δ_N in PROSE ≥ 0.15 → verse licenses a slant that prose does not.P3— the invention number, split (amendment A11).yes_anyover the fourNONEcells, with every named pair rescored bytools/rhyme_pairs.pyand the rate split into (i) the named pair isSTRICT/NEARby the rule — the seat found an unplanted relation the rule agrees with — and (ii) the named pair isNONEby the rule. Only (ii) is reported under the word manufacture, and only (ii) carries the registered claim at ≥ 0.30, which is the receiving-end face offramework/v0.2§7.14 and §7.20. Every pair is printed.P4— the three printed pairs. Δ_N restricted toPN4,PN5,PN6. n = 9 cells per cell type; no bar is set and none can be, because three loci cannot carry one. Reported because those three are the pairs the published sentence names.P5— grading, secondary. Share ofhits calledSTRENGTH: half, by level.-
P6— the unprompted rate (stage U). Share of items where the open prompt names both planted words, by cell. Descriptive; it decides whether the prompted figures need the "only when hunted for" caveat or not. -
P7— the orthography probe (stage P), and it can veto the whole reading.hit(PHON)— six full rhymes spelled unlike — againsthit(ORTH)— six rule-NONEpairs spelled alike. Registered: ifhit(ORTH) > hit(PHON), every figure in this run is disclosed as measuring orthographic resemblance, and the phonetic reading ofM1,P1andF1is withdrawn. The covariate is reported too: mean shared final-letter run is 2.63 atSTRICT, 1.60 atNEAR, 0.13 atNONEin the main item set, so the confound is real and large and is not being guessed at. P8— across-locus consistency (amendment A8). Δ per locus, the count of loci with Δ > 0, and a sign test across them. One lexical pairing can carry a spurious effect; twenty pairings sharing nothing but the sound relation cannot without twenty coincidences. This, and not the interaction alone, is what separates the sound relation from the particular pairing — see §12.
Uncertainty (amendment A3, narrowed by A10). The primary statement is a bootstrap over loci — 10,000 resamples with replacement, seed 20260822, percentile interval on Δ. Loci are the resampled unit; seats and cells are not.
What that interval is (A10). The loci are constructed, not sampled from a population of English
phrase-ends, so the interval describes instability over these loci and nothing wider. Every
conclusion under P1 and F1 is stated about these constructed loci and these pairs; a wider
reading is marked in the text as an inference, never as a measurement.
A within-locus label-reshuffle reference distribution (20,000 draws, same seed) is also computed and is reported under that name. It is not a randomisation p-value: the levels are determined by the words, not assigned at random, so no permutation of them can be one. No criterion above depends on it.
7. Controls and their criteria
C1— naturalness. From stage R, mean rank per core cell (1 = most natural). With the crossover the confound it guards is narrower than before: a main effect of the filler word now cancels in Δ, soC1is asked only whether theaXcell specifically — the one cell carrying a relation — is ranked better than the other three by more than 0.30 rank points. If it is,P1is reported as possibly confounded by an interaction between the two words' naturalness. A confound in the opposite direction strengthens a null and is reported as such.C2— fluency. From stage S. 4 of 4 planted breaks must be caught. Any cell judged not fluent drops its whole locus from the primary; the drop is reported with the reason.C3— pool provenance. How manyXandYfillers are attested in the S211 blind-seat pool (materials/pool.json), by stratum. No bar; it is what licenses the matched claim at the sixPNloci.C4— mechanical, run by the verifier before dispatch and again after. All 138 planted relations recomputed fromtools/rhyme_pairs.py, and the 2 × 3 structure re-checked cell by cell: exactly oneNEARper locus, ataX; exactly oneSTRICTper three-level locus, ataZ; every other cellNONE. Any mismatch and the run does not go.C5— filler shape, descriptive. Mean filler length in characters and syllables by cell.C6— unplanted phrase-end relations (amendment A5). The screen in §3.4, plus a registered sensitivity analysis recomputing Δ_N, Δ_F and every criterion withPR6,PN3,VE1,VE8andVE11dropped. Both figures are reported side by side, and if they disagree about which criterion fires, the sensitivity figure governs and the disagreement is the headline.
8. Failure criteria — what would make this run worthless
Written before dispatch, per charter §6:
M1fails. Three seats cannot recover a full rhyme at a phrase-end above the floor → the task does not measure what it is named for and nothing else is reported as a finding.C1fires on theaXcell. The primary is disclosed as possibly confounded.- Dead-cell rate above 10% after the single re-dispatch → reported as a pilot.
- The floor is at ceiling. If
hit(bX)orhit(aY)exceeds 0.60, the seats are naming the two phrase-end words whatever they are, there is no floor left, andP1andF1are both withheld. - A stratum collapses. If
C2drops more than 5 of 25 loci, the run is underpowered andF1may not fire. - The crossover degenerates. If
hit(bX) > hit(aX)andhit(bY) > hit(aY), the fixed word is driving everything and the interaction is not interpretable as a sound effect; reported as such.
9. Budget
Pre-flight built from max_tokens plus the reasoning cap, not from an assumed output length —
note (abc), the one estimate this project has overrun failed exactly there.
| stage | calls | cap (out) | worst-case |
|---|---|---|---|
| stage D, 138 items × 3 seats | 414 | 400 | $0.91 |
| stage P, 12 probe items × 3 seats (amendment A6) | 36 | 400 | $0.08 |
stage U, 57 items × 1 seat (P1) |
57 | 520 | $0.19 |
| stage R, 25 loci × 2 seats | 50 | 400 | $0.11 |
stage S, 142 items × 1 seat (P2) |
142 | 400 | $0.24 |
| pre-run critic, 2 rounds | 2 | 8500 | $0.13 |
| re-dispatch allowance, 15% at the doubled cap | ~100 | — | $0.22 |
| total | ~800 | $1.88 |
Declared experiment ceiling $1.90. Runner ceiling $1.55. Stop-loss $1.35 — the runner halts and writes what it has if billed cost passes the stop-loss.
Expected actual, from S211's measured per-call rates (P1 $0.00127, P2 $0.00079): about $0.85.
The ceiling is a guard, not a forecast.
Today's UTC ledger stands at $1.067599850 of $5.00 after S211, leaving $3.932400150. The ceiling fits and leaves $2.03.
QR has not been dispatched by this project since selection and its billed rate is unknown. Per
note (bqk) the runner prices it from its own first twenty stage-D calls and, if the measured rate
would carry the run past the runner ceiling, drops it and completes on P1+P2 — with the
arithmetic written down at the moment of the change and nothing already bought re-bought. That
decision is made before any stage-U, R or S cell exists, so no outcome can be steered by it.
10. What is deliberately not measured
- Recognition. S211 found 6 of 14 answers naming the Gulistan. It is not probed here because it cannot help a seat: the planted level is constructed by this design and is not a property of any published text. Stated so that its absence is a decision and not an oversight.
- Contamination against a published hand.
T-gulistan-R42-v1measurescleanagainst Eastwick 1852 — 0 shared 7-grams, 0 twelve-grams, longest common run 6 tokens, computed after the freeze. It is reported as provenance. No claim in this design turns on independence from a published rendering, so the standing selection gate is satisfied by the declaration and the measurement rather than by a comparison.
11. The amendments this design carries
Round 1 of the pre-run critic returned NEEDS-REDESIGN, 5 findings, 2 BLOCKING, before any data
call. All five are accepted; critic-response.md carries the reasoning and the one refusal.
| # | severity | amendment |
|---|---|---|
| A1 | BLOCKING | reader struck from every prediction and consequence; the handbook change, if F1 fires, carries its warrant on its face |
| A2 | BLOCKING | the design becomes a 2 × 3 crossover and the primary becomes an interaction; 69 items → 138 |
| A3 | MAJOR | the permutation p-value is demoted to a named reference distribution; bootstrap over loci becomes the uncertainty statement |
| A4 | MAJOR | stage U added — the unprompted rate is bought instead of assumed; "upper bound" withdrawn |
| A5 | MAJOR | every passage screened for unplanted phrase-end relations |
Round 2 also returned NEEDS-REDESIGN, 5 findings, 2 BLOCKING — and buying it was right: it
found the confound that would have made the whole run about spelling, and it caught a false statement
in one of round 1's refusals.
| # | severity | amendment |
|---|---|---|
| A6 | BLOCKING | stage P, the orthography/phonology probe — 12 items, 36 calls, with a registered veto (P7) over the phonetic reading of every figure in the run |
| A7 | BLOCKING | shared final-letter run recorded on all 138 items as a covariate; hit within the NONE cells reported against it |
| A8 | BLOCKING | per-locus Δ and a sign test across loci (P8). The critic's remedy — comparable lexical pairings in both NEAR and NONE cells — is refused in writing: no word pair can be simultaneously a rhyme and not a rhyme, so the confound is a property of the question. See §12. |
| A9 | MAJOR | the primary moves to the 20 unflagged loci; the round-1 refusal's stated ground was factually wrong and the critic corrected it |
| A10 | MAJOR | Δ and its interval declared descriptive of these constructed loci; no class-level threshold claim |
| A11 | MAJOR | P3 split by rescoring every named pair; only rule-NONE pairs are called manufactured |
12. The one thing no version of this design can fix
The sound relation is present in exactly one cell, and that cell is also the only one containing that particular pair of words. No stimulus can separate them, because a pair cannot be simultaneously a rhyme and not a rhyme. The crossover removes every main effect of either word; what remains is the interaction of the two, and a sound relation is one kind of interaction between two words.
What the design substitutes for a fix is replication across lexis. Twenty loci, twenty unrelated
pairs of English words, nothing in common but the planted relation. If Δ is carried by one or two
loci it is a fact about those words; if it is broad and consistent in sign, the alternative
explanation needs twenty independent coincidences. P8 reports exactly that, and the result page
states the limit in these words rather than implying the crossover disposed of it.