Repository path: wiki/findings/results/RS-20260822b-echo-threshold.md · rendered 2026-09-09
Page metadata (front matter)
| type | result |
|---|---|
| id | RS-20260822b-echo-threshold |
| status | frozen |
| created | 2026-08-22 |
| updated | 2026-08-22 |
| senses | style-correspondence |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-echo-threshold.md, workshop/experiments/E-20260822b-echo-threshold/design.md, workshop/experiments/E-20260822b-echo-threshold/critic-response.md, workshop/translations/gulistan/R42-v1/translation.md, workshop/translations/gulistan/R41-v1/translation.md, workshop/regimes/R42-declared-endrhyme.md, tools/rhyme_pairs.py, framework/v0.2/README.md, wiki/findings/results/RS-20260822-synonym-reach.md, wiki/goodness-senses.md |
A slant chime is registered — but not the three slants I published as unhearable, and not by ear alone
ARM-echo-threshold step 1, track T3. E-20260822b-echo-threshold. 773 bodies, 0 dead in the
primary stage, $1.195830900 against a declared ceiling of $1.90 ($1.064523400 of it data, $0.131308 the two critic rounds). Verifier 332 checks, 0 failures,
4 mutation tests. Pre-run critic two rounds, 10 findings, 4 BLOCKING, nine accepted and one
refused in writing — and round 2 caught the confound that would have made the whole run about
spelling.
Tier D NOT PASSED. No jury was involved and nothing here is a quality claim about any rendering.
internal-judgment-only.
1. What was under test — a sentence of my own
RS-20260822-synonym-reach §4 printed seven English pairs that the project's mechanical rule counts
as chimes and asserted, in the lead's voice, with nothing behind it:
No reader hears any of them as a chime.
framework/v0.2 §7.25 item 3 rests on that sentence. It is the reason a practitioner is told that
the relaxed form of the enumeration instruction points at a property of English rather than at his
author. This run was built to find out whether it is true.
2. The headline
It is too strong as a general claim, and exactly right about the three pairs it named.
| Δ (interaction) | 95% CI | loci | |
|---|---|---|---|
M1 full rhyme — the gate |
+0.578 | [0.311, 0.800] | 15 |
P1 slant — the primary |
+0.350 | [0.183, 0.517] | 20 |
P4 the three pairs §4 printed |
0.000 | — | 3 |
M1PASSES at +0.578 against a registered bar of +0.40. The task detects a full rhyme.P1HOLDS at +0.350 against a registered bar of +0.25, with the interval clear of zero. A slant at a phrase-end is registered, andRS-20260822§4's sentence andframework/v0.2§7.25 item 3 are softened by name in §7.26.F1does not fire. The registered strengthening criterion required Δ < 0.10; the run went the other way and the handbook does not get the hard rule it was set up to earn.- And
P4is 0.000. On scattered / suspended, creatures / entities and spread / tend — the three pairs the published sentence actually named — the slant buys nothing: +0.000, −0.333, +0.333. The sentence was right about its own examples and wrong as a generalisation, which is a more useful correction than either "it stands" or "it falls".
Raw cell means, primary locus set: aZ 0.800 · aX 0.600 · aY 0.267 · bY 0.167 ·
bX 0.150 · bZ 0.133. The floor sits near 0.15 and is nowhere near the 0.60 saturation that
failure criterion 4 would have withheld the run for.
3. The finding a translator can use, and it is about verse
P2 holds, and it is the largest single split in the run.
| Δ slant | 95% CI | loci | |
|---|---|---|---|
| VERSE — couplets and quatrains | +0.542 | [0.250, 0.792] | 8 |
| PROSE — parallel cola | +0.222 | [0.056, 0.389] | 12 |
| difference | +0.320 | (bar +0.15) |
At a verse line-end a slant does most of the work a full rhyme does; at a prose colon-end it does about two-fifths of it. The two strata are the same translator, the same work, the same session and the same mechanical rule; what differs is that the verse line-end is a position where a rhyme is expected and the prose colon-end is not. This is the first quantity the project has that separates the two, and it says the answering move a translator has been treated as one move is two.
And the grading is unambiguous. Of the 37 slant detections that carried a strength verdict,
35 were called half; of 46 full-rhyme detections, 1 was. The seats do not mistake a slant
for a rhyme. They register it and they name it as a half-chime — which is what a translator is
buying, and is worth less than a rhyme and more than nothing.
4. The finding that qualifies everything above
The seats read text, not sound, and they call a spelled-alike non-rhyme an echo almost as readily as a real one. Stage P, twelve dedicated items, added on the round-2 critic's BLOCKING 1:
| arm | example pairs | hit |
|---|---|---|
PHON — full rhymes spelled unlike |
blade / weighed, near / austere, just / nonplussed | 0.778 |
ORTH — non-rhymes spelled alike |
sword / word, beard / heard, cough / dough | 0.722 |
The registered veto — if hit(ORTH) > hit(PHON), the phonetic reading of every figure is
withdrawn — does not fire, by 0.056. It very nearly did. And by seat it is not close to uniform:
| seat | PHON |
ORTH |
|---|---|---|
P1 |
1.000 | 0.500 |
P2 |
1.000 | 0.833 |
QR |
0.333 | 0.833 |
QR is doing orthography and not phonology — it calls sword / word and beard / heard echoes
and blade / weighed and near / austere not. One of the three seats is, on this task, a
spelling-matcher.
Two things pull the other way, and they are why P1 is reported rather than withheld.
- Within the item set, detection does not rise with spelling overlap. Grouping the 20 primary loci by the shared final-letter run of their slant pair: 0 letters → Δ +0.333 (n=4), 1 → +0.250 (n=8), 2 → +0.444 (n=9), 3 → +0.000 (n=2), 4 → +0.167 (n=2). The two most spelled-alike slants are the two that buy the least. (Declared post-hoc; the cell counts are small.)
- Inside the 94
NONEcells the covariate does behave: hit is 0.141 at 0 shared letters, 0.233 at 1, 0.667 at 2 (n=3). Orthography is doing real work at the floor — it is just not what is producing Δ.
What this licenses, stated exactly. The slant is registered; whether it is heard or seen this run cannot say, and one of its three instruments is demonstrably reading with its eyes. For a translator that distinction is not academic: a chime that works on the page may not survive being read aloud. The handbook sentence in §7.26 is written in the word registered, not heard.
5. The controls, including the one that failed
C1naturalness runs against the finding, which strengthens it. Mean rank of the four core cells (1 = most natural, 4 = least):bY1.84 ·aY2.14 ·bX2.82 ·aX3.09. The cell carrying the slant is ranked the least natural English of the four, by 0.82 of a rank. The registered confound criterion fires only ifaXis ranked better; it is ranked worse, so detection is happening despite the item reading least naturally, not because of it.C2, the fluency screen, FAILED as an instrument and its exclusions are not applied. It caught 3 of 4 planted breaks against a registered bar of 4 of 4 — it passed "the footing must be laid before the therefore" as fluent — and it judged 76 of 138 ordinary items not fluent, on grounds that are about register, not grammar: "Archaic and literary phrasing, not modern or ordinary English", "uses archaic, poetic phrasing". The Gulistan's English is elevated and the screen was asked for ordinary. This is a declared departure from the registered rule. Applying it as written would drop 22 of 25 loci and leave the primary uncomputable. What follows is that the item set's fluency is unverified, and that goes in §7 as a limit rather than being smoothed over.C4mechanical: 138 of 138 planted relations recomputed and matched, and the crossover re-checked cell by cell — exactly oneNEARper locus ataX, exactly oneSTRICTataZ, every other cellNONE.C5filler shape: mean filler length 5.74 / 5.88 / 5.70 characters atSTRICT/NEAR/NONE. The levels are not distinguished by word length.C3pool provenance: 11 of 25 slant fillers and 9 of 25NONEfillers are attested in S211's blind-seat lists, including all six at thePNloci, which is what makesP4a like-for-like test of the printed sentence.- Dead cells: 0 of 414 in stage D. Stage U lost 8 of 57 (14%) to truncation and is reported with that caveat.
6. P3 — do the seats manufacture echoes? Less than the bar, but they do it
Across the 94 NONE cells the seats answered yes at 0.248. They named 70 pairs there, of which
63 (90%) the rule scores NONE — genuinely manufactured — giving a manufacture rate of
0.223 against a registered bar of 0.30. The claim does not fire.
The manufactured pairs are worth printing, because they are what a hand hunting for echoes talks itself into: ability / kindness · thought / drowned · spread / nurse · green / tangle · power / constancy · remedy / retreat · age / name.
The other 7 are the opposite case and are equally worth having — pairs the seats found that the rule
agrees with and the design had not planted: dear / fear (STRICT), friend / end (STRICT),
rulers / scholars (N1), said / heeded (N2). Five of the seven fall in the five loci the
unplanted-echo screen had already flagged, which is the screen working.
7. P6 — the unprompted rate, bought instead of assumed
The round-1 critic struck the phrase upper bound as an assumption dressed as a measurement, so it
was bought. P1, 57 items, the same passages under an open prompt that never mentions sound:
| cell | names both planted words | mentions sound at all |
|---|---|---|
aZ full rhyme |
0.389 | 0.667 |
aX slant |
0.267 | 0.333 |
aY floor |
0.188 | 0.313 |
Even a full rhyme at a line-end is spontaneously remarked on in fewer than two cases in five, against 0.800 when the seat is told to look. The prompted figures are roughly twice the unprompted ones at every level, so the direction the withdrawn phrase assumed is right — and now measured. The slant's unprompted margin over the floor is +0.079, which on n = 15–16 per cell is not distinguishable from nothing.
8. What this changes in the handbook
framework/v0.2 §7.26, and it is a correction of my own sentence plus two new facts:
RS-20260822§4's "No reader hears any of them as a chime" is withdrawn as written, and §7.25 item 3 is softened. A slant at a phrase-end is registered at +0.350 above a matched floor.- The sentence was right about its own three examples. scattered / suspended, creatures / entities, spread / tend buy 0.000. The instruction that follows is a distinction, not a permission: some slants carry and these do not, and the handbook should say which rather than licensing the class.
- Position decides most of it. A slant at a verse line-end is worth +0.542 and at a prose colon-end +0.222. If you are answering saj' in prose, a slant is worth about a third of a rhyme; in verse it is worth most of one.
- What you buy is named a half-chime, not a rhyme — 35 of 37 times.
- Registered, not necessarily heard. One instrument in three is matching spelling. Say registered; do not say heard.
9. Limits
- Three model seats, no human reader, and none is available to buy (
NEXT.md, named-not-built). Every figure is a figure about models reading English. The whole item set and every seat verdict are inworkshop/experiments/E-20260822b-echo-threshold/so that a human can check any cell — which is the most this project can do here, and is not the same as having checked it. - The estimand cannot be pure sound, and no design could make it so. The relation is present in exactly one cell, and that cell is also the only one holding that particular pair of words; a pair cannot be simultaneously a rhyme and not a rhyme. The crossover removes every main effect of either word. What is left is replication across lexis: 14 of 20 loci positive, 4 zero, 2 negative, twenty unrelated word-pairs with nothing in common but the planted relation. The round-2 critic's demand for comparable lexical pairings in both cells was refused in writing on exactly this ground.
C2failed and the item set's fluency is therefore unverified (§5).- 20 constructed loci from one work in one language pair. They are not drawn from a population of English phrase-ends, so the bootstrap interval describes instability over these loci and nothing wider. Any reading beyond them is inference, not measurement.
- The NEAR classes are unbalanced — N1 ×5, N2 ×17, N3 ×3 — structurally, because
N1arises almost only between polysyllables andSTRICTbetween polysyllables is nearly unreachable. Δ by class: N2 +0.353, N3 +0.222, N1 +0.200; all positive, none carrying a claim alone. - The phrase-end screen is punctuation-based while the prompt leaves phrase undefined, so it bounds only the risk it can see.
- The label-reshuffle reference distribution returned 0.0072 for Δ_N and 0.0009 for Δ_F. It is not a randomisation p-value — the levels are determined by the words, not assigned at random — and no criterion here rests on it.
- The lead wrote every item and holds the hypothesis. The seats are blind, the levels are mechanical, and the pre-run critic was bought twice before any data; that is the mitigation, not a cure.
- A useful by-product, and a caution for the project's own instruments: of 30 classic English
orthographic traps,
tools/rhyme_pairs.pyscores 25 asNEAR— near / bear, call / shall, comb / bomb, good / blood. TheNEARclass and the spelled-alike class very largely coincide, which is why stage P could only be built from a narrow residue. - Tier D NOT PASSED; every self-assessment
provisional.
10. The wire to the translation limb
T-gulistan-R42-v1 — Sa'di's دیباچه verse, 79 bayts, 979 Persian words → 1,508 English words,
rendered under the new R42 and frozen at 3108ee96 before this design was written — is where the
question came from and where the answer lands. The regime forces a graded verdict at every bayt, and
the hand ended with 43 STRICT, 4 NEAR, 7 NONE and 2 the dictionary could not grade, plus
eight AVAILABLE-REFUSED where a chime was reachable and was declined.
Four of the span's rhyme positions are slants, and this run prices them. All four stand at verse line-ends — b036–b037, b051, b095, b106–b107 — which is the position where a slant is worth +0.542, most of a rhyme. That is the arm's answer to the translator's question, and it says the four slants were worth writing. The reading does not transfer to the prose span rendered the same day: at a colon-end the same move is worth +0.222.
It also says something about the eight refusals. Every one was declined under rule 2, six of them because the chime needed an end-filler the Persian has not got. At +0.542 a verse slant is cheap and carries; an invented filler is neither. The refusals look better after the measurement than they did before it.