Repository path: workshop/experiments/E-20260824-eye-or-ear/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260824-eye-or-ear |
| status | frozen |
| created | 2026-08-24 |
| updated | 2026-08-24 |
| senses | style-correspondence |
| provisional | true |
| internal-judgment-only | true |
| links | wiki/arms/ARM-echo-threshold.md, wiki/findings/results/RS-20260822b-echo-threshold.md, framework/v0.2/README.md, tools/rhyme_pairs.py, config/models.md, workshop/translations/gulistan-bab1d/R46-v1/translation.md, workshop/regimes/R46-eye-or-ear.md, wiki/method-notes.md |
E-20260824 — is the chime in the sound or in the look? ARM-echo-threshold step 2
Frozen before dispatch. The translator's log of T-gulistan-bab1d-R46-v1 was frozen at commit
ee2608df, before this file existed.
AMENDED 2026-08-24 after the round-1 critic pass (critic-v1.json, NEEDS-REDESIGN, 12 findings —
ten accepted in whole or part, one refused in writing, one partly refused). The amendments A1–
A11 and the reasoning for each are in critic-response.md; they are applied below and marked
where they bite. The design was re-submitted for round 2 after they were applied.
ROUND 2 RETURNED NEEDS-REDESIGN AGAIN, and this run goes ahead under that verdict, which is
stated here rather than buried. A12–A18 apply what round 2 found and could be fixed. Its
remaining BLOCKING finding cannot be fixed: it asks for a manipulation that varies spelling while
holding pronunciation genuinely constant, and §2 and §3.2 of this design are the demonstration that
English does not supply one at a phrase-end — that is the run's own first result, not an
oversight. Its other unmet findings — no human readers, no corpus sampling frame, no independent
phonetic coding, seats chosen on prior results — are this project's standing limits, recorded in
config/models.md and carried on the face of every instruction (§11). Every one of them is printed
in the result page's limits. The primary is restated by A14 as an association, not a causal
decomposition of look against sound.
1. The question, and why it is not a question about our instruments
framework/v0.2 §7.26 item 5 tells a practitioner, in as many words, that this handbook cannot tell
him whether the chime he has written works on the ear or only on the page:
The instruments read text. On a dedicated twelve-item probe they called non-rhymes that are spelled alike — sword / word, beard / heard, cough / dough — an echo at 0.722, against 0.778 for real rhymes spelled unlike; one of the three was doing spelling and not sound at all. A chime that works on the page may not survive being read aloud, and this handbook cannot yet tell you which of yours is which.
The subject-rule sentence (wiki/tracks.md §The subject rule), written before the design:
What this unit teaches about translating literature: whether an English pair standing at two phrase-ends is registered because of how the two words sound or because of how they look — and therefore whether a translator answering a source rhyme should care that his half-chime also looks alike on the page, which is a choice he makes at the desk between two candidate words.
That is a question about the reader at the receiving end and about a decision the translator takes,
not about tools/rhyme_pairs.py, whose arithmetic is not in doubt and which here defines the
independent variable rather than being measured by it. The seats' reading mode is a limit, named
in §7 and bounded by a registered control (P3 below), not the subject.
2. What the naive design would have measured, and why it was thrown away
The first build of build.py asked for a filler alike in spelling and unrelated in sound — the
plain-English reading of eye rhyme. Twenty-one of twenty-four loci were dropped by the mechanical
check, all for one reason: tools/rhyme_pairs.py scores love/move, said/maid, done/stone,
blood/wood, height/weight, bury/jury, break/streak as NEAR-N2 — pararhyme: the final
vowels differ and the coda after them is identical.
An English eye rhyme is, in almost every case, also a consonance. The two are the same pair seen
twice, and the same is true of two of the three examples §7.26 item 5 prints by name: sword /
word and beard / heard share the coda R D / D, and the S212 probe's "non-rhymes spelled alike"
arm was therefore carrying a real sound relation. §7.26 item 5's inference — one seat was doing
spelling and not sound — may still hold for that seat, but the arm it was inferred from is not the
clean orthographic contrast it was taken to be. That is this design's first finding and it is a
gate, not the unit; it forced the design below and it is reported in the result page's §2.
3. Materials
130 items over 17 loci (A7: L18 and L19 dropped by hand after the mechanical check
passed them), built and verified by build.py, manifest at materials/items.json.
Each locus is one short English passage with two parallel phrase-ends. One end takes the
fixed word, the other the filler; every cell of a locus is the same string with exactly one
word changed, and the articles do not vary across fillers.
3.1 The main crossover — 14 loci, 8 cells each, 112 items
| filler X | filler V | filler W | filler Y | |
|---|---|---|---|---|
| fixed a | ORTH — NEAR-N2, run ≥ 3 |
CONS — NEAR-N2, run ≤ 2 |
RHYME — STRICT, run ≤ 2 |
NONE |
| fixed b | NONE |
NONE |
NONE |
NONE |
run is the number of final letters the two rhyme-bearers share (identical bearers → excluded).
The ≥ 3 / ≤ 2 split is RS-20260822b stage P's own: its PHON arm sits at 1, 1 and 2 shared final
letters, its ORTH arm at 4, 4 and 4.
X and V carry the same phonetic relation to the fixed word and differ only in whether the pair
also looks alike. Their difference is the contribution of spelling and nothing else planted in the
design varies with it. The crossover then cancels any main effect of either word, exactly as in
E-20260822b: X, V, W, Y each stand once with a and once with b. Written out, the Y
terms cancel and the primary is [hit(aX) − hit(aV)] − [hit(bX) − hit(bV)] — the X-against-V
difference minus the same difference where neither carries a planted relation, which is what the b
column measures and subtracts.
A3 (round-1 critic findings 3 and 4, partly accepted). What the b column does NOT subtract is a
difference that exists only relative to a: within the N2 class blood / wood and
blood / cloud are not equally close in sound, and the closer one may also be the spelled-alike one.
The primary is therefore orthographic likeness PLUS any residual within-N2 phonetic-distance
difference that travels with it, and every instruction it licenses says so.
A7 materials criterion, now binding rather than descriptive: at every locus X and V are
monosyllabic with primary stress on the final syllable. L18 (mountain, first-syllable stress
against monosyllables) and L19 (disyllabic machine against monosyllabic moon) were dropped for
failing it; the rule cannot see a stress or syllable-count mismatch and passed them both.
Two loci, so the shape is legible:
L1— "What ran in the street that night was the {fixed}; and what he saw when he lifted his eyes was the {filler}." fixed a = blood, b = ash; X = wood (N2, run 3), V = cloud (N2, run 1), W = mud (STRICT, run 1), Y = stone.
L12— "What frightened him most about the tower was its {fixed}; and all that the boy had between his hands was a {filler}." fixed a = height, b = age; X = weight (N2, run 5), V = boat (N2, run 1), W = kite (STRICT, run 0), Y = sack.
Three of nineteen candidate loci were dropped by the check and not repaired — L4, L13, L16
— with the failing pair printed in materials/items.json. L4 and L16 failed because their X
came back NONE rather than N2; they are re-used, in a different role, in §3.2.
3.2 The NEAR-EYE sub-arm — 3 loci, 6 cells each, 18 items, carrying no claim
sword / word, heard / beard and word / lord are the residue: alike in spelling at the rime and
scored NONE by the rule. Cells X (NEAR-EYE), W (RHYME), Y (NONE) crossed with fixed
a / b. Two of the three are the pairs framework/v0.2 §7.26 item 5 prints by name.
A2 (round-1 critic finding 1, accepted). These are NOT sound-unrelated and the sub-arm was
renamed for it. All three end in /rd/ or /d/; the tool's NONE verdict is an artefact of its coda
rule, which compares the coda after the final vowel and treats CMUdict's ER as a vowel, so
word W ER D has coda D while sword S AO R D has coda R D. P4 is withdrawn as a
claim and the sub-arm is descriptive only.
What that leaves, and it is this run's answer to the question §7.26 item 5 was really asking.
Between the 21 loci the first build dropped and this finding: A16 (round-2 critic finding 7,
accepted) — this is a failure to find, not a claim of non-existence. Under
tools/rhyme_pairs.py, CMUdict pronunciations, and the candidate-generation procedure of build.py,
the project could not construct a pair alike in spelling at the rime and unrelated in sound that
two words could both stand at a phrase-end for. No corpus was searched and no human phonologist
adjudicated. And the tool's r-coloured misclassification cuts towards scarcity, not away from it:
the three pairs it did admit turned out on inspection to share a coda. The worry is your chime
working on the page rather than the ear? is therefore hard to pose in the pure form it assumes.
3.3 The two mechanical censuses, at $0
M-avail— the availability half of the arm's question, run on the frozen English ofT-gulistan-bab1d-R46-v1bysweep.pyunder the rules stated at the head of that file: how often does a chime of each kind sit at two adjacent phrase-ends of a translator's actual English? Already run; the figure is in the result page.M-census— over a declared list of classic English orthographic traps, what fraction of spelled-alike pairs areSTRICT,N2, orNONEby sound? This is the number that decides whether the naive eye-rhyme question is answerable at all.census.py, list fixed before it was scored.
4. Procedure
Three stages. All calls temperature 0, one item per call, isolated context, dispatched in one
shuffled stream per stage with seed 20260824. max_tokens 1600 reasoning / 700 body, the setting
that ran 316 calls with 0 dead in S216.
Stage D — prompted detection, sound-framed (130 items × 3 seats = 390 calls)
The seat is shown one passage and asked, in a prompt identical for every item:
Read the passage below aloud in your head. Do the two phrase-endings chime — that is, do the last words of the two phrases echo one another in sound? Answer
yesorno. Ifyes, name the two words and say whether the echo isfullorhalf.
The prompt names sound and asks for an inward reading aloud, and that is deliberate: it is the
form of the practitioner's own question. A neutral prompt would very likely raise the ORTH rate,
and §7 says so as a limit. It is also E-20260822b stage D's framing, which is what makes the two
runs comparable.
Scoring, RESTRUCTURED BY A12 (round-2 critic finding 3, accepted). Two outcomes, and the
judgment is now the primary:
say= 1 iff the seat answersyes. This is the primary outcome. It cannot be reached by copying a visually similar pair out of the passage.hit= 1 iffsay = 1and the seat names both planted phrase-end words (case-insensitive, either order). Reported besidesaythroughout.
A claimed orthographic effect must appear in say. An effect present only in hit is reported
as an endpoint-naming effect and licenses no handbook sentence — that is exactly the artefact the
finding names. A strength verdict of full or half is recorded where given and is descriptive. Unparseable or empty → one re-dispatch,
then dropped and counted dead.
Stage N — the same question, neutrally framed (A4; 56 items × 2 seats = 112 calls)
Added on round-1 critic finding 5, accepted. A null under a prompt that says in sound and
read it aloud in your head licenses only a claim about that instruction. Stage N puts the four
cells that enter the primary — aX, aV, bX, bV of the 14 main loci — to P1 and P2 under a
prompt that names neither sound nor spelling and asks for no inward reading:
Read the passage below. Do the two phrase-endings echo one another?
Δ_ORTH − Δ_CONS is computed under both framings and both are reported.
Stage R — naturalness ranking, control C1 (17 loci × 2 seats = 34 calls)
The four (three, in the sub-arm) a cells of each locus shown together, letters shuffled per (locus,
seat); the seat ranks them for how naturally the English reads, saying nothing about sound. Seats
P1 and P2.
Seats
P1 = openai/gpt-5.6-terra, P2 = google/gemini-3.6-flash, QR = qwen/qwen3.7-max
(config/models.md). P4 and P5 are out on any task shape (notes (bne), (bps)); GL is
out on LONG prompts; P3 is priced from a probe or not used.
The primary is computed on P1 and P2 only, and this is registered here, before dispatch.
The ground is RS-20260822b stage P, which measured P1 at 1.000 phonetic / 0.500 orthographic and
P2 at 1.000 / 0.833, against QR at 0.333 / 0.833 — a demonstrated spelling-matcher on this
task, method note (bqy). QR is run anyway, as P3, a declared positive control (§6).
5. Estimands
Write hit(f g) for the mean hit over (locus, seat) cells with fixed word f and filler g.
Every quantity is an interaction, so a main effect of either word cancels:
Δ_RHYME = [hit(aW) − hit(aY)] − [hit(bW) − hit(bY)] Δ_CONS = [hit(aV) − hit(aY)] − [hit(bV) − hit(bY)] Δ_ORTH = [hit(aX) − hit(aY)] − [hit(bX) − hit(bY)] Δ_EYE = [hit(aX) − hit(aY)] − [hit(bX) − hit(bY)], on the three sub-arm loci
Intervals are bootstrap percentile over loci, 10,000 resamples, seed 20260824; the locus is the resampling unit, seats and cells are not.
A13 (round-2 critic finding 4, accepted in part). What the interval is and is not. Fourteen
hand-built loci are not a sample of English phrase-ends and two deterministic seats are not a sample
of readers. The interval is a dispersion statistic over this item set, and every criterion in §6 is
a decision rule about this item set — not an inference about English, about readers, or about
models in general. No sentence this run writes may be phrased as one.
6. Predictions and criteria, registered
M1— the gate, now two-part (A9, round-1 critic finding 11, accepted). (i) interaction: Δ_RHYME ≥ +0.30 (RS-20260822bmeasured a full rhyme at a phrase-end at +0.578); (ii) direct: rawhit(aW)≥ 0.60 and rawhit(aY)≤ 0.35. Both must hold — the interaction alone can fail or pass on an arbitrarybW-against-bYdifference. IfM1fails the task does not detect a rhyme and nothing else in this run is reportable as a finding.P1— THE PRIMARY: does spelling add anything once the sound relation is held constant? Δ_ORTH − Δ_CONS. The lead's registered prediction, made before dispatch and recorded here so it can fail: ≥ +0.15 — the lead expects the look to add, because §7.26 item 5's probe suggested the seats read by eye.- ≥ +0.15 in
say, with the interval clear of zero →A14(round-2 findings 1 and 2, accepted as far as they can be): the claim is an ASSOCIATION, not a decomposition. What may be said is: among pairs carrying the same rule-class sound relation to the fixed word, the ones that also look alike are registered more often here — whose causes this design cannot separate into orthography and residual within-N2phonetic distance. The handbook consequence is still real and still narrow: a translator choosing between twoN2candidates has a reason to prefer the spelled-alike one. §7.26's word registered stands as written. - |Δ_ORTH − Δ_CONS| < 0.10 → spelling contributes nothing measurable at a phrase-end.
A4(finding 5, accepted): this does NOT discharge §7.26 item 5. Under the sound-framed prompt it narrows the hedge to when a reader is attending to sound; whether it narrows further is what stageNanswers, and the handbook may go no further than the neutral figure allows. - Between 0.10 and 0.15, or an interval spanning zero → reported as indeterminate, no handbook change.
P2— is a pararhyme registered at all? Δ_CONS ≥ +0.20. New: §7.26'sNEARclass pooledN1,N2andN3and no figure exists forN2alone. This is the number a translator writing road / blade actually needs.P3— the ITEM-SET sensitivity check (A8, round-1 critic finding 9, accepted, and its status is downgraded). Δ_ORTH − Δ_CONS computed onQRalone.QRis not a power statement aboutP1andP2— a different model reading differently proves nothing about theirs. It asks a narrower question that it can answer: can these materials produce an orthographic advantage in any seat, including one this project has measured reading by eye? IfQRreturns ≥ +0.15 while the reporting seats return < 0.10, the item set is shown able to express the effect. IfQRreturns < 0.10 too, the run reports that the item set produced no orthographic effect in the one further seat tested — one seat, one outcome, one contrast (A18, round-2 finding 11) — which is evidence about the items and is stated as such, never as a null about reading.P4— WITHDRAWN as a claim (A2). TheNEAR-EYEfigure overE1–E3is printed besideRS-20260822bstage P's 0.722 so the two can be read together, and nothing rests on it.P5— strength verdicts. Thefull/halfsplit by level, descriptive, as §7.26 item 4.P6— bare assent (A5, round-1 critic finding 6, accepted in part). TheANSWER: yesrate, reported beside the naming-requiredhit, so that said yes and found the right pair are separable. The finding's stronger form — that visual string-matching rather than heard chime produces the effect — is the effect under test, not a confound, and is refused as one.
7. Controls
C1naturalness, TWO-SIDED (A6, round-1 critic finding 7, accepted). From stage R. It fires ifaXandaVdiffer in mean naturalness rank by more than 0.5 in EITHER direction — awkwardness can make a phrase-end salient just as fluency can make a relation easy to accept.A15(round-2 critic finding 6, accepted): ifC1fires, the primary is reported as disclosed-confounded and NO handbook sentence is written from it — a fired confound control may not coexist with a positive conclusion. ThatC1does not cover thebcells, rates side-by-side rather than in isolation, and uses two of the responding seats as raters, is a cost decision and is a named limit.C2mechanical. Every planted relation recomputed fromtools/rhyme_pairs.pybyverify.pyand the 2 × 4 structure re-checked cell by cell: exactly oneN2-run≥3 per locus ataX, oneN2-run≤2 ataV, oneSTRICTataW, every other cellNONE. Any mismatch and the run does not go.C3unplanted phrase-end relations — a LIMITED diagnostic (A18, round-2 finding 11). It catches unplanted relations between adjacent phrase-end words and nothing else: not internal phonological priming, not semantic association, not shared morphology, not template attention. Every carrier is segmented by the rule ofsweep.pyand every unplanted adjacent phrase-end pair scored. Loci carrying one are named, and a registered sensitivity analysis recomputes every criterion with them dropped.C4filler shape, descriptive only — it establishes nothing about equivalence. Mean filler length in characters and syllables by cell.C6endpoint identification, added free byA18(round-2 finding 9). The rate at which each cell's answer names the two planted phrase-end words, printed per cell, so that a failure of thebbaseline or of endpoint identification in any cell is visible rather than assumed away.C5floor.hitover the fourbcells andaY.
8. Failure criteria — what would make this run worthless
M1fails. No finding is reported.- Floor saturation: mean
hitover the fourbcells andaY> 0.60. Every Δ is then compressed against a ceiling and the run is withheld. C1fires onaX. The primary is disclosed as possibly confounded by wording.P3returns < 0.10 as well. The primary is reported as no orthographic effect in any seat, including a measured spelling-matcher — a statement about the item set — and never as a null about how readers read.- Dead cells > 5% of stage D. Reported with the caveat on every figure.
9. Contamination
The lead's English of the Gulistan span is the translation limb and no figure in this study is
computed from it except the descriptive M-avail census, which is a hand-level figure and is never
summed with the seats'. The carriers of §3 are constructed English written for this design and render
no source. The lead is not the independent third translator of anything here.
The measurement was run after the log froze, per CLAUDE.md: tools/dependence_check.py, the lead's
whole English of tales 20–25 against Eastwick 1852's text of the same six tales (Internet Archive
OCR, $0). 2 shared 7-grams, 0 shared 12-grams, longest common run 7 tokens — clean. Both
figures are reported, per the standing rule that run length alone is a poor proxy.
10. Budget
Pre-flight, built from max_tokens and not from an assumed output length (note (abc)). Stage D
bodies are one sentence in and one short answer out.
| stage | calls | worst case |
|---|---|---|
| pre-run critic, 2 rounds | 2 | $0.20 |
| D — detection, 130 × 3 | 390 | $1.95 |
N — neutral framing, 56 × 2 (A4) |
112 | $0.55 |
| R — naturalness, 17 × 2 | 34 | $0.30 |
| retry allowance | — | $0.40 |
| declared ceiling | 538 | $3.40 |
Per-call rates are read from a 6-call probe before the run, never from config/models.md
(S216). Declared de-scope, in order — RE-ORDERED BY A17 (round-2 critic finding 10, accepted). Stage
N was first to go and is now last, because without it no claim beyond the exact sound-framed
instruction may be made at all. If the probe puts the run above the ceiling: (1) drop QR from
the 18 NEAR-EYE items, which carry no claim; (2) cut stage R to P1 alone, with C1
computed from one seat and reported as single-seat on the face of the control; (3) drop QR
from stage D entirely, losing P3; (4) only then drop stage N — and if it is dropped,
every claim is restricted to the exact sound-framed instruction and no sentence about ordinary or
neutral reading is written. Nothing else changes: the estimands, criteria and interpretations above
hold in every de-scoped version.
11. Warrant, to be carried on the face of every instruction (A11)
n constructed loci in one language, k model seats, no human reader, Tier D NOT PASSED. Two
deterministic seats are not a sample of readers; the bootstrap resamples hand-built loci and creates
no population-level support for a claim about readers. Every sentence this run puts into
framework/v0.2 carries that line, in the form §7.26 already uses.